Band-type brake analysis method and system

By improving the brake analysis method and system of the YOLOv8m network, the horn rod, chain and brake cylinder head in the railway image are detected, the risk of brake is judged and early warning is issued, which solves the problem of stoppage, conflict and derailment caused by the brake shoe holding the wheels in the railway, improves the detection accuracy and efficiency, and ensures railway safety.

CN120032328AActive Publication Date: 2025-05-23LIAONING QIHUI ELECTRONIC SYST ENG CO LTD

Patent Information

Application Number
CN202510499156.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

During the railway operation, the vehicle is carrying the gate due to the locking of the brake shoe during the hump slip, causing accidents such as stoppage, conflict and derailment, which affects the operation efficiency and poses safety hazards.

Method used

A brake analysis method and system based on the improved YOLOv8m network is adopted. By obtaining the railway incoming vehicle images, axle count data and vehicle number tags, combined with the space pyramid pooling module, part self-attention module and coordinate attention mechanism, the horn rod, chain and brake cylinder head in the image are detected and analyzed, and the brake risk type is judged and the early warning is issued.

Benefits of technology

Real-time dynamic detection and early warning of railway locking risks has been achieved, detection accuracy and efficiency have been improved, stoppages, conflicts and derailment accidents in railway operations have been significantly reduced, and train operations have been ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032328A_ABST
    Figure CN120032328A_ABST
Patent Text Reader

Abstract

The invention discloses a band-type brake analysis method and system, and relates to the technical field of railway band-type brake detection. The method comprises the following steps: acquiring a railway coming train image, axle counting data and a train number label; determining carriage number information according to the axle counting data and the number label; performing target detection on the railway coming vehicle image based on a target detection model of a first improved YOLOv8m network, and calculating the length of a piston rod when the image contains the piston rod and a brake cylinder head; when the length of the piston rod exceeds a preset length threshold value, a wind embracing risk early warning is given out; when the image contains a chain, performing chain segmentation on the chain detection frame area image based on an instance segmentation model of a second improved YOLOv8m-Seg network; processing the chain segmentation result to obtain the chain bending degree; and when the bending degree is greater than a preset bending threshold value, sending out chain embracing risk early warning. The band-type brake detection precision reaches the millimeter level, peak-crossing humping of a band-type brake vehicle can be effectively prevented, and train operation safety is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of railway brake detection, and in particular to a brake analysis method and system. Background Art

[0002] During the railway hump sliding operation, if the brake shoes hold the wheels and brake, the vehicle will operate with the brakes on. The sliding process of the vehicle with brakes will cause it to stop on the way, or even collide or derail, which will seriously affect the railway operation efficiency and may even cause safety accidents.

[0003] Intelligent identification of railway brakes is an important technology in the field of railway transportation safety. Braking refers to the situation when the hump is dismantled and the vehicle to be released is not ventilated as required, and the manual brake machine is not released as required, so that all the brake shoes of the vehicle are still close to the wheel tread. The brakes will cause the vehicle to stop on the way, leading to frequent accidents such as collisions and derailments. The brakes used by railway freight cars to stop and control speed during operation are divided into two categories: one is air braking and the other is manual braking. Air braking is to inject air into the air cylinder to push the piston so that the bellows rod extends and drives the brake shoe to hold the wheel, and the brake shoe rubs against the wheel to perform railway braking; manual braking is to manually turn the manual brake machine, so that the manual brake chain pulls the brake shoe to hold the wheel, and the brake shoe rubs against the wheel to perform railway braking. Wind hold refers to the situation where, when a vehicle is slid out during a hump operation, the air is not exhausted as required, causing the brake shoes to hold the wheels, causing the vehicle to operate with the brakes on. Chain hold refers to the situation where, when a vehicle is slid out during a hump operation, the manual brake is not released as required, causing the brake shoes to hold the wheels, causing the vehicle to operate with the brakes on.

[0004] In recent years, with the continuous emergence of new models in the field of deep learning, vehicle intelligent detection technology based on computer vision has also made great progress. In the field of object detection, mainstream algorithms are mainly divided into two categories: (1) Single-stage model: This type of model simplifies the object detection task into a regression and classification problem, and directly predicts the category and bounding box of the object by inputting the image. Due to its simple structure, the single-stage model has a high detection speed and is suitable for real-time detection scenarios and devices with limited computing power. (2) Two-stage model: Unlike the single-stage model, the two-stage model divides the detection task into two stages. The first stage generates candidate regions, and the second stage further classifies and regresses these candidate regions. By gradually refining the target features, the model performs better in small target detection and complex scenes. However, the multi-stage calculation characteristics make the model structure more complex, which correspondingly increases the computational overhead.

[0005] Therefore, it is worth studying how to use deep learning to dynamically detect and warn of railway brake risks in real time. Summary of the invention

[0006] In view of the above problems, the present invention proposes a brake analysis method and system to try to solve or alleviate one or more of the above problems.

[0007] According to one aspect of the present invention, a brake analysis method is provided, the method comprising: Obtain the collected railway vehicle images, axle counting data, and vehicle number labels; Determine the carriage number information based on the axle counter data and the carriage number label; Processing and analyzing the railway vehicle image based on the improved YOLOv8m network to obtain analysis results; including: performing target detection on the railway vehicle image based on the target detection model of the first improved YOLOv8m network; analyzing the detected target to determine the brake risk type; Display the image processing results, analysis results and corresponding carriage number information.

[0008] Furthermore, the improvements of the first improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; the targets include: a chain, a bellows rod or a brake cylinder head.

[0009] Furthermore, the target detection model based on the first improved YOLOv8m network performs target detection on the railway vehicle image, including: inputting the railway vehicle image into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local image features and perform feature fusion; then, aggregating features of different scales through the spatial pyramid pooling module, including: processing features of different scales through convolutional layers, batch normalization and SiLU activation functions to obtain feature maps with half the number of channels; passing through multiple maximum pooling layers in sequence to gradually reduce the size of the feature map; splicing the feature maps after pooling of different scales in the channel dimension; then, through a partial self-attention module, evenly dividing the feature map aggregated by the spatial pyramid pooling module into two parts, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between query vectors, key vectors and value vectors; the other part of the feature map is connected through jump connections The neck network is used to fuse, enhance and process the features extracted by the backbone network; the coordinate attention mechanism is used to generate a feature map with enhanced position information, including: first, performing global average pooling on the features output by the neck network in the horizontal direction and the vertical direction respectively, and extracting compressed features in different directions; then, these features are fused through a set of convolutions with shared weights; the fused features are re-divided into two parts in the horizontal direction and the vertical direction, and the attention matrix is ​​generated by 1×1 convolution and Sigmoid activation function respectively; finally, the features output by the neck network are weighted by the attention matrix through element-wise multiplication to obtain a feature map with enhanced position information; the head network is used to convert the feature map with enhanced position information into the final target detection result.

[0010] Furthermore, the analysis of the detected target to determine the brake risk type includes: when the detected target includes a bellows rod and a brake cylinder head, determining whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; when it is determined to be a valid bellows rod, calculating the length of the bellows rod; comparing the length of the bellows rod with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold When a chain is detected, a chain embrace risk warning is issued; when the detected target contains a chain, the chain detection frame area image is extracted, and the chain detection frame area image is segmented based on the instance segmentation model of the second improved YOLOv8m-Seg network to obtain a chain mask image; for the chain mask image, the parabola fitting algorithm is used to calculate the chain centerline, and the ratio of the chain arc depth to the chain chord length is used to evaluate the chain bending degree; when the bending degree is greater than the preset bending threshold, a chain embrace risk warning is issued.

[0011] Furthermore, the improvements of the second improved YOLOv8m-Seg network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m-Seg backbone network; adding a multi-scale sequence fusion module between the backbone network and the neck network; adding a feature fusion module between the neck network and the head network; the instance segmentation model based on the second improved YOLOv8m-Seg network performs chain segmentation on the chain detection frame area image, including: inputting the chain detection frame area image into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion ; Then, the features of different scales are aggregated through the spatial pyramid pooling module, including: processing the features of different scales through convolutional layers, batch normalization and SiLU activation functions to obtain feature maps with half the number of channels; passing through multiple maximum pooling layers in sequence to gradually reduce the size of the feature maps; splicing the feature maps after pooling at different scales in the channel dimension; then, through the partial self-attention module, the aggregated feature maps are evenly divided into two parts, one part of the feature maps enters the self-attention module to perform global information modeling through matrix operations between query vectors, key vectors and value vectors; the other part of the feature maps are fused with the output of the self-attention module through jump connections; The features extracted by the backbone network are fused with a multi-scale sequence fusion module, including: extracting feature maps of levels P2 to P5 from the backbone network, preliminarily fusing low-level local detail information and high-level global semantic information through element splicing after resizing, and generating a comprehensive description of the target features; then, channel dimension reduction is performed through 1×1 convolution, and the features are mapped to a higher-level feature space; the edge details and context information of the target are extracted under different receptive fields using the RepVGG module with convolution kernels of 3 and 5; the features extracted by the backbone network are fused, enhanced and processed using the neck network; the features extracted by the multi-scale sequence fusion module are fused, enhanced and processed using the feature fusion module The obtained fusion features are fused with the feature map of the P3 level of the neck network, including: splicing the two parts of the feature map in the channel dimension; generating a feature weight set through 3×3 convolution and Sigmoid activation function; adjusting the features to be fused according to the feature weight set; fusing the adjusted two parts of the features through element addition; using the head network to convert the feature map fused by the feature fusion module into the final instance segmentation result; wherein the head network includes a segmentation head and a prediction head, the segmentation head is used to generate a high-resolution native mask of the target, and the prediction head is used to predict target attributes, and the target attributes include the target box position, category information, and mask coefficients related to segmentation.

[0012] Furthermore, the formula for evaluating the degree of bending of the chain according to the ratio of the chain arc depth to the chain chord length is: ; in, represents the curvature of the chain centerline, represents the chord length of the center line of the chain, Indicates the maximum depth of the centerline of the chain compared to the chord length.

[0013] Furthermore, the method of determining the carriage number information according to the axle counting data and the carriage number label includes: a railway train includes wheels of four axles, and the distance between the wheel of the first axle and the wheel of the second axle, the distance between the wheel of the second axle and the wheel of the third axle, and the distance between the wheel of the third axle and the wheel of the fourth axle are different; the axle counting data includes the time when each wheel passes through the magnetic steel A and the time when each wheel passes through the magnetic steel B; the speed of each wheel is calculated according to the time difference and the spacing between the magnetic steel A and the magnetic steel B; the time difference when the magnetic steel B of two adjacent wheels passes through is multiplied by the average speed of the two wheels to obtain the distance between the two wheels; the distance is matched with the distance between the wheels of each axle to determine the axle numbers of the first axle, the second axle, the third axle and the fourth axle of the railway carriage and the time when they pass through the axle counting magnetic steel B, thereby determining the passing time of the railway carriage; the carriage number information is determined according to the carriage passing time and the carriage number label.

[0014] According to another aspect of the present invention, a brake analysis system is proposed, which includes a data acquisition end, a service end and a user end; the data acquisition end includes a camera component, an axle counter magnetic steel component and a radio frequency identification module; the camera component is used to collect railway vehicle images; the axle counter magnetic steel component is used to collect axle counter data; the radio frequency identification module is used to collect vehicle number labels; the service end includes an image processing module, a brake analysis module, and a car matching module; the image processing module is used to process railway vehicle images, including a target detection submodule, the target detection submodule is used to perform target detection on railway vehicle images based on a target detection model of a first improved YOLOv8m network, the target including a chain, a bellows rod or a brake cylinder head; the brake analysis module is used to perform brake analysis according to the output result of the image processing module to determine the brake risk type; the car matching module is used to determine the car number information according to the axle counter data and the car number label; the user end includes a display module, the display module is used to display the output result of the image processing module, the analysis result output by the brake analysis module, and the corresponding car number information output by the car matching module.

[0015] Furthermore, the improvements of the first improved YOLOv8m network in the target detection submodule include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; the image processing module also includes a target segmentation submodule, and the target segmentation submodule is used to extract the chain detection box area image when the detection result of the target detection submodule contains a chain, and perform chain segmentation on the chain detection box area image based on the instance segmentation model of the second improved YOLOv8m-Seg network; wherein the improvements of the second improved YOLOv8m-Seg network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m-Seg backbone network; adding a multi-scale sequence fusion module between the backbone network and the neck network; between the neck network and the head network. Add a feature fusion module; the brake analysis module includes a wind brake analysis submodule and a chain brake analysis submodule; the wind brake analysis submodule is used to determine whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head when the output result of the target detection submodule is that the railway vehicle image contains a bellows rod and a brake cylinder head; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is within a certain pixel range around the bellows rod; when it is determined to be When the bellows rod is effective, the length of the bellows rod is calculated; the length of the bellows rod is compared with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, a wind hug risk warning is issued; the chain hug analysis submodule is used to calculate the segmentation result output by the target segmentation submodule using a parabola fitting algorithm to obtain the chain centerline, and evaluate the degree of bending of the chain according to the ratio of the chain arc depth to the chain chord length; when the bending degree is greater than the preset bending threshold, a chain hug risk warning is issued.

[0016] Furthermore, the data acquisition end also includes a PLC control unit, which is used to control the train caisson equipment; the service end also includes: a data storage module, a user management module, and an equipment control module; the data storage module is used to store the collected data and data processing results; the user management module is used to manage user-related information; the equipment control module is used to perform data docking and control with the PLC control unit; the user end also includes an alarm module, which is used to send an alarm signal when the analysis result is a wind risk warning or a chain risk warning.

[0017] The beneficial technical effects of the present invention are: The present invention proposes a brake analysis method and system. In the data acquisition part, the vehicle information is automatically identified through automatic vehicle detection and automatic door opening and closing caisson control, thereby realizing automated operation; image data is automatically collected after the vehicle arrives, providing high-quality data input for subsequent processing; in the detection part, YOLOv8m is used as the core algorithm, and customized improvements are made according to task requirements; in the detection output part, by comparing with the set threshold, it is determined whether there is a potential risk, and an early warning prompt is issued in time. Among them, the detection part focuses on target detection, and adds a spatial pyramid pooling module and some self-attention modules to the original YOLOv8m backbone network; adds a coordinate attention mechanism after the neck network and before the head network to meet the high requirements for precise positioning in the task. By adaptively adjusting the feature weights, the model can focus on the target boundary more accurately, thereby achieving efficient detection and positioning of the chain; for target segmentation, for the special form of the handbrake chain of the human brake machine, adds a spatial pyramid pooling module and some self-attention modules to the original YOLOv8m-Seg backbone network; adds a multi-scale sequence fusion module between the backbone network and the neck network; adds a feature fusion module between the neck network and the head network. By effectively fusing global context information and position information, the model's ability to analyze complex forms at different scales is enhanced, thereby achieving accurate segmentation of the target.

[0018] Thanks to the above optimization and improvement, the accuracy of brake detection of the present invention has reached the millimeter level, among which the recognition rate of bellows rod is as high as 99.9%, and the recognition rate of hand brake chain has also reached 96%. Compared with the traditional manual or low-precision detection scheme, the present invention significantly improves the detection accuracy and efficiency, and reaches the leading level in accuracy and automation; the present invention is applied to railway brake detection, through real-time detection and dynamic early warning, it can effectively prevent the braked vehicles from slipping during peak hours, thereby reducing the labor intensity of railway operators, improving work efficiency, and ensuring the safety of train operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, in which: Figure 1 is a flow chart of a brake analysis method according to an embodiment of the present invention; Figure 2 1 is an example diagram of the distance between the hump-pushing operation carriage and each axle of the front of the vehicle in an embodiment of the present invention; Figure 3 It is a structural schematic diagram of a target detection model based on a first improved YOLOv8m network in an embodiment of the present invention; Figure 4It is a structural schematic diagram of an instance segmentation model based on the second improved YOLOv8m-Seg network in an embodiment of the present invention; Figure 5 is a business flow chart of brake analysis in an embodiment of the present invention; Figure 6 It is a structural schematic diagram of a brake analysis system according to an embodiment of the present invention; Figure 7 It is a system architecture diagram of a brake analysis system described in an embodiment of the present invention. DETAILED DESCRIPTION

[0021] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.

[0022] The embodiment of the present invention provides a brake analysis method, such as Figure 1 As shown, the method includes: S1, acquiring the collected railway vehicle image, axle counting data, and vehicle number label; S2, determining the vehicle number information of the carriage according to the axle counting data and the vehicle number label; S3, processing and analyzing the railway vehicle image based on the improved YOLOv8m network to obtain the analysis result; specifically including: S31. Perform target detection on the collected railway vehicle image based on the target detection model of the first improved YOLOv8m network, wherein the target includes a chain, a bellows rod or a brake cylinder head; S32. Analyze the detected target to determine the brake risk type; including: S321. When the image contains a bellows rod and a brake cylinder head, determine whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head; the judgment standard for a valid bellows rod is: the brake cylinder head in the image is within a certain pixel range around the bellows rod; S322. Calculate the length of the bellows rod when it is determined to be a valid bellows rod; compare the bellows rod length with a preset length threshold, and when the bellows rod is When the bellows length exceeds the preset length threshold, a wind hug risk warning is issued; S323, when the image contains a chain, the chain detection frame area image is extracted, and the chain detection frame area image is segmented based on the instance segmentation model of the second improved YOLOv8m-Seg network to obtain a chain mask image; S324, for the chain mask image, the chain centerline is calculated using a parabola fitting algorithm, and the degree of bending of the chain is evaluated based on the ratio of the chain arc depth to the chain chord length; when the degree of bending is greater than the preset bending threshold, a chain hug risk warning is issued; S4, the image processing results, analysis results and the corresponding carriage number information are displayed.

[0023] The method starts from S1. In S1, the collected railway vehicle image, axle counting data, and vehicle number label are obtained.

[0024] According to the embodiment of the present invention, when a train arrives on the railway, the starter magnet detects the train arrival signal and notifies the PLC control unit to control the caisson equipment to open the cover. At the same time, the car number detection device (RFID) starts to read the car number, the video service starts to collect video, and the axle counting magnet starts to count the axles to obtain the axle counting data. When the carriage passes the detection camera, the detection image will be pulled.

[0025] Then S2 is executed, in which the carriage number information is determined according to the axle counting data and the car number label, including: the railway train includes wheels of 4 axles, and the distance between the wheel of the first axle and the wheel of the second axle, the distance between the wheel of the second axle and the wheel of the third axle, and the distance between the wheel of the third axle and the wheel of the fourth axle are different; the axle counting data includes the time when each wheel passes through the magnetic steel A and the time when each wheel passes through the magnetic steel B; the speed of each wheel is calculated according to the time difference and the spacing between the magnetic steel A and the magnetic steel B; the time difference when the magnetic steel B of two adjacent wheels passes through is multiplied by the average speed of the two wheels to obtain the distance between the two wheels; the distance is matched with the distance between the wheels of each axle to determine the axle numbers of the first axle, the second axle, the third axle and the fourth axle of the railway carriage and the time when they pass through the axle counting magnetic steel B, thereby determining the passing time of the railway carriage; the carriage number information is determined according to the carriage passing time and the car number label.

[0026] According to the embodiment of the present invention, the number of wheels of the carriage and the front of the train are fixed, and the range of the distance difference between each wheel is also within a fixed range, and there is no overlapping distance range. Therefore, the distance between the two wheels can be obtained by multiplying the time when the vehicle wheel presses the magnet by the speed. The number of axle counts of the carriage passing the magnet can be calculated by the direction, distance and current axle count. Specifically, two axle count magnets are installed on the rails, namely magnet A and magnet B. When the wheel passes through magnet A, the time of magnet A is obtained, and when the wheel passes through magnet B again, the time of magnet B is obtained. If the order of the vehicle passing is from magnet A to magnet B, it means that it is driving forward, otherwise it is driving in the reverse direction; the speed of the current wheel is obtained by (magnet B time-magnet A time) / the distance between the two magnets, and the distance between the two vehicles can be obtained by multiplying (magnet B time of the second wheel-magnet B time of the first wheel) by the average speed of the two wheels.

[0027] As an example, Figure 2As shown, the wheel distance parameters of ordinary carriages are as follows: railway freight carriages have four wheels, and the distance from the first wheel to the second wheel is between 1 meter and 2 meters; the distance from the second wheel to the third wheel is between 4.5 meters and 20 meters; the distance from the third wheel to the fourth wheel is between 1 meter and 2 meters; the distance from the fourth wheel to the first axle wheel of the next car is between 2 meters and 4.5 meters. The wheel distance parameters of train locomotives are as follows: the locomotive has 6 wheels, and the distance from the first wheel to the second wheel is between 1 meter and 2 meters; the distance from the second wheel to the third wheel is between 1 meter and 2 meters; the distance from the third wheel to the fourth wheel is between 4.5 meters and 20 meters; the distance from the fourth wheel to the fifth wheel is between 1 meter and 2 meters; the distance from the fifth wheel to the sixth wheel is between 1 meter and 2 meters.

[0028] The business logic of the train carriage wheel axle counting is as follows: During the railway peak pushing operation, the carriage and the locomotive will push over the axle counting magnet. Therefore, when a new carriage arrives, the current axle count is 6, and the peak pushing operation interval is generally more than 10 minutes. Therefore, when the first axle of the carriage presses over the axle counting magnet, the calculated wheelbase must be greater than 45,000 mm. When the first axle of the carriage presses over the magnet, because the distance between the 6th axle of the previous locomotive and the 1st axle of the carriage is greater than 45,000 mm, the axle count needs to be +1, because the current axle count is the 6th axle of the previous locomotive, so the current axle count +1 is 1. When the second axle of the carriage presses over the magnet, because the wheelbase between the 1st and 2nd axles of the carriage is between 1,000 mm and 2, the axle count needs to be +1, because the current axle count is 1, so the current axle count +1 is 2. When the third axle of the carriage passes through the magnetic steel, because the wheelbase between the 2nd and 3rd axles of the carriage is between 4500 mm and 20000 mm, the axle count needs to be +1, because the current axle count is 2, so the current axle count +1 is 3. When the fourth axle of the carriage passes through the magnetic steel, because the wheelbase between the 3rd and 4th axles of the carriage is between 1000 mm and 2000 mm, the axle count needs to be +1, because the current axle count is 3, so the current axle count +1 is 4. When the first axle of the second carriage passes through the magnetic steel, because the wheelbase between the 4th axle of the carriage and the 1st axle of the next carriage is between 2000 mm and 4500 mm, the axle count needs to be +1, because the current axle count is 4, so the current axle count +1 is 1. In this way, when the wheels of each carriage pass through the magnetic steel, the axle count data of each carriage will be calculated, and the carriages can be cut apart section by section according to the time when the 1st and 4th axles pass through the magnetic steel.

[0029] The business logic of the train locomotive wheel axle counting is as follows: During the railway peak operation, when the 4th axle of the last carriage passes the axle counting magnet, the current axle count is 4. When the 1st axle of the locomotive presses over the magnet, because the distance between the 4th axle of the carriage and the 1st axle of the locomotive is between 2000 mm and 4500 mm, the current axle count is +1, because the current axle count is the 4th axle of the previous carriage, so the current axle count is 1. When the 2nd axle of the locomotive presses over the magnet, because the wheelbase between the 1st and 2nd axles of the carriage is between 1000 mm and 2000 mm, the axle count needs to be +1, because the current axle count is 1, so the current axle count +1 is 2. When the 3rd axle of the locomotive presses over the magnet, because the wheelbase between the 2nd and 3rd axles of the carriage is between 1000 mm and 2000 mm, the axle count needs to be +1, because the current axle count is 2, so the current axle count +1 is 3. When the 4th axle of the front of the train presses over the magnetic steel, because the wheelbase between the 3rd and 4th axles of the carriage is between 4500 mm and 20000 mm, the axle count needs to be +1, because the current axle count is 3, so the current axle count +1 is 4. When the 5th axle of the front of the train presses over the magnetic steel, because the wheelbase between the 4th and 5th axles of the carriage is between 1000 mm and 2000 mm, the axle count needs to be +1, because the current axle count is 4, so the current axle count +1 is 5. When the 6th axle of the front of the train presses over the magnetic steel, because the wheelbase between the 5th and 6th axles of the carriage is between 1000 mm and 2000 mm, the axle count needs to be +1, because the current axle count is 5, so the current axle count +1 is 6. In this way, after the 6th axle of the front of the train passes the magnetic steel, the peak pushing operation is completed.

[0030] The vehicle parking and reversing business processing is as follows: After the vehicle stops, the axle count and axle count time of the previous wheel remain unchanged. When the vehicle stops and starts, when the axle count magnet is pressed, the direction can be obtained according to the order of pressing axle count magnet A and axle count magnet B. If it is in the forward direction, the axle count +1 will continue to process the car axle count business logic; if it is in the reverse direction, the axle count -1 will continue to process the car axle count business reverse logic, that is, the car axle count business is the current axle count +1 becomes the current axle count -1, if it is 1 axle, the next axle count starts from 4 and processes the axle count business in the order of 4, 3, 2, 1. The abnormal data correction process is as follows: When the current vehicle axle is pressed and rubbed repeatedly on two magnets, it may cause confusion in the axle count data of one car. If the vehicle is traveling in the forward direction, when the 4th wheel of the current car passes the magnet to the 1st wheel of the next car, if the distance between the two wheels is between 2000 mm and 4500 mm, the current axle count is forcibly modified to 1, and the subsequent vehicles passing by are recounted according to normal driving, so as to ensure that the axle count of the vehicle behind the disordered car is calculated correctly. If the vehicle is traveling in the reverse direction, when the 1st wheel of the current car passes the magnet to the 4th wheel of the next car, if the distance between the two wheels is between 2000 mm and 4500 mm, the current axle count is forcibly modified to 4, and the subsequent vehicles passing by in reverse are recounted in the reverse direction according to normal driving, so as to ensure that the axle count of the vehicle behind the car is calculated correctly.

[0031] Then, S3 is executed. In S3, the railway vehicle image is processed and analyzed based on the improved YOLOv8m network to obtain the analysis result.

[0032] According to an embodiment of the present invention, first, in S31, the target detection model based on the first improved YOLOv8m network performs target detection on the collected railway vehicle image to determine whether there is a chain, a bellows rod or a brake cylinder head in the image. Among them, the process of training the target detection model based on the first improved YOLOv8m network includes: S311, obtaining a training data set. The training and testing of the target detection model uses a self-made data set, which is aggregated from the field collection data of multiple stations and divided into a training set and a test set in a ratio of 8:2. The data set is finely annotated by a professional team according to strict annotation standards to ensure the high quality of the data and the accuracy of the annotation. In the data collection process, the actual environmental differences of different stations are fully considered to enhance the model's adaptability to diverse scenes. For example, the collection tasks cover a variety of light conditions, including strong light during the day, low light at night, and shadow areas; the diversity of equipment; and the complexity of the background, including raindrops, leaves and debris occlusion, etc. S312, preprocessing the training data. In order to further improve the generalization ability of the model for diverse scenes, data enhancement operations are performed on the training data. These enhancements include rotation, flipping, scaling, brightness adjustment, etc., so that the model can adapt to complex actual environments. S313: Input the preprocessed training data set into the target detection model based on the first improved YOLOv8m network for training to obtain a trained target detection model.

[0033] In order to balance the accuracy and inference speed of the model, YOLOv8m, which is more mature in technology among single-stage models, is selected as the basic framework of the model. As a member of the YOLO series, YOLOv8m is also composed of three networks: backbone, neck and head. Specifically, the backbone network effectively captures the local and global information of the image by stacking the C2f feature extraction module designed with multiple gradient flows; the neck network adopts the idea of ​​path aggregation network, and extracts effective information from multiple levels of feature maps by contracting and expanding the path, thereby enhancing the detection ability of small objects and multi-scale targets; the head network is responsible for completing the final target detection task, including category prediction, bounding box regression, and target confidence estimation.

[0034] In complex scenes, since the background and the target share similar texture or color features, the feature extraction process of the model is easily disturbed, resulting in false detection and missed detection. In order to better meet the detection requirements of the task target, the target detection model proposed in this paper is targetedly improved and optimized based on YOLOv8m. The specific architecture of the improved model is as follows Figure 3As shown in the figure, the improved model structure still consists of three parts: the backbone, the neck and the head. The improvements include: introducing the spatial pyramid pooling module and the partial self-attention module in the fourth stage of the backbone network; introducing the coordinate attention mechanism in the third stage of the neck network. The purpose of this is to further enhance the model's ability to accurately locate the target boundary, so that it can better adapt to the actual needs of the system for detection accuracy while maintaining efficient reasoning speed.

[0035] The backbone network is responsible for extracting multi-level and multi-scale universal features from the input image to form a feature map with high-level semantic information to provide support for subsequent detection tasks. First, the C2f feature extraction module, as the basic unit of the network, is responsible for extracting local features at each stage and enhancing the expression of multi-scale information through feature fusion. This module adopts the CSP structure design and divides the input feature map into two branches through 1×1 convolution: one branch is connected through multiple residual modules. The first branch is processed, and the other branch directly performs convolution operations. Subsequently, the outputs of the two branches are spliced ​​through the Concat operation, and after batch normalization (BN) and SiLU activation function processing, the features are finally sorted through the convolution operation to obtain the final output. By stacking multiple C2f modules, the network can learn rich feature expressions at different levels, while optimizing the fusion of high-level semantic information and low-level detail information at different stages. The formula of the C2f feature extraction module can be expressed as follows:

[0036] in, is the input feature, is the output feature, represents 1×1 convolution, Represents the channel splitting operation, is the residual module, Represents a feature concatenation operation. It is the intermediate feature of the module processing process and is the output obtained after N times of residual module processing.

[0037] Subsequently, the spatial pyramid pooling module (Spatial Pyramid Pooling-Fast, SPPF) aggregates features of different scales, so that the network can better retain semantic information when processing complex scenes and generate richer feature representations. The SPPF module first processes the input feature map through convolutional layers, batch normalization (BN) and SiLU activation functions to obtain a feature map with half the number of channels. Then, the feature map passes through three 5×5 maximum pooling layers in sequence, and the size of the feature map gradually decreases after each pooling. The output of each pooling layer will be used as the input of the next pooling. This serial pooling process can effectively capture information at different scales and improve the model's ability to capture details and global information in the image. Finally, by splicing the feature maps after pooling at different scales in the channel dimension, the SPPF module can integrate feature information from different scales to form a multi-scale feature representation. The formula of the spatial pyramid pooling module can be expressed as follows:

[0038] in, is the input feature, is the output feature of the module, represents a 5×5 maximum pooling operation, Represents the output obtained after k times of maximum pooling.

[0039] When the model separates the target and background features, it often inevitably introduces unnecessary background information to interfere with the target features, resulting in false detection and missed detection. To solve this problem, the usual practice is to introduce an attention mechanism in the deep layer of the model to highlight the target features. However, the existing mainstream attention mechanisms are mostly implemented through convolution operators. The inherent limitations of convolution operators make it difficult to establish long-distance dependencies between features, which is precisely the key to accurately locate the target. In contrast, the Transformer structure can efficiently capture global long-range dependencies with the self-attention mechanism. However, the computational complexity and memory usage of the Transformer are high, which often brings a large time overhead in real-time reasoning tasks. Therefore, a partial self-attention module (PSA) is introduced into the backbone network. The design of the PSA module aims to effectively enhance the network's ability to model long-distance dependencies while avoiding the computational overhead caused by the global self-attention mechanism. Specifically, the input feature map is evenly divided into two parts, so that the complexity of the self-attention calculation is controlled within a low range. Part of the feature map is input into the self-attention module for global information modeling, capturing long-distance dependencies and enhancing the contextual understanding of the target. Through matrix operations between the query vector (Query), key vector (Key) and value vector (Value), key features are extracted to capture the long-range correlation between features; the other part of the feature map is fused with the output of the self-attention module through jump connections. In addition, in order to further improve the reasoning efficiency, the PSA module optimizes the dimensions of the query vector and key vector in the self-attention mechanism, setting their dimensions to half of the value vector, thereby reducing the amount of calculation. At the same time, the module uses batch normalization (BN) instead of layer normalization (LN) for standardization, thereby improving the running speed and stability of the model. This module is placed after the fourth stage with the lowest resolution in the model, focusing on extracting and combing key information in high-level abstract features. Due to the low feature resolution of the fourth stage, the quadratic complexity of the self-attention calculation is significantly reduced, so the overall reasoning speed can still meet the real-time requirements. The formula of some self-attention modules can be expressed as follows:

[0040] in, is the input feature, is the output feature, Represents the Transformer module. Through the above operations, the long-range dependencies between features can be effectively captured without significantly increasing the computational overhead, so as to better process the semantic information in the deep layer of the network.

[0041] The neck network is used to integrate the features extracted from the backbone network for fusion, enhancement and processing to help the model better detect objects of different scales. Specifically, the neck network receives the features of the P3, P4 and P5 levels in the backbone network as input, first enhances the low-level features through the bottom-up information transmission path, and then fuses them with the high-level features to ensure that the network can obtain rich multi-level feature representations. Subsequently, the network diffuses the detail information to the deeper levels of the network through the top-down information transmission path, enhances the interaction between high-level semantic features and low-level detail features, and further improves the model's sensitivity to details. Combining bottom-up and top-down information flow ensures that the network still maintains efficient detection performance in complex scenes, especially showing significant advantages in multi-scale object detection tasks. In order to further improve the model's ability to extract target location information, the coordinate attention mechanism (CA) is introduced in the output part of the third stage of the neck network and before the head network to enhance the ability to accurately locate the target boundary. The coordinate attention mechanism extracts compressed features in different directions by performing global average pooling in the horizontal and vertical directions respectively; then, these features are subjected to information interaction and integration through a set of convolutional layers with shared weights to improve the expressiveness of the features; the fused features are re-divided into two parts in the horizontal and vertical directions, and the corresponding attention matrices are generated through 1×1 convolution and Sigmoid activation function respectively; finally, the original features are weighted by the attention matrix through element multiplication to enhance the position information in the feature map, effectively improving the model's positioning accuracy for the target boundary. The formula of the coordinate attention mechanism can be expressed as follows:

[0042] in, is the input feature, is the output feature, and Represents the average pooling operation in the horizontal and vertical directions of the feature, is the Sigmoid activation function.

[0043] The head network is responsible for converting the feature maps processed from the trunk and neck into the final target detection results. Specifically, the network contains 3 detection heads, each of which can be divided into two parts, each consisting of 2 3×3 convolutions and one 1×1 convolution, which are used to predict coordinate regression information and category confidence information respectively.

[0044] Furthermore, in order to prevent the model from overfitting, regularization techniques such as weight decay and Dropout were introduced during the model training process to improve the robustness of the model.

[0045] Then, in S32, the detected target is analyzed to determine the type of brake risk; first, S321, when the image contains a bellows rod and a brake cylinder head, it is determined whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head; the judgment criterion for a valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; S322, when it is determined to be a valid bellows rod, the length of the bellows rod is calculated; the length of the bellows rod is compared with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, a wind brake risk warning is issued.

[0046] Specifically, the brake cylinder of a railway freight car is the main component of railway brakes. The air pressure is converted into mechanical thrust through the action of the brake valve, and the brake pads are pushed to press the wheels to produce a braking effect. When braking, the bellows rod will be pushed out of the brake cylinder under the action of air pressure. If the bellows rod and the brake cylinder head are detected, it is further determined whether the brake cylinder head is within 50 pixels around the bellows rod. If this condition is met, the subsequent calculation will continue; otherwise, the detection of this frame will end. When the positional relationship between the bellows rod and the brake cylinder meets the requirements, it is determined whether there is a complete brake cylinder head in the image according to the aspect ratio of the brake cylinder. If a complete brake cylinder head is detected, the brake cylinder size corresponding to the bottom component of the vehicle model is matched according to the model output, and the length of the bellows rod is automatically calibrated and calculated. Since the diameter of the brake cylinder head is a known fixed value, the actual physical length corresponding to a single pixel can be calculated by the model output result. By further analyzing the number of detected bellows rod pixels, its length can be accurately calculated. The specific formula is:

[0047] in, and are the actual lengths of the bellows rod and the diameter of the brake cylinder head, respectively. and The length of the bellows rod and the diameter of the brake cylinder head are the pixel values, respectively. Subsequently, it is determined whether the length exceeds the set threshold, and a notification is sent to the main program based on the judgment result. If it exceeds, an early warning notification is issued. If the complete brake cylinder head is not detected, the detection of the current frame will end.

[0048] Furthermore, for the analysis results of multiple frames of images, the following validity verification can be formulated according to the position and proportion of the brake cylinder and the bellows rod: 1) At least 3 air cylinder data verification: Because the speed of railway freight cars during pushing operations is relatively slow, multiple images will be generated when the train brake cylinder passes through the chain detection equipment. Setting 3 consecutive brake cylinder verifications can remove false alarm data caused by misidentification; 2) Valid range verification: The bellows rod is generally more than 50 mm, which is greater than 50 pixels when converted into pixels. The image recognition area is defined and the maximum range of the image is set to be reduced by 50 pixels. For example, the image pixels , the valid range is 50 to 1230 in width and 50 to 670 in height. If the bellows rod appears in the valid range, it is deemed as valid data, otherwise it is invalid data, which can prevent other edge devices from identifying it as a bellows rod problem; 3) Verification of the ratio of the brake cylinder to the bellows rod: The width of the bellows rod is generally less than 1 / 2 of the air cylinder. If the width of the bellows rod is less than 1 / 2 of the air cylinder, it is valid data, otherwise it is invalid data, which can prevent the problem of the bellows rod being out of proportion to the air cylinder. 4) Verification of the overlap ratio of the brake cylinder and the bellows rod: The air cylinder and the bellows rod are connected together and will not overlap. The overlap ratio is set not to exceed 50% of the bellows rod area, which can effectively prevent the problem of fuzzy identification of the brake cylinder and the bellows rod; 5) Verification of the proportion of valid data: In a carriage, the number of bellows rods exceeding the threshold / the total number of brake cylinder data identified accounts for more than 50% as valid data to prevent problems caused by angle recognition of the bellows rod being too long; 6) Threshold verification: When the above verification is passed, the longest data of the bellows rod is obtained and compared with the set threshold of 65 pixels. If the bellows rod length exceeds 65, it is a warning data, and a voice and page warning is issued. That is: for the processing results of multiple frames of images, determine whether to issue a wind risk warning based on the above rules.

[0049] Then, S323, when the image contains chains, the chain detection box area image is extracted, and the chain detection box area image is chain segmented based on the instance segmentation model of the second improved YOLOv8m-Seg network; wherein, the process of training the instance segmentation model based on the second improved YOLOv8m-Seg network includes: S3231, obtaining a training data set; S3232, preprocessing the training data; S3233, inputting the preprocessed training data set into the instance segmentation model based on the second improved YOLOv8m-Seg network for training, and obtaining a trained instance segmentation model.

[0050] Since the local area of ​​slender targets is extremely narrow and occupies only a small number of pixels, it is easily obscured by complex background interference during feature transfer. In addition, the aspect ratio of slender targets is much larger than that of conventional targets. Limited by the receptive field of the detector, it is difficult to generate complete target features, which affects the model's accurate segmentation of slender targets. In view of the special morphology of slender targets, the instance segmentation model proposed in this paper is improved and optimized based on the YOLOv8m-Seg network. The specific model structure is as follows: Figure 4As shown. Specifically, a spatial pyramid pooling module and a partial self-attention module are added to the original YOLOv8m-Seg backbone network; a multi-scale sequence fusion module is added between the backbone network and the neck network; and a feature fusion module is added between the neck network and the head network. The multi-scale sequence fusion module generates a global information representation with more details by aggregating feature maps of different scales in the backbone network, thereby enhancing the network's sensitivity to target edge details and morphological changes; in the third stage of the neck network, the global context information is effectively fused with the position information through the feature fusion module and passed down layer by layer, so that the downstream segmentation head can use more accurate information to complete fine segmentation.

[0051] The backbone network plays a core role in the task. It generates feature maps with high-level semantic information by extracting multi-level and multi-scale common features from the input image, providing a solid foundation for subsequent tasks. First, the C2f feature extraction module, as the basic unit of the network, is responsible for extracting local features at each stage and enhancing the expression of multi-scale information through feature fusion. Subsequently, the features from different scales are aggregated through the spatial pyramid pooling module to enhance the network's semantic understanding ability when processing complex scenes. The SPPF module first processes the input feature map through convolutional layers, batch normalization (BN) and SiLU activation functions to obtain a feature map with half the number of channels. Then, the feature map passes through three 5×5 maximum pooling layers in sequence, and the size of the feature map is gradually reduced after each pooling. The output of each pooling layer is used as the input of the next pooling. This serial pooling process can effectively capture information at different scales and improve the model's ability to capture details and global information in the image. Finally, by splicing the feature maps after pooling at different scales in the channel dimension, the SPPF module can integrate feature information from different scales to form a multi-scale feature representation. A partial self-attention module is introduced in the backbone network. The input feature map is evenly divided into two parts. One part of the feature map is input into the self-attention module for global information modeling to capture long-distance dependencies and enhance the contextual understanding of the target. Global information modeling is performed through matrix operations between query vectors, key vectors, and value vectors to capture long-range correlations between features; the other part of the features is fused with the output of the self-attention module through jump connections.

[0052] In order to further improve the segmentation accuracy, a multi-scale sequence fusion module (MSF) is designed to generate global context information by aggregating multi-scale features, effectively compensating for the defects caused by the limitation of the receptive field. Specifically, this module extracts feature maps of levels P2 to P5 from the backbone network, and after resizing, it preliminarily fuses low-level local detail information and high-level global semantic information through element splicing to generate a comprehensive description of the target features. To avoid information loss during excessive upsampling or downsampling of features, feature maps of all levels are adjusted to the same size as the P3 level feature map. Subsequently, channel dimension reduction is performed through 1×1 convolution, and the features are mapped to a higher-level feature space to simplify the computational complexity and retain key information. On this basis, the RepVGG module with convolution kernels of 3 and 5 is used to extract the edge details and context information of the target under different receptive fields, further enhancing the expression ability of the global context and ensuring that the network can accurately capture the complete features of the target under complex backgrounds. The formula of the multi-scale sequence fusion module can be expressed as follows:

[0053] in, represents the feature map of the i-th level, is the output feature of the module, Represents the RepVGG module.

[0054] The neck network consists of a bottom-up expansion path and a top-down contraction path, which is used to integrate, enhance and process the features extracted from the backbone network to help the model better detect objects of different scales. The neck network receives the features of the P3, P4 and P5 levels in the backbone network as input. It first enhances the low-level features through a bottom-up expansion path, and then fuses them with high-level features to ensure that the network can obtain rich multi-level feature representations. The neck network diffuses detail information to deeper levels of the network through a top-down contraction path, enhances the interaction between high-level semantic features and low-level detail features, and further improves the model's sensitivity to details. The combination of bottom-up and top-down information flow ensures that the neck network shows significant advantages in the segmentation of slender targets.

[0055] Subsequently, the feature fusion module (FFM) is used to fuse the feature maps from the multi-scale sequence fusion module and the P3 level of the neck network expansion path to enhance the network's accurate segmentation ability. Specifically, the module first splices the two parts of the feature map in the channel dimension to preliminarily achieve information complementarity. Then, a feature weight set is generated through 3×3 convolution and Sigmoid activation function to reflect the importance of each feature. According to the feature weight set, the features to be fused are adjusted separately to enhance the response of key features and suppress redundant information. Finally, the two parts of the features are fused by element-wise addition to ensure that the final output has a stronger feature expression ability, thereby improving the network's segmentation effect on slender targets. The formula of the feature fusion module is:

[0056] in, and is the feature to be fused, is the output feature, is the Sigmoid activation function.

[0057] The head network is responsible for converting the feature map fused by the feature fusion module into the final instance segmentation result. The head network consists of a segmentation head and three prediction heads. The main task of the segmentation head is to generate a high-resolution native mask of the target. It consists of multiple convolutional layers, which gradually extract and restore spatial information to generate accurate segmentation results. First, the input feature map is processed using 3×3 convolution to extract spatial information and enhance the detail expression of the feature. Subsequently, the feature map is upsampled by transposed convolution to restore the spatial resolution of the image, thereby retaining the fine-grained information of the target. Finally, the channel is reduced in dimension by a 1×1 convolution layer to generate a high-resolution target mask. The prediction head is designed for feature maps at different levels, mainly used to predict the attributes of the target, including the location of the target box, category information, and mask coefficients related to segmentation. First, the perception of the target area is enhanced by 3×3 convolution, and then the target category classification, bounding box regression, and mask coefficient prediction are completed by 1×1 convolution. Finally, the segmentation mask of the target is generated based on the output of the segmentation head and the prediction head.

[0058] Then, S324, the chain mask image is calculated using a parabola fitting algorithm to obtain the chain centerline, and the degree of bending of the chain is evaluated based on the ratio of the chain arc depth to the chain chord length; when the degree of bending is greater than a preset bending threshold, a chain hugging risk warning is issued.

[0059] Specifically, the brake chain of railway freight cars is the main component of railway manual brakes. The brake chain is stirred by the winch, so that the brake chain is tightened and the brake pads are pushed to press the wheels to produce a braking effect. When braking, the brake chain at the bottom of the carriage will be tightened and the curvature of the brake chain will become smaller. The chain segmentation result is approximated by a parabola, and the curvature of the brake chain is calculated based on the ratio of the arc depth of the brake chain to the chord length of the chain. By analyzing the segmentation results, the curvature calculated by the segmentation results is compared with the preset threshold, and the curvature change is used to identify potential risk situations. All segmented images under the same regression frame are regarded as the same target to avoid the influence of occlusion. In order to accurately obtain the geometric characteristics of the target, a parabola fitting algorithm based on the least squares method is used to obtain the centerline of the chain: Then, the curvature of the centerline of the chain is calculated using a mathematical derivation method to quantitatively reflect the degree of deformation of the chain. The specific formula is:

[0060] in, is the curvature of the chain centerline, represents the chord length of the center line of the chain, Indicates the maximum depth of the centerline of the chain compared to the chord length.

[0061] Furthermore, for the analysis results of multiple frames of images, the validity judgment rules of chain-lock detection data are designed as follows: 1) At least 3 brake chain data verification: Because the speed of railway freight cars during pushing operations is relatively slow, multiple pictures will be generated when the train brake chain passes through the chain-lock detection equipment. Setting 3 consecutive brake chain verifications can remove false alarm data caused by misidentification. 2) Valid data ratio verification: In a carriage, the number of brake chains with a curvature greater than 0.06 (relatively curved) / the number of brake chains accounts for more than 25% of the total number of brake chains, which is invalid data; the number of brake chains with a curvature less than 0.035 (relatively straight) / the number of brake chains accounts for less than 25% of the total number of brake chains, which prevents false alarms when the brake chain is relatively straight at a certain angle and relatively curved at other viewing angles. 3) Threshold verification: When the above verification is passed, obtain the picture with the longest brake chain chord length, use the brake chain curvature to compare with the threshold of 0.035, and issue a voice and page warning if it is less than the threshold. That is, for the processing results of multiple frames of images, determine whether to issue a chain lock risk warning based on the above rules.

[0062] Then, S4 is executed, in which the image processing result, the analysis result and the corresponding carriage number information are displayed.

[0063] The following is an example of a complete process. Figure 5As shown, when a car comes, the starter magnet detects the incoming signal and notifies the PLC control unit to control the caisson equipment to open the cover. At the same time, the vehicle number detection device starts to read the vehicle number, the video service starts video acquisition, and the axle counting magnet starts axle counting. When the carriage passes the detection camera, the detection image will be pulled and the brake detection service will perform identification and detection. If the image does not contain the detection target, it will exit and proceed to the next image detection. If the target is detected, YOLOv8m is used as the core algorithm of the detection task to perform intelligent detection on the brake cylinder, bellows rod and brake chain. The brake cylinder is detected according to the brake cylinder detection model. Partial self-attention modules and coordinate attention mechanisms are introduced to cope with the high requirements for precise positioning in the task. By adaptively adjusting the feature weights, the model can focus on the target boundary more accurately, match the brake cylinder size of the corresponding vehicle bottom component for automatic calibration, and calculate the length of the bellows rod; the brake chain is detected according to the chain holding detection model, and the ROI area is obtained according to the model output and used as the input of the segmentation model. The obtained segmentation results will be approximated by a parabola, and the curvature of the chain will be evaluated based on the ratio of the chain arc depth to the chain chord length. If the test result meets the set threshold, the wind-holding bellows threshold is set to 65, and the chain-holding curvature threshold is set to 0.035. A message is sent to the main program, and the test result is pushed to the brake detection business unit. The detection unit summarizes the submitted test results according to the axle counting data of the axle counting magnet, and simultaneously obtains the vehicle number data uploaded by the vehicle number recognition device in the time period from axle 1 to axle 4 of the carriage as the current carriage number. After the detection of the same carriage is completed, the current carriage data is checked for validity. More than three brake cylinder images are tested for valid data, and the proportion of brake chain curvature greater than 0.06 is less than 25% and the proportion of curvature less than 0.035 is greater than 75%. If there is no valid data, the non-warning pictures and videos are directly saved, and the data is pushed to the large screen for display. If the data is valid and exceeds the threshold, a warning picture, carriage video and warning voice are generated, and the warning voice is pushed to the hump building and the work site to let the operators stop the car to deal with the brake problem.

[0064] The target detection model and segmentation model in the present invention use YOLOv8m as the core algorithm, and are customized and improved according to task requirements. For target detection, a spatial pyramid pooling module and a partial self-attention module are added to the original YOLOv8m backbone network; a coordinate attention mechanism is added after the neck network and before the head network to cope with the high requirements for precise positioning in the task. By adaptively adjusting the feature weights, the model can focus on the target boundary more accurately, thereby achieving efficient detection and positioning of the chain; for target segmentation, for the special form of the hand brake chain of the human brake machine, a spatial pyramid pooling module and a partial self-attention module are added to the original YOLOv8m-Seg backbone network; a multi-scale sequence fusion module is added between the backbone network and the neck network; a feature fusion module is added between the neck network and the head network, and the model's ability to analyze complex forms at different scales is enhanced by effectively fusing global context information and position information, thereby achieving accurate segmentation of the target.

[0065] Another embodiment of the present invention provides a brake analysis system, such as Figure 6As shown, the system includes: a data acquisition terminal 610, a service terminal 620 and a user terminal 630; the data acquisition terminal 610 includes a camera component 6110, an axle counter magnetic steel component 6120 and a radio frequency identification module 6130; the camera component 6110 is used to collect railway vehicle images; the axle counter magnetic steel component 6120 is used to collect axle counter data; the radio frequency identification module 6130 is used to collect vehicle number tags; the service terminal 620 includes an image processing module 6210, a brake analysis module 6220, and a car matching module 6230; the image processing module 6210 is used to process railway vehicle images, including a target detection submodule 62110 and a target segmentation submodule 62120; the target detection submodule The block 62110 is used to perform target detection on the railway vehicle image based on the target detection model of the first improved YOLOv8m network to determine whether there is a chain, a bellows rod or a brake cylinder head in the image; the target segmentation submodule 62120 is used to extract the chain detection frame area image when the output result of the target detection submodule 62110 is that the image contains a chain, and perform chain segmentation on the chain detection frame area image based on the instance segmentation model of the second improved YOLOv8m-Seg network; the brake analysis module 6220 is used to perform brake analysis according to the output result of the image processing module 6210 to determine whether to perform a brake warning, including the wind lock analysis submodule 62210 and the chain lock analysis submodule 62210. Analysis submodule 62220; wind embrace analysis submodule 62210 is used to determine whether the bellows is a valid bellows based on the positional relationship between the bellows and the brake cylinder head when the output result of the target detection submodule 62110 is that the railway vehicle image contains a bellows and a brake cylinder head; the judgment standard of the valid bellows is: the brake cylinder head in the image is located within a certain pixel range around the bellows; when it is determined to be a valid bellows, the bellows length is calculated; the bellows length is compared with a preset length threshold, and when the bellows length exceeds the preset length threshold, a wind embrace risk warning is issued; the chain embrace analysis submodule 62220 is used to use parabola simulation to the chain segmentation result output by the target segmentation submodule 62120 The chain centerline is obtained by the combined algorithm, and the bending degree of the chain is evaluated according to the ratio of the chain arc depth to the chain chord length; when the bending degree is greater than the preset bending threshold, a chain lock risk warning is issued; the car matching module 6230 is used to determine the car number information according to the axle counter data and the car number label; the user end 630 includes a display module 6310 and an alarm module 6320; the display module 6310 is used to display the output result of the image processing module 6210 and the analysis result output by the brake analysis module 6220, as well as the corresponding car number information output by the car matching module 6230; the alarm module 6320 is used to issue an alarm signal when the analysis result is a wind lock risk warning or a chain lock risk warning.

[0066] In this embodiment, preferably, the data acquisition end 610 also includes a PLC control unit 6140, which is used to control the train caisson equipment; the server end 620 also includes: a data storage module 6240, a user management module 6250, and an equipment control module 6260; the data storage module 6240 is used to store the collected data and data processing results; the user management module 6250 is used to manage user-related information; and the equipment control module 6260 is used to perform data docking and control with the PLC control unit 6140.

[0067] In this embodiment, preferably, the improvements of the first improved YOLOv8m network in the target detection submodule 62110 include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; the improvements of the second improved YOLOv8m-Seg network in the target segmentation submodule 62120 include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m-Seg backbone network; adding a multi-scale sequence fusion module between the backbone network and the neck network; adding a feature fusion module between the neck network and the head network.

[0068] According to an embodiment of the present invention, Figure 7As shown in the figure, the business functions of the system include: 1) Core functions: used to manage the registration management, log management, database connection, etc. of sub-services; 2) Image recognition service: build a recognition model through deep learning, parse the camera video stream, identify the extension length of the brake cylinder piston stroke, the hand brake chain of the manual brake machine, and the bending degree of the brake cylinder chain, and set the width ratio of the brake cylinder and the bellows rod within 0.5; the brake cylinder and the bellows rod are in the same centerline position; the bellows rod is within 50 pixels of the edges of the image and more than 3 brake cylinder images are detected; the proportion of the curvature of the brake chain greater than 0.06 is less than 25% and the proportion of the curvature less than 0.035 is greater than 75% as valid data rules, combined with the set thresholds of windage 65 and chainage 0.035, it is judged as a brake vehicle if the threshold is exceeded; 3) Equipment control service: data docking with the PLC control unit, which can be checked through the PLC control unit View and control each hardware device; 4) Video acquisition service: Automatically capture the video stream of the camera, and generate car videos and train videos according to the vehicle; 5) File storage service: Provide brake warning pictures and video files for storage and access functions; 6) Vehicle number recognition service: Through the vehicle number recognition hardware device, RFID is used to identify the bottom label of the vehicle, and the vehicle number, model and other information are parsed to achieve automatic matching with the detected vehicle, providing basic data for real-time warning and data statistics; 7) Warning service: When a vehicle is coming, after the wheel presses over the starting magnetic steel, the PLC obtains the signal of the incoming vehicle, notifies the system to open the caisson cover, and the image service starts the detection work. During the detection process, the wind-locked and chain-locked cars that exceed the threshold after image analysis are pushed to the large screen page for warning, and the warning recognition picture is displayed on the large screen page, and the alarm position is marked, so that the Camel Peak shift staff can intuitively view the detection results. According to the car and the detection type, the intelligent synthesis voice is pushed to the business site for voice broadcast. After hearing the voice on site, the vehicle can stop and use the intercom to communicate with the Camel Peak shift staff to confirm whether the brake is engaged. If the brake is engaged, the brake is processed. The system can also be configured with real-time reminder devices such as voice alarm lights and large screens at the station that can display early warning data to remind on-site personnel to pay attention to running vehicles; 8) Authority verification: Filter the station hump data according to authority, allowing different managers to see the data within their authority. The authority is divided into administrator authority, station administrator authority, and hump duty officer authority.Administrators have all system permissions and can set system parameters to assign permissions to relevant business personnel, etc.; station administrators can view multiple hump data permissions and access data statistics pages, conduct big data analysis to warn vehicle models and times, etc., and guide and communicate with hump site personnel to pay attention to operating specifications based on data analysis results; hump duty officers only need to view their own hump data permissions, view the warning information in real time, confirm whether to brake, and communicate with on-site operators in real time to handle warning issues; 9) Braking business function: integrate detection data, generate specified WEB API interfaces according to customer needs to facilitate viewing of relevant content on the front-end large screen; 10) Management background: system permissions, personnel division, management thresholds and other information settings can be set; 11) Data visualization large screen: the page displays the latest brake warning data, the latest vehicle passing data, real-time monitoring data, visualized data analysis, historical brake warning data and historical vehicle passing data, etc.

[0069] For the undetailed parts of the brake analysis system according to the embodiment of the present invention, please refer to the above detailed description of the method embodiment.

[0070] The technical effect of the present invention was further verified by experiments. The experiment used multiple data sets collected on site for training and testing. The wind embrace detection data set contains two types of images, the gate cylinder and the bellows rod, totaling 17,251 images; the chain embrace detection data set contains 16,051 images, and the segmentation data set contains 14,832 images. The system model is built based on the PyTorch deep learning framework, the batch size is set to 16, and the SGD optimizer with an initial learning rate of 0.01 and a momentum of 0.937 is used for a total of 300 cycles. In order to ensure the stability and convergence speed of the training process, a learning rate scheduling strategy is adopted to reduce the learning rate by 10 times every 100 cycles to achieve a balance between model exploration and convergence. In order to achieve the optimal detection accuracy, the data enhancement was carefully adjusted and optimized. The final data enhancement parameters of the detection model and the segmentation model are shown in Tables 1 and 2.

[0071] Table 1 Data augmentation parameters of the detection model

[0072] Table 2 Data augmentation parameters of segmentation model

[0073] 1) The experimental verification of the wind embrace detection model is as follows. In order to verify the advanced nature of the detection model, the wind embrace detection model proposed in the present invention is compared with advanced methods in the field, including Faster R-CNN, SSD, YOLOv5m, YOLOv8m, etc. Faster R-CNN is a two-stage target detection model, and the other methods are single-stage target detection models. In order to ensure a fair comparison, the public codes of these methods are used to reproduce their network structures, and the models are trained and evaluated in the same training environment. The experiment used the same hyperparameter settings, data sets, and evaluation indicators to ensure the comparability of the results. As shown in Table 3, the model proposed in the present invention achieved the highest evaluation index in the comparison with other target detection methods, proving its advantage in target detection accuracy.

[0074] Table 3 Comparison of the present invention with other target detection methods

[0075] In order to verify the effectiveness of the introduced modules, an ablation experiment was conducted. The experiment used YOLOv8m as the baseline model, and added partial self-attention modules to the backbone network and coordinate attention mechanisms to the head network in turn to verify the contribution and effectiveness of each module. The experimental results are shown in Table 4. Among them, the partial self-attention module effectively enhanced the model's understanding of the global context and improved the detection accuracy; the coordinate attention mechanism further improved the ability to locate the target position, especially in complex backgrounds. These results show that the design of each module has practical value, and its combination can significantly improve the overall model performance, verifying the effectiveness and rationality of the method.

[0076] Table 4 Experimental results

[0077] In order to improve the inference speed, the detection model was converted from PyTorch to TensorRT to make full use of hardware acceleration. The inference GPU deployed on site is NVIDIA GeForce RTX 4070, which supports FP16 and FP32 calculations, both of which provide 29.15 TFLOPS of computing performance. Through TensorRT optimization, the inference speed of the model has been significantly improved, and the inference time of a single image (including pre- and post-processing) is about 12ms. In order to further improve the processing efficiency, multi-threaded inference is adopted during the deployment process to enhance the concurrent processing capability. Finally, the optimized and deployed Fengbao recognition system can process the image data captured by multiple cameras in real time at a speed of 125 frames per second in actual operation on site, meeting the needs of real-time monitoring and detection. After testing, the Fengbao recognition model can accurately identify and locate targets, quickly determine the state changes of key components, and meet the needs of industrial-grade high-precision detection; the detection accuracy reaches ±5mm, showing its good performance and reliability.

[0078] 2) The experimental verification of the chain hug detection model is as follows. In order to verify the effectiveness of the segmentation model, the instance segmentation model proposed in the present invention is comprehensively compared with a variety of mainstream methods in the field, including Mask R-CNN, YOLACT, YOLOv5m-seg and YOLOv8m-seg. To ensure the fairness and authority of the comparison, the public codes of these methods are used to reproduce their network structures, and the models are trained and evaluated under the same training environment. The experiment strictly unified the hyperparameter settings, data sets and evaluation indicators to ensure the comparability of the results. As shown in Table 5, the proposed segmentation model outperforms other instance segmentation methods in all evaluation indicators, which fully demonstrates its advantages in instance segmentation tasks.

[0079] Table 5 Comparison of the segmentation model of the present invention with other instance segmentation methods

[0080] In order to verify the effectiveness of the introduced modules, an ablation experiment was conducted. The experiment used YOLOv8m-seg as the baseline model, and added the multi-scale sequence fusion module and the feature fusion module in turn on it to verify the contribution and effectiveness of each module. The experimental results are shown in Table 6. Among them, the multi-scale sequence fusion module significantly enriches the diversity of feature expression by modeling and fusing information of different scales, so that the model has higher recognition ability and robustness when dealing with complex scenes or diverse targets. The feature fusion module effectively fuses features from different sources, and transmits key information downward through the contraction path, further enhancing the transmission and representation capabilities of semantic information, thereby optimizing the performance of the model in the segmentation task. These results show that the design of each module has practical value, and its combination can significantly improve the overall model performance, verifying the effectiveness and rationality of the method of the present invention.

[0081] Table 6 Experimental results

[0082] After comprehensive testing, the chain lock identification model has demonstrated excellent performance, with a detection accuracy of over 95%. Regardless of complex environments or changing conditions, the model can efficiently and accurately identify and detect brake risks, providing a strong guarantee for the safe operation of trains.

[0083] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments may still be modified, or some or all of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A brake analysis method, characterized in that: include: Obtain the collected railway vehicle images, axle counting data, and vehicle number labels; Determine the carriage number information based on the axle counter data and the carriage number label; Processing and analyzing the railway train image based on the improved YOLOv8m network to obtain the analysis result; including: performing target detection on the railway train image based on the target detection model of the first improved YOLOv8m network, and the improvements of the first improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; the target includes: a chain, a bellows rod or a brake cylinder head; the target detection model based on the first improved YOLOv8m network performs target detection on the railway train image, including: inputting the railway train image into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; then, aggregating features of different scales through the spatial pyramid pooling module; then, through the partial self-attention module, averaging the feature maps aggregated by the spatial pyramid pooling module The network is evenly divided into two parts, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between the query vector, the key vector and the value vector; the other part of the feature map is fused with the output of the self-attention module through jump connections; the neck network is used to fuse, enhance and process the features extracted by the backbone network; the coordinate attention mechanism is used to generate a feature map with enhanced position information, including: first, performing global average pooling on the features output by the neck network in the horizontal and vertical directions respectively, and extracting compressed features in different directions; then, these features are fused through a set of convolutions with shared weights; the fused features are re-divided into two parts in the horizontal and vertical directions, and the attention matrix is ​​generated by 1×1 convolution and Sigmoid activation function respectively; finally, the features output by the neck network are weighted by the attention matrix through element multiplication to obtain a feature map with enhanced position information; the feature map with enhanced position information is converted into the final target detection result by using the head network; Analyze the detected targets and determine the brake risk type; Display the image processing results, analysis results and corresponding carriage number information.

2. A brake analysis method according to claim 1, characterized in that: Aggregating features of different scales through the spatial pyramid pooling module includes: processing features of different scales through convolutional layers, batch normalization and SiLU activation functions to obtain feature maps with half the number of channels; passing through multiple maximum pooling layers in sequence to gradually reduce the size of the feature map; splicing the feature maps after pooling of different scales in the channel dimension; The dimension of the query vector and the key vector in the partial self-attention module is half of the value vector, and batch normalization is used instead of layer normalization for normalization.

3. A brake analysis method according to claim 1 or 2, characterized in that: The analyzing the detected target to determine the brake risk type includes: when the detected target includes a bellows rod and a brake cylinder head, determining whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; when it is determined to be a valid bellows rod, calculating the length of the bellows rod; comparing the length of the bellows rod with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, A chain hugging risk warning is issued; when the detected target contains a chain, the chain detection frame area image is extracted, and the chain detection frame area image is segmented based on the instance segmentation model of the second improved YOLOv8m-Seg network to obtain a chain mask image; for the chain mask image, a parabola fitting algorithm is used to calculate the chain centerline, and the degree of bending of the chain is evaluated according to the ratio of the chain arc depth to the chain chord length; when the degree of bending is greater than the preset bending threshold, a chain hugging risk warning is issued.

4. A brake analysis method according to claim 3, characterized in that: The improvements of the second improved YOLOv8m-Seg network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m-Seg backbone network; adding a multi-scale sequence fusion module between the backbone network and the neck network; adding a feature fusion module between the neck network and the head network; the instance segmentation model based on the second improved YOLOv8m-Seg network performs chain segmentation on the chain detection box area image, including: The chain detection frame area image is input into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local features of the image and perform feature fusion; then, aggregating features of different scales through the spatial pyramid pooling module, including: processing features of different scales through convolutional layers, batch normalization and SiLU activation functions to obtain feature maps with half the number of channels; passing through multiple maximum pooling layers in sequence to gradually reduce the size of the feature map; splicing the feature maps after pooling at different scales in the channel dimension; then, through the partial self-attention module, the aggregated feature map is evenly divided into two parts, one part of the feature map is global information modeled through matrix operations between query vectors, key vectors and value vectors; the other part of the feature map is fused with the output of the self-attention module through jump connections; The features extracted by the backbone network are fused using a multi-scale sequence fusion module, including: extracting feature maps of levels P2 to P5 from the backbone network, initially fusing low-level local detail information and high-level global semantic information through element splicing after resizing, and generating a comprehensive description of the target features; then, performing channel dimension reduction through 1×1 convolution and mapping the features to a higher-level feature space; using the RepVGG module with convolution kernels of 3 and 5, extracting edge details and contextual information of the target under different receptive fields; The neck network is used to fuse, enhance and process the features extracted by the backbone network; The fusion features extracted by the multi-scale sequence fusion module are fused with the feature map of the neck network P3 level by using the feature fusion module, including: splicing the two parts of the feature map in the channel dimension; generating a feature weight set by 3×3 convolution and Sigmoid activation function; adjusting the features to be fused respectively according to the feature weight set; fusing the adjusted two parts of the features by element addition; The head network is used to convert the feature map fused by the feature fusion module into the final instance segmentation result; wherein the head network includes a segmentation head and a prediction head, the segmentation head is used to generate a high-resolution native mask of the target, and the prediction head is used to predict the target attributes, the target attributes including the target box position, category information and mask coefficients related to segmentation.

5. A brake analysis method according to claim 4, characterized in that: The formula for evaluating the degree of chain bending based on the ratio of the chain arc depth to the chain chord length is: ; in, represents the curvature of the chain centerline, represents the chord length of the center line of the chain, Indicates the maximum depth of the centerline of the chain compared to the chord length.

6. A brake analysis method according to claim 1, characterized in that: Determining the carriage number information according to the axle counting data and the carriage number label includes: A railway train comprises wheels on four axles, and the distances between the wheels of the first axle and the wheels of the second axle, the distances between the wheels of the second axle and the wheels of the third axle, and the distances between the wheels of the third axle and the wheels of the fourth axle are different; the axle counting data comprises the time when each wheel passes through a magnetic steel A and the time when each wheel passes through a magnetic steel B; the speed of each wheel is calculated based on the time difference and the spacing between the magnetic steel A and the magnetic steel B; the time difference when the magnetic steel B of two adjacent wheels passes through is multiplied by the average speed of the two wheels to obtain the distance between the two wheels; the distance is matched with the distance between the wheels of each axle to determine the axle numbers of the first axle, the second axle, the third axle and the fourth axle of the railway carriage and the time when they pass through the axle counting magnetic steel B, thereby determining the railway carriage passing time; the carriage number information is determined based on the carriage passing time and the carriage number label.

7. A brake analysis system, characterized in that: Including data collection end, service end and user end; The data acquisition terminal includes a camera assembly, an axle counter magnetic steel assembly and a radio frequency identification module; the camera assembly is used to collect images of incoming railway vehicles; the axle counter magnetic steel assembly is used to collect axle counter data; the radio frequency identification module is used to collect vehicle number tags; The server includes an image processing module, a brake analysis module, and a carriage matching module; the image processing module is used to process the railway vehicle image, and includes a target detection submodule, and the target detection submodule is used to perform target detection on the railway vehicle image based on the target detection model of the first improved YOLOv8m network, and the target includes a chain, a bellows rod or a brake cylinder head; The improvements of the first improved YOLOv8m network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; the target detection model based on the first improved YOLOv8m network performs target detection on the railway vehicle image, including: inputting the railway vehicle image into the backbone network for feature extraction, including: using multiple C2f feature extraction modules in the backbone network to extract local image features and perform feature fusion; then, aggregating features of different scales through the spatial pyramid pooling module; then, through the partial self-attention module, evenly dividing the feature map aggregated by the spatial pyramid pooling module into two parts, one part of the feature map enters the self-attention module to perform global information modeling through matrix operations between the query vector, the key vector and the value vector; the other part of the feature map is fused with the output of the self-attention module through a jump connection; The features extracted by the backbone network are fused, enhanced and processed by the neck network; a feature map with enhanced position information is generated by using the coordinate attention mechanism, including: first, global average pooling is performed on the features output by the neck network in the horizontal and vertical directions respectively to extract compressed features in different directions; then, these features are fused through a set of convolutions with shared weights; the fused features are re-divided into two parts in the horizontal and vertical directions, and an attention matrix is ​​generated by 1×1 convolution and Sigmoid activation function respectively; finally, the features output by the neck network are weighted by the attention matrix through element multiplication to obtain a feature map with enhanced position information; the feature map with enhanced position information is converted into the final target detection result by using the head network; the brake analysis module is used to perform brake analysis according to the output result of the image processing module to determine the type of brake risk; the car matching module is used to determine the car number information according to the axle counter data and the car number label; The user terminal includes a display module, which is used to display the output result of the image processing module, the analysis result output by the brake analysis module, and the car number information output by the corresponding car matching module.

8. A brake analysis system according to claim 7, characterized in that: The improvements of the first improved YOLOv8m network in the target detection submodule include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m backbone network; adding a coordinate attention mechanism after the neck network and before the head network; The image processing module also includes a target segmentation submodule, which is used to extract a chain detection box area image when the detection result of the target detection submodule contains a chain, and perform chain segmentation on the chain detection box area image based on the instance segmentation model of the second improved YOLOv8m-Seg network; wherein the improvements of the second improved YOLOv8m-Seg network include: adding a spatial pyramid pooling module and a partial self-attention module to the original YOLOv8m-Seg backbone network; adding a multi-scale sequence fusion module between the backbone network and the neck network; adding a feature fusion module between the neck network and the head network; The brake analysis module includes a wind lock analysis submodule and a chain lock analysis submodule; the wind lock analysis submodule is used to determine whether the bellows rod is a valid bellows rod based on the positional relationship between the bellows rod and the brake cylinder head when the output result of the target detection submodule is that the railway vehicle image contains a bellows rod and a brake cylinder head; the judgment standard of the valid bellows rod is: the brake cylinder head in the image is located within a certain pixel range around the bellows rod; when it is determined to be a valid bellows rod, the length of the bellows rod is calculated; the length of the bellows rod is compared with a preset length threshold, and when the length of the bellows rod exceeds the preset length threshold, a wind lock risk warning is issued; the chain lock analysis submodule is used to calculate the segmentation result output by the target segmentation submodule using a parabola fitting algorithm to obtain the chain centerline, and evaluate the degree of bending of the chain according to the ratio of the chain arc depth to the chain chord length; when the bending degree is greater than the preset bending threshold, a chain lock risk warning is issued.

9. A brake analysis system according to claim 7 or 8, characterized in that: The data acquisition end also includes a PLC control unit, which is used to control the train caisson equipment; the service end also includes: a data storage module, a user management module, and an equipment control module; the data storage module is used to store the collected data and data processing results; the user management module is used to manage user-related information; the equipment control module is used to perform data docking and control with the PLC control unit; the user end also includes an alarm module, which is used to send an alarm signal when the analysis result is a wind risk warning or a chain risk warning.

Citation Information

Patent Citations

  • Train brake chain state recognition method and device, equipment and storage medium

    CN114758150A

  • Intelligent pre-detection and alarm system for hump humping vehicle

    CN115402381A

  • Intelligent pre-detection and alarm system for hump humping vehicle

    CN115923875A

  • SAR image aircraft target detection method based on improved YOLOv5

    CN116630798A

  • YOLOv7 pavement garbage detection method based on improved CA attention mechanism

    CN117523493A

Cited By

  • Image detection method for truck brake chain tensioning fault

    CN120496003A