Railway safety detection method based on multi-modal sensor
Through multimodal sensor fusion technology, the limitations of driver visual judgment in railway transportation are solved, accurate identification and timely alarm of rail foreign objects are achieved, and the safety and reliability of railway transportation are improved.
Patent Information
- Application Number
- CN202510467213.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, there are limitations in relying on drivers to visually judge track safety in railway transportation, resulting in fatigue and attention reduction, and it is difficult to identify safety hazards at long distances or at night in a timely manner.
Using multimodal sensors combined with deep learning models, a variety of data is obtained and fusion processed through telephoto cameras, short-focus cameras, infrared cameras and range-finding radars, identify orbital foreign objects and divide warning levels.
It realizes accurate identification and timely alarm of rail foreign objects in complex environments, and improves the safety and reliability of railway transportation, especially in limited line of sight or complex weather conditions.
Smart Images

Figure CN120372356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of track safety detection, and particularly to a railway safety detection method based on multimodal sensors. Background Art
[0002] In the railway transportation system, ensuring track safety is of utmost importance. The driver needs to be highly concentrated during driving to judge that there are no sundries on the railway line interfering with the driving safety of the train. However, relying solely on the driver's visual judgment has many limitations. Long-time driving can easily cause driver fatigue, resulting in a decline in their reaction speed and attention. During long-distance driving, the driver may experience visual fatigue and mental slackness. Even under normal sight conditions, some less obvious potential safety hazards may be missed. Moreover, the sight is limited at night and it is not easy to observe at a long distance, making dangerous events extremely likely to occur. Summary of the Invention
[0003] The present invention aims to solve the deficiencies of the prior art and provides a railway safety detection method based on multimodal sensors. By using a multimodal visual detection scheme, it effectively makes up for the deficiencies of a single detection method, avoids the situation where it is difficult for the human eye to distinguish clearly at a long distance and the light is insufficient at night, and can accurately identify and detect track foreign objects, better ensuring the safety of railway transportation.
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] A railway safety detection method based on multimodal sensors, and the specific detection steps are as follows:
[0006] S1. Start the media service, and then start the program;
[0007] S2. After the media server is started, start the long-focus camera program, short-focus camera program, infrared camera program, and ranging radar program, and start the detection function;
[0008] S3. Obtain the long-focus camera target recognition result, short-focus camera target recognition result, infrared camera target recognition and track segmentation result, and radar ranging result respectively;
[0009] S4. Integrate the target detection and track segmentation results of the long-focus camera, short-focus camera, and infrared camera in step S3 and the ranging result of the ranging radar;
[0010] S5. Distinguish different alarm levels for the integration result in step S4 according to the warning regulations, and display the result after dividing the warning levels.
[0011] In step S1, when starting the media service, initialize the media service. The initialization settings include network parameter configuration, data storage path setting, and communication protocol initialization with other devices.
[0012] In step S3, the target recognition algorithm is based on a deep learning model. Perform target detection algorithm processing on the images taken by three cameras to obtain recognition results for identifying target objects on the railway; obtain the segmentation results of the track line through semantic segmentation algorithm processing.
[0013] To improve the accuracy of the track line segmentation and the reliability of the area division, preprocess the collected images. Use a detail enhancement algorithm to increase the contrast between the edge information of the track line and the background, making the track edge feature information more abundant. That is, the enhancement algorithm is:
[0014] Among them, f(x,y) is the original image, g(x,y) is the sharpened image, c is a flag parameter, calculate the edge gradient of the sharpened image, and further extract the edge features through the gradient value, as shown in the following formula:
[0015]
[0016] Among them, operator A and operator B are the x and y direction operators of the edge gradient respectively, grande is the gradient, I is the preprocessed image, and φ is the set gradient threshold.
[0017] In step S4, fuse the target detection and track segmentation results of the long - focal - length camera, short - focal - length camera, and infrared camera, as well as the ranging results of the ranging radar, including fusing the data obtained by different sensors to obtain comprehensive detection information. In the fusion process, use a single mapping matrix to calculate the position mapping relationship, as shown in the following formula: p′ i =Hp i ; where, p i is the pixel coordinate point of the long - focal - length camera or the pixel coordinate point of the infrared camera p i =(x i ,y i ,1) T , p′ i is the pixel coordinate point p′ mapped to the short - focal - length interface i =(x′ i ,y′ i ,1) T , H is a 3×3 single mapping matrix Calculate matrix H according to the mapped coordinate points.
[0018] In the step S5 of classifying the warning level according to the detection result, the warning level is classified according to the preset level classification rules based on factors such as the distance, type, and quantity of the detected target.
[0019] The classification of the warning level divides the alarm level according to the segmented track line area. The area inside the track line and a part of the area outside the track line is the alarm area. Different alarm levels are set according to different distances. When the distance between the target and the track is ≤ 50m, it is a first-level alarm, that is, emergency braking is triggered; when it is 50 - 100m, it is a second-level alarm, that is, the driver is prompted; when it is > 100m, it is a third-level alarm, that is, monitoring and recording.
[0020] The beneficial effects of the present invention are as follows: The present invention uses multi-modal data complementary enhancement. The long-focus camera captures small targets at a long distance, combines with the field of view of the short-focus camera to achieve "far-near" space coverage; the infrared camera detects living targets (such as animal or personnel intrusion) based on the thermal radiation characteristics, which is complementary to the visible light information of the optical camera to improve the target recognition rate in night or haze scenes. The data fusion algorithm is used, and the mapping matrix is used to map the detection results of the long-focus camera and the infrared camera to the display interface of the short-focus camera. The present invention is assisted by track segmentation. Based on the semantic segmentation algorithm, the track region ROI is extracted to exclude interference targets in non-track regions. Brief Description of the Drawings
[0021] Figure 1 It is the process framework diagram of the present invention;
[0022] Figure 2 It is the track line segmentation area in the detection area of the present invention;
[0023] Figure 3 It is the track line extension area in the detection area of the present invention;
[0024] Figure 4 It is the actual detection result of the present invention.
[0025] Hereinafter, the embodiments of the invention will be described in detail with reference to the drawings. Detailed Embodiments
[0026] The present invention will be further described below with reference to the drawings and embodiments:
[0027] A railway safety detection method based on multi-modal sensors, and the specific detection steps are as follows:
[0028] Step S1: The program starts, and the media server starts.
[0029] Initialize the media server, configure the network communication protocol, ensure the real-time transmission and storage of sensor data, and transmit the target video collected by the backend to the front-end display through the media server. As the core hub of the system, the media server realizes the unified access, efficient transmission, and reliable storage of multi-modal data, and provides real-time interaction capabilities for the front-end, ensuring full-link traceability and visual monitoring of defect detection results.
[0030] Steps S2 - S3: Start multiple sensor detection programs and output detection results and track line segmentation results.
[0031] Start the long-focus camera program, short-focus camera program, infrared camera program, and ranging radar program, and start the detection function; transmit the video frames to be detected by different sensors into different detection models through the media server, detect the video frames to be detected by different sensors through step S3, and transmit the detection results to the next step.
[0032] To improve the accuracy of track line segmentation and the reliability of area division, preprocess the collected images, and use a detail enhancement algorithm to improve the contrast between the edge information of the track line and the background, making the track edge feature information more abundant. The enhancement algorithm is as follows:
[0033]
[0034] Among them, f(x, y) is the original image, g(x, y) is the sharpened image, c is a flag parameter, calculate the edge gradient of the sharpened image, and further extract the edge features through the gradient value, as shown in the following formula:
[0035]
[0036] Among them, operator A and operator B are the x and y direction operators of the edge gradient respectively, grande is the gradient, I is the preprocessed image, and φ is the set gradient threshold. Figures 2 - 3 is the track line segmentation result, where Figure 2 is the track line segmentation result, Figure 3 is the track line extended area.
[0037] Steps S4 - S5: Fusion of multi-sensor detection results and division of alarm levels.
[0038] Fuse the detection results of the long-focus camera and the infrared camera into the detection results of the short-focus camera, that is, draw the result bounding boxes of the long-focus camera and the infrared camera on the display interface of the short-focus camera. The fusion process uses a single mapping matrix to calculate the position mapping relationship, as shown in the following formula:
[0039] p′ i =Hp i
[0040] Among them, p i is the pixel coordinate point of the long - focal camera or the pixel coordinate point p of the infrared camera i =(x i , y i , 1) T , p' i is the pixel coordinate point p' mapped to the short - focal interface i =(x' i , y' i , 1) T , H is a 3×3 single - mapping matrix Calculate the matrix H according to the mapping coordinate points. The fusion result is as Figure 4 shown. The inspector draws a label box, where the label box contains different warning - level numbers and corresponding label - box colors, and contains the distance information of the detection target from the locomotive head, giving the driver intuitive detection information through the color and alarm - level information.
[0041] Embodiment 1: Monitoring along railway lines
[0042] The present invention is applicable to railway special lines and freight yard scenarios. Currently, during the operation of railway special lines and freight yards, the monitoring of the front - running situation mainly relies on the driver's direct observation. However, this method has many limitations. On the one hand, long - time driving by the driver is likely to cause visual fatigue, and in the face of complex weather conditions such as heavy rain, thick fog, sand and dust, the line of sight is blocked, making it difficult to clearly and accurately observe the road conditions ahead. On the other hand, the railway special - line and freight - yard environments are complex, with many visual blind spots. Relying solely on the driver's naked eye observation, it is very likely to miss some potential safety hazards, such as sudden failures of loading and unloading equipment in the freight yard, foreign object intrusion on the railway special line, etc. The railway safety detection method based on multi - modal sensors of the present invention emerges under this background. In the railway special line, by installing devices such as high - definition cameras and millimeter - wave radars at the front end of the train, the high - definition camera captures the image information of the front track and the surrounding environment in real time, and the millimeter - wave radar monitors the distance, speed and other data between the train and the front objects in real time.
[0043] Embodiment 2: Mine transportation
[0044] In the mine transportation environment, ensuring the safe and unobstructed track is crucial for the efficient operation of the entire mine and the safety of personnel. However, the complex and changeable conditions in the mine may lead to potential safety hazards such as foreign objects on the track and unauthorized entry of personnel, and it is difficult for traditional detection methods to comprehensively and timely detect these potential risks. Based on this, the multi-modal safety detection method of the present invention plays a key role in the mine transportation scenario. Based on the fused comprehensive feature vector, combined with the safety standards and actual situation of mine transportation, a safety judgment model is established. For example, when the size of a foreign object exceeds a certain threshold and the distance from the vehicle is less than the safe distance, it is determined that there is a serious safety hazard; when a person is detected on the track and the vehicle is about to approach, it is also determined to be a dangerous situation. By analyzing and judging each feature in the comprehensive feature vector, it is determined whether there are safety threats such as foreign objects and personnel on the track, and the level of the threat is evaluated.
[0045] The invention has been described above in an exemplary manner with reference to the accompanying drawings. Obviously, the specific implementation of the invention is not limited by the above methods. As long as various improvements are made by adopting the method concept and technical solution of the invention, or directly applied to other occasions without improvement, they are all within the protection scope of the invention.
Claims
1. A railway safety detection method based on multi-modal sensors, characterized in that, The specific detection steps are as follows: S1. Start the media service and then start the program; S2. After the media server starts, start the long - focal - length camera program, short - focal - length camera program, infrared camera program, and ranging radar program; S3. Obtain the target recognition results of the long - focal - length camera, short - focal - length camera, target recognition and track segmentation results of the infrared camera, and the ranging results of the radar respectively; S4. Integrate the target detection and track segmentation results of the long - focal - length camera, short - focal - length camera, and infrared camera, and the ranging results of the ranging radar in step S3; S5. Distinguish different alarm levels for the integration result in step S4 according to the warning regulations, and display the result after dividing the warning level.
2. The railway safety detection method based on multimodal sensors according to claim 1, wherein, In step S1, while starting the media service, initialize the media service. The initialization settings include network parameter configuration, data storage path setting, and initialization of communication protocols with other devices.
3. A railway safety detection method based on a multi-modal sensor according to claim 1, characterized in that, In step S3, the target recognition algorithm is based on a deep - learning model. The target detection algorithm is used to process the images taken by the three cameras to obtain recognition results for identifying target objects on the railway; the track line segmentation results are obtained by semantic segmentation algorithm processing.
4. The railway safety detection method based on a multi-modal sensor according to claim 3, wherein, To improve the accuracy of the track line segmentation and the reliability of the area division, the collected images are preprocessed, and a detail enhancement algorithm is used to increase the contrast between the edge information of the track line and the background, making the track edge feature information richer. That is, the enhancement algorithm is as follows: Among them, f(x,y) is the original image, g(x,y) is the sharpened image, c is a flag parameter. Calculate the edge gradient of the sharpened image, and further extract the edge features through the gradient value, as shown in the following formula: Among them, operator A and operator B are the x and y direction operators of the edge gradient respectively, grande is the gradient, I is the pre - processed image, and φ is the set gradient threshold.
5. A railway safety detection method based on a multi-modal sensor according to claim 4, characterized in that, In the step S4, the object detection and orbit segmentation results of the long-focus camera, short-focus camera, and infrared camera and the ranging results of the ranging radar are fused, including fusing the data obtained by different sensors to obtain comprehensive detection information. In the fusion process, a single mapping matrix is used to calculate the position mapping relationship, as shown in the following formula: p′ i = Hp i ; Among them, p i is the pixel coordinate point of the long - focal camera or the pixel coordinate point p of the infrared camera i =(x i , y i , 1) T , p' i is the pixel coordinate point p' after being mapped to the short - focal interface i =(x' i , y' i , 1) T , and H is a 3×3 single - mapping matrix Calculate the matrix H according to the mapped coordinate points 6. The railway safety detection method based on a multi-modal sensor according to claim 5, wherein, In step S5, for the step of dividing the warning level of the detection result, according to the factors of the distance, type, and quantity of the detected target, the warning level is divided according to the preset level - division rules.
7. A railway safety detection method based on a multi-modal sensor according to claim 6, characterized in that, The division of the warning level divides the alarm level according to the segmented track line area. The area inside the track line and a part of the area outside the track line is the alarm area. Different alarm levels are set according to different distances. When the distance between the target and the track ≤50m, it is a first - level alarm, that is, emergency braking is triggered; when it is 50 - 100m, it is a second - level alarm, that is, the driver is prompted; when it is >100m, it is a third - level alarm, that is, monitoring and recording.