Railway pedestrian warning system and method based on improved YOLOv7 algorithm
Patent Information
- Application Number
- CN202510375655.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2045-03-27
AI Technical Summary
[0004]本发明旨在至少解决现有技术中存在的技术问题之一;为此,本发明提出了基于改进YOLOv7算法的铁路行人警戒系统及方法,用于解决现有安防系统在复杂城际铁路场景中存在检测精度不够、误报率高且动态风险量化能力不足的技术问题
[0062]In terms of detection accuracy and robustness, this invention achieves a balanced optimization of detection efficiency and accuracy through improvements to the YOLOv7 algorithm. Replacing the original backbone network with FasterNet significantly reduces the number of model parameters, and combining it with the CARAFE lightweight upsampling operator enhances computational efficiency, making it particularly suitable for deployment on edge devices. The introduced SimAM attention mechanism, through parameter-free feature enhancement, effectively improves pedestrian recognition accuracy in complex backgrounds at railway platforms (such as billboards, stacked luggage, and other distractions). The EIOU loss function, through optimization of center point distance and aspect ratio, reduces the false negative rate in densely populated scenarios, enhancing the model's robustness.
Smart Images

Figure CN120299063B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of railway safety and involves deep learning technology, specifically a railway pedestrian warning system and method based on the improved YOLOv7 algorithm. Background Technology
[0002] Intercity railways, as an important mode of transportation, connect urban transportation networks, providing people with a fast and efficient way to travel. However, pedestrian safety on intercity railways has always been a major concern. The presence of pedestrians on railway lines can lead to serious safety accidents, such as pedestrians accidentally entering the tracks, crossing the tracks, or waiting for transportation on the tracks. These situations can result in collisions, endangering pedestrian lives, and also bring certain safety hazards and operational pressures to railway operations and management.
[0003] However, given the complex environment of railway platforms, with numerous obstructions such as billboards and haphazardly stacked luggage, as well as the dense crowds, traditional detection technologies struggle to accurately identify pedestrians, leading to missed or false detections and failing to reliably guarantee pedestrian safety. Furthermore, existing systems lack effective integration of multi-dimensional data, including real-time train location and pedestrian movement trajectories, making it difficult to dynamically quantify pedestrian risk levels or intelligently adjust warning zones based on train operation status. This results in delayed and inaccurate warnings when sudden dangerous behavior occurs, hindering timely responses. Summary of the Invention
[0004] This invention aims to solve at least one of the technical problems existing in the prior art; to this end, this invention proposes a railway pedestrian warning system and method based on the improved YOLOv7 algorithm, which is used to solve the technical problems of insufficient detection accuracy, high false alarm rate and insufficient dynamic risk quantification capability of existing security systems in complex intercity railway scenarios.
[0005] To achieve the above objectives, a first aspect of the present invention provides a railway pedestrian warning system based on an improved YOLOv7 algorithm, comprising:
[0006] Zone division module: used to dynamically divide the warning zone according to the train's GPS positioning and train timetable, resulting in several zones; wherein the several zones include danger zones, buffer zones and safety zones;
[0007] Data processing module: used to identify pedestrians in the monitored area using a high-resolution camera with a pedestrian target detection model and obtain pedestrian coordinates; wherein, the pedestrian target detection model is built based on the improved YOLOv7 algorithm;
[0008] Risk scoring module: used to calculate the pedestrian dynamic potential energy field and potential energy gradient in real time based on pedestrian coordinates and several areas, and obtain the risk score based on the potential energy gradient and pedestrian motion vector;
[0009] Tiered early warning module: used to provide tiered early warnings and voice alerts to pedestrians based on risk scores.
[0010] Furthermore, the dynamic division of warning zones based on train GPS positioning and train timetables includes:
[0011] A1, Obtain the train's latitude and longitude coordinates (x, y, y) in real time based on the train's GPS positioning. train ,y train ), velocity v train And obtain the estimated arrival time t of the next train from the train timetable. ETA ;
[0012] A2, according to the formula R = R base +v train ·(t ETA -t current The dynamic expansion radius R is obtained by calculating k; where R base t represents the base radius. current This represents the current time, and k represents an adjustment coefficient used to control the expansion speed.
[0013] A3, Before the train enters the platform, the restricted area is divided as follows:
[0014] The area with a radius of R centered on the train is designated as the danger zone;
[0015] The ring-shaped area extending D distance from the danger zone is designated as a buffer zone.
[0016] The remaining areas are designated as safe zones;
[0017] A4, When a train enters the platform, the restricted area is divided as follows:
[0018] The area a distance 'a' on both sides of the railway track centerline is designated as a danger zone.
[0019] The area at a distance of (a, b) on both sides of the railway track centerline is divided into a buffer zone;
[0020] The remaining areas are designated as safe zones;
[0021] A5 maps the coordinates of danger zones, buffer zones, and safe zones to the field of view of a high-resolution camera, generating a virtual electronic fence.
[0022] It should be noted that if the divided areas overlap, the overlapping area shall be classified as the higher-level area, and the order of the area levels is: danger area > buffer area > safe area; for example, if the danger area and the buffer area overlap, the overlapping part shall be classified as the danger area.
[0023] During the dynamic zone delineation process, before the train enters the platform, the radius of the danger zone can be dynamically adjusted based on factors such as train speed, estimated arrival time, and current time. When the train enters the platform, different zones are defined according to the distance to both sides of the track centerline. By flexibly adjusting the warning range, various practical factors during train operation are fully considered, effectively improving the safety guarantee capability for railway pedestrians in different scenarios. Furthermore, by mapping the area coordinates to the camera's field of view to generate a virtual electronic fence, it facilitates subsequent personnel monitoring and early warning in conjunction with video surveillance.
[0024] Furthermore, the pedestrian target detection model is built based on the improved YOLOv7 algorithm, including:
[0025] B1, an improvement on the original YOLOv7 network:
[0026] The backbone feature extraction network of the original YOLOv7 network was replaced with the lightweight network FasterNet;
[0027] The upsampling operator of the original YOLOv7 network was replaced with the lightweight upsampling operator CARAFE to reduce model parameters and improve the model's feature extraction capability.
[0028] Introducing the SimAM attention mechanism enhances the model's ability to perceive people in complex backgrounds near intercity railways;
[0029] In the post-processing part, the robustness of the model in dense locations near intercity railways is further enhanced based on the EIOU loss function, resulting in the YOLOv7-FasterNet network model.
[0030] B2. Retrieve video footage from cameras on platforms near the intercity railway, use a frame-sampling algorithm every second to obtain pedestrian image data, and manually filter the images to obtain an image set.
[0031] B3, use annotation tools to annotate the image set to obtain the original dataset containing image and pedestrian labels;
[0032] B4. After augmenting the original dataset according to a preset ratio, it is divided into training set, validation set and test set according to a preset split ratio.
[0033] B5. Pre-train the YOLOv7-FasterNet network using the ImageNet dataset, save the pre-trained weights, and obtain the pre-trained model;
[0034] B6. The parameters of the pre-trained model are fine-tuned using the training set, validation set, and test set to obtain a pedestrian target detection model that meets the preset accuracy requirements.
[0035] In the YOLOv7-FasterNet network model, the FasterNet lightweight backbone network reduces redundant computation through partial convolution (PConv) technology, effectively adapting to the computing power limitations of edge devices such as NVIDIA Jetson Xavier NX. Secondly, the CARAFE dynamic upsampling operator improves the feature pyramid fusion efficiency and pedestrian detection accuracy by predicting a dedicated convolution kernel for each location. Simultaneously, the SimAM parameterless attention mechanism adaptively allocates weights through an energy function, suppressing interference information and reducing the model's false detection rate in complex backgrounds. Furthermore, by applying the EIoU loss function, combining overlap loss, center distance, and aspect ratio constraints, the model's localization accuracy in dense scenes is enhanced, providing an efficient solution for pedestrian safety protection on intercity railways.
[0036] Furthermore, the deployment process of the pedestrian target detection model includes:
[0037] After configuring the deep learning environment on NVIDIA Jetson Xavier NX, the pedestrian target detection model is deployed on NVIDIA Jetson Xavier NX;
[0038] TensorRT is used to accelerate inference of the deployed pedestrian target detection model, thereby improving the real-time processing performance of the pedestrian target detection model.
[0039] Furthermore, the step of calculating the pedestrian's dynamic potential energy field in real time based on the pedestrian's coordinates and several regions includes:
[0040] The pedestrian target detection model is used to obtain the pedestrian's coordinates (x, y) at time t in real time. t ,y t );
[0041] According to the formula The dynamic potential energy field U(x) of the pedestrian was calculated. t ,y t ); where t acc k represents the cumulative time a pedestrian spends in the monitored area. z (t) represents the dynamic potential energy intensity in region z at time t, where T is the dynamic potential energy intensity. Y λ1 represents the time series of pedestrians entering the buffer zone in history, λ2 represents the historical memory weight, and λ3 represents the time decay coefficient.
[0042] Furthermore, the dynamic potential energy intensity kz The methods for obtaining (t) include:
[0043] The danger zone is marked as set Ω. R The buffer region is marked as set Ω. Y Mark the safe area as set Ω G ;
[0044] Where, k R k Y k G Let k represent the basic potential energy in the danger zone, buffer zone, and safe zone, respectively, and k R >k Y >k G α(t) represents the time sensitivity coefficient, according to the formula The calculation yields t. train N(t) represents the arrival time of the next train, and N(t) represents the number of pedestrians entering the buffer zone Ω. Y The cumulative number of times, β represents the cumulative effect coefficient of the buffer zone.
[0045] Potential energy fields are used to describe the potential energy of an object at different locations in space. In this invention, the position and behavior of pedestrians are mapped onto an energy field, with higher energy indicating greater risk. Therefore, a dynamic potential energy field constructs a multi-dimensional risk perception model through spatial partitioning quantization, enhanced temporal sensitivity, and the cumulative effect of historical behavior. Its mathematical expression combines the concept of physical potential energy with pedestrian behavior, enabling the system to accurately identify high-risk scenarios (such as pedestrians rapidly approaching a dangerous area while a train is approaching), thereby triggering tiered warnings and improving the intelligent level of safety protection on intercity railways.
[0046] Furthermore, the formula for calculating the potential energy gradient is:
[0047] Furthermore, the risk score obtained based on the potential energy gradient and pedestrian motion vector includes:
[0048] According to the formula The pedestrian motion vector is calculated; where Δt represents the preset time interval.
[0049] According to the risk scoring formula: The risk score S is calculated; where η1 represents the position term coefficient, η2 represents the acceleration term coefficient, and |||| represents the modulus.
[0050] In the risk scoring formula, the pedestrian movement vector predicts behavioral trends by quantifying the magnitude and direction of speed (risk increases as the pedestrian approaches a high potential energy region); the potential energy gradient reflects the spatial abrupt changes in the danger zone (a steep gradient represents the risk boundary effect); and the acceleration term captures the drastic degree of speed change (such as sudden acceleration suggesting an intrusion intention). Through a real-time updated dynamic potential energy field model, continuously integrating train position, time decay factor, and pedestrian behavior data, the scoring mechanism can accurately respond to instantaneous state changes—automatically strengthening the potential energy field intensity when a train approaches and dynamically correcting the gradient distribution when the pedestrian's trajectory deviates, achieving precise spatiotemporal matching of risk warnings.
[0051] Furthermore, the step of providing graded early warnings and voice alerts to pedestrians based on risk scores includes:
[0052] Set a first-level threshold A and a second-level threshold B, where A > B;
[0053] When the risk score S∈(0,B], a green safety signal is sent to the staff.
[0054] When the risk score S∈(B,A], a yellow warning signal is sent to the staff.
[0055] When the risk score S∈[A,+∞), a red danger signal is sent to the staff and a broadcast voice alarm is issued.
[0056] A second aspect of the present invention provides a railway pedestrian warning method based on an improved YOLOv7 algorithm, comprising:
[0057] S1. Based on the train's GPS positioning and train timetable, the warning zone is dynamically divided into several zones; wherein the several zones include danger zones, buffer zones and safety zones.
[0058] S2, using a high-resolution camera with a pedestrian target detection model to identify pedestrians in the monitored area and obtain pedestrian coordinates; wherein, the pedestrian target detection model is built based on the improved YOLOv7 algorithm;
[0059] S3 calculates the dynamic potential energy field and potential energy gradient of pedestrians in real time based on pedestrian coordinates and several areas, and obtains a risk score based on the potential energy gradient and pedestrian motion vector.
[0060] S4 provides graded warnings and voice alerts to pedestrians based on risk scores.
[0061] Compared with the prior art, the beneficial effects of the present invention are:
[0062] In terms of detection accuracy and robustness, this invention achieves a balanced optimization of detection efficiency and accuracy through improvements to the YOLOv7 algorithm. Replacing the original backbone network with FasterNet significantly reduces the number of model parameters, and combining it with the CARAFE lightweight upsampling operator enhances computational efficiency, making it particularly suitable for deployment on edge devices. The introduced SimAM attention mechanism, through parameter-free feature enhancement, effectively improves pedestrian recognition accuracy in complex backgrounds at railway platforms (such as billboards, stacked luggage, and other distractions). The EIOU loss function, through optimization of center point distance and aspect ratio, reduces the false negative rate in densely populated scenarios, enhancing the model's robustness.
[0063] In terms of dynamic risk quantification, this invention constructs a risk scoring model by integrating multi-dimensional data such as real-time train location and pedestrian movement trajectory to achieve dynamic quantification of risk level; and intelligently adjusts the warning area range based on train operation status to avoid resource waste while ensuring safe coverage, thereby improving the timeliness of response to sudden dangerous behaviors.
[0064] At the system deployment level, this invention is based on the hardware adaptation solution of NVIDIA Jetson Xavier NX and the TensorRT acceleration engine to improve the model inference speed, which can meet the needs of real-time video stream processing. In addition, the hierarchical early warning mechanism implements differentiated alarm strategies based on risk scores, which reduces the false alarm rate while ensuring that high-risk events can trigger strong warnings, forming a complete early warning system from safety reminders to emergency alarms. Attached Figure Description
[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0066] Figure 1 A schematic diagram of the technical process of the railway pedestrian warning system based on the improved YOLOv7 algorithm provided by the present invention;
[0067] Figure 2 A schematic diagram of the railway pedestrian warning system framework based on the improved YOLOv7 algorithm provided by the present invention;
[0068] Figure 3 A schematic diagram illustrating the workflow of the region division module provided by this invention;
[0069] Figure 4 The framework diagram of the pedestrian target detection model based on the improved YOLOv7 algorithm provided by this invention is shown. Detailed Implementation
[0070] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0071] Please see Figures 1-4 The first aspect of this invention provides a railway pedestrian warning system based on an improved YOLOv7 algorithm, comprising:
[0072] Zone division module: used to dynamically divide the warning zone based on the train's GPS positioning and train timetable, resulting in several zones; these zones include danger zones, buffer zones, and safety zones.
[0073] Data processing module: used to identify pedestrians in the monitored area using a high-resolution camera with a pedestrian target detection model and obtain pedestrian coordinates; the pedestrian target detection model is built based on the improved YOLOv7 algorithm;
[0074] Risk scoring module: used to calculate the pedestrian's dynamic potential energy field and potential energy gradient in real time based on the pedestrian's coordinates and several areas, and to obtain a risk score based on the potential energy gradient and the pedestrian's motion vector;
[0075] Tiered early warning module: used to provide tiered early warnings and voice alerts to pedestrians based on risk scores.
[0076] It should be noted that the region division module, data processing module, risk scoring module, and graded early warning module of the present invention are communicatively connected.
[0077] In this embodiment, the region division module achieves real-time dynamic region division by integrating train dynamic information and track geographic data. The specific steps are as follows:
[0078] First, real-time latitude and longitude coordinates (x, y, t) are obtained through train GPS positioning. train ,y train ), running speed v train And extract the estimated arrival time t of the next train from the timetable. ETA ;
[0079] Next, set the base radius R. base Adjustment coefficient k (used to control the expansion speed, set based on historical experience), extension distance D (used to control the width of the buffer zone), and track safety distances a and b;
[0080] Then use the common R = R base +vtrain ·(t ETA -t current )·k calculates the dynamic expansion radius R of the hazardous area, where t current The current time;
[0081] Next, the areas are divided according to the train's operating location:
[0082] Before the train enters the platform: (GPS positioning can be used to determine whether the train has entered the platform)
[0083] Danger zone: A circular area with radius R centered on the train;
[0084] Buffer zone: A ring-shaped area extending D meters beyond the danger zone;
[0085] Safe Zone: The rest of the area;
[0086] When the train enters the platform:
[0087] Danger zone: a meters on each side of the track centerline;
[0088] Buffer zone: a to b meters on either side of the track centerline;
[0089] Safety zone: The remaining platform area.
[0090] Then, the geographical coordinates of the danger zone, buffer zone, and safe zone are converted into pixel coordinates of the camera's field of view to generate a virtual electronic fence, which is then overlaid onto the monitoring screen.
[0091] When the defined areas overlap, they are covered according to priority: danger zone > buffer zone > safe zone. For example, if the danger zone of one train overlaps with the safe zone of another train, the overlapping area is the danger zone.
[0092] The zone division module enhances the accuracy and adaptability of railway pedestrian safety protection through a spatiotemporally coupled intelligent zone division strategy. Based on the train's real-time speed and remaining time, the module dynamically expands the radius of the danger zone. Before entering the station, a dynamic circular area is used to cover potential risks in the direction of train travel. After entering the station, the division switches to a fixed distance on both sides of the track centerline, adapting to dense pedestrian traffic on the platform. Simultaneously, combined with the train timetable, the module dynamically adjusts the potential energy intensity of the danger zone through a time-sensitive coefficient, causing the risk level to gradually increase as the train approaches. Furthermore, the module integrates GPS positioning, track geographic data, and camera field of view to achieve precise mapping from physical space to the virtual monitoring area, ensuring a high degree of consistency between the electronic fence and the actual scene, thus guaranteeing the safety of railway pedestrians.
[0093] In the region segmentation module, assuming a certain intercity railway station, a train approaches the platform at 80 km / h, with an estimated arrival time t. ETA =14:05:00, current time tcurrent =14:00:00, then t ETA -t current =5min=0.0833h, set R base =50m, k=0.03;
[0094] Then calculate the dynamic radius: R = 50 + 80 × 0.0833 × 0.03 × 1000 = 249.92 m ≈ 250 m; (where multiplying by 1000 indicates converting the unit km to m)
[0095] Before entering the station: The danger zone is a circular area with a radius of 2500 meters centered on the train; the buffer zone is a ring-shaped area extending 50 meters outward (D=50m);
[0096] After entering the station: the danger zone is 3 meters on each side of the track centerline; the buffer zone is an area of 3 to 8 meters.
[0097] Next, the coordinates of the above area are mapped onto the platform camera screen, and red (danger) and yellow (buffer) warning boxes are overlaid.
[0098] To improve pedestrian safety on intercity railways, an effective pedestrian monitoring and warning system is needed. Traditional pedestrian detection algorithms often suffer from low accuracy, high false alarm rates, and insufficient detection capability for small targets in complex intercity railway scenarios. Therefore, in the data processing module, a pedestrian target detection model based on an improved YOLOv7 algorithm is constructed and deployed on high-resolution cameras to identify pedestrians in real time within the monitored area, obtaining their coordinates and providing a data foundation for subsequent risk scoring.
[0099] In this embodiment, the specific improvement and deployment process of the YOLOv7 algorithm may include the following steps:
[0100] S1: Obtain video data of pedestrians near the intercity railway platform by accessing the camera, write video frame-segmentation code to perform video frame-segmentation operation to obtain image data and filter it, removing images with high repetition rate and keeping images with low repetition rate.
[0101] The filtered images were augmented using data augmentation code, expanding the original images by a 1:7 ratio. Augmentation methods included flipping, rotating, cropping, scaling, translation, Gaussian noise, and mosaicking. The dataset was then divided into training, validation, and test sets in a 6:2:2 ratio.
[0102] S2: Lightweight improvements to the YOLOv7 algorithm yield the YOLOv7-FasterNet network model:
[0103] The backbone feature extraction network of the original YOLOv7 algorithm is replaced with the lightweight network FasterNet;
[0104] Replace the upsampling operator with the lightweight upsampling operator CARAFE to reduce model parameters and improve the model's feature extraction capability;
[0105] Introducing the SimAM attention mechanism enhances the model's ability to perceive people in complex backgrounds near intercity railways;
[0106] Furthermore, the robustness of the model in dense locations near intercity railways is further enhanced in the post-processing part based on Soft-EIOU-NMS.
[0107] Specifically, to minimize model parameters and address issues such as poor real-time performance at edges due to excessive parameter count, this embodiment replaces the backbone feature extraction network in the YOLOv7 object detection network with the lightweight FasterNet network. Compared to the traditional YOLOv7 backbone convolutional network, FasterNet consists of three main parts: a base network module, a feature fusion module, and an upsampling module, which can be divided into four levels. Each stage is equipped with a set of FasterNet blocks, and an embedding or merging layer is placed before each stage. The last three layers are used for feature classification. Inside each FasterNet block, the structure is a partial convolutional (PConv) layer followed by two pointwise convolutional (PWConv) layers, and normalization (BN) and activation (ReLU) layers are placed only after the intermediate layers. This design can preserve feature diversity and achieve lower latency.
[0108] For example, if given a set of dimensions H in ×W in ×C in The feature map is used as input data, where H in W in C in The height, width, and number of channels of the input feature map are respectively represented by H. out ×W out ×C out The relationship between the input and output layers of a standard convolutional layer with n convolutional kernels of k×k×c is as follows:
[0109]
[0110] C in =c
[0111] C out =n
[0112] In the formula: p represents the patch; c represents the number of channels in the convolution kernel; n represents the number of convolution kernels; w×h represents the convolution dimension; s represents the stride;
[0113] The formulas for calculating the computational cost (FLOPs) and parameter count (params) of this operation are as follows:
[0114] FLOPs=C in ×k 2 ×C out ×W in ×H in
[0115] params = C out ×(k 2 ×C in +1)
[0116] In the formula: FLOPs represents the number of floating-point operations, used to measure the computational complexity of the model; C in Indicates the number of input channels; k represents the size of the standard convolutional kernel; C out This represents the number of output channels, which is also the number of convolution kernels, n.
[0117] The FasterNet module discovered that the main reason for the low FLOPs problem caused by the existing operator DWConv is frequent memory access. Therefore, it proposed Partial Convolution (PConv) to address the redundancy of features in convolutional neural networks. The purpose of Partial Convolution is to reduce both memory access and computational redundancy. Its FLOPs1 calculation method is shown below:
[0118]
[0119] If we compare the two, we can see that some convolutions in FasterNet can significantly reduce computation, thus achieving the goal of a lightweight network.
[0120]
[0121] Since this invention needs to be deployed on edge devices, in this embodiment, the ordinary upsampling operator in the original YOLOv7 algorithm is replaced with the lightweight upsampling operator CARAFE.
[0122] The CARAFE lightweight upsampling operator is mainly divided into two modules: the upsampling kernel prediction module and the feature recombination module.
[0123] The upsampling kernel prediction module consists of three steps: feature map channel compression, content encoding and upsampling kernel prediction, and upsampling kernel normalization.
[0124] In feature map channel compression, for an input feature map of shape H×W×C, a 1×1 convolution is first used to compress its number of channels to C. m The main purpose of this step is to reduce the amount of computation in subsequent steps;
[0125] In content encoding and upsampling kernel prediction, for the compressed input feature map in the first step, a k-kernel is used. encoder ×k encoder Convolutional layers are used to predict the upsampling kernel, with C input channels. m The number of output channels is Then, by unfolding in space, we obtain the shape as The upsampling kernel;
[0126] In the upsampling kernel normalization, the upsampling kernel obtained in the second step is normalized using softmax so that the sum of the convolution kernel weights is 1.
[0127] Feature Reconstruction Module: For each position in the output feature map, map it back to the input feature map and extract the k-th position centered thereon. up ×k up The region is calculated by taking the dot product of the predicted upsampling kernel at that point, and the predicted upsampling kernel at that point, to obtain the output value. Different channels at the same location share the same upsampling kernel.
[0128] The area around intercity railways is complex and densely populated. To improve the model's ability to perceive people in time and space, the SimAM attention mechanism was used.
[0129] Given a query sequence Q and a key-value pair sequence K, the attention mechanism determines the attention weights by calculating the similarity between the query sequence and the key sequence. The specific process consists of the following three steps: For each element q in the query sequence Q, calculate its cosine similarity with each element k in the key sequence K; for each query sequence element q, normalize its similarity with the key sequence element k to obtain the attention weight; use the attention weights to perform a weighted summation on the value sequence V to obtain the final attention representation.
[0130] Next, regarding the loss function, general loss functions do not take into account the orientation between the ground truth bounding box and the predicted bounding box, resulting in slower convergence. This embodiment introduces EIoU into the improved YOLOv7 model to improve training speed, including the elimination of overlap loss (L... IoU ), center distance loss (L dis ) and width and height loss (L asp Introducing a suitable loss function, its calculation formula is as follows:
[0131]
[0132] In the formula: IoU represents IoU loss, ρ 2 (b,b gt ρ represents the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box. 2 (w,w gt ), ρ 2 (h,h gt ) represent the Euclidean distance between the width and height of the predicted bounding box and the width and height of the ground truth bounding box, respectively; where c represents the diagonal length of the smallest bounding rectangle that can contain both the predicted bounding box and the ground truth bounding box. and It represents the width and height of the smallest outer bounding box that covers both boxes.
[0133] By following the steps above, we can obtain the most effective pedestrian object detection network, YOLOv7-FasterNet.
[0134] S3: In terms of model training, we first pre-trained the model using the ImageNet dataset, then used the collected pedestrian dataset near the intercity railway to train the improved model and fine-tuned its parameters. After training, we obtained the final weights and various training data evaluation metrics. At the same time, we tested the improved model on the test set to obtain a pedestrian target detection model that meets the preset accuracy requirements.
[0135] S4: By analyzing the comparison between the model parameters and the computing power of the edge device, in this embodiment, NVIDIA Jetson Xavier NX is selected as the intelligent hardware processing device at the edge of the data processing module.
[0136] The pedestrian object detection model was deployed on the edge embedded device NVIDIA Jetson Xavier NX, and the hardware device was configured with a deep learning environment. At the same time, TensorRT was used to accelerate the inference of the deployed object detection model, further improving the real-time processing performance of the lightweight YOLOv7 model on the edge device.
[0137] The data processing module obtains the pedestrian coordinates (x) t ,y t After that, the data is transmitted to the risk scoring module for real-time pedestrian motion feature recognition and dynamic potential energy field calculation. Then, the risk score is calculated based on the potential energy gradient and pedestrian motion vector to quantify the risk. The specific steps are as follows:
[0138] First, mark the danger zone as set Ω. R The buffer region is marked as set Ω. Y Mark the safe area as set Ω G Based on the pedestrian's current location, dynamically calculate the potential energy intensity k. z (t):
[0139] When pedestrians are in a danger zone, k z (t)=k R ·α(t), where This means that the closer the time is to the train's arrival, the stronger the potential energy;
[0140] When pedestrians are in the buffer zone, k z (t)=k Y ·(1+β·N(t)), where β represents the cumulative effect coefficient of the buffer zone, determined based on historical experience, and N(t) represents the number of pedestrians entering the buffer zone Ω. Y The more times a pedestrian enters the buffer zone, the stronger the potential energy.
[0141] When pedestrians are in a safe area, k z (t)=k G ;
[0142] Where, k R k Y k G Let k represent the basic potential energy in the danger zone, buffer zone, and safe zone, respectively, and k R >k Y >k G ;
[0143] Next, the dynamic potential energy field of the pedestrian is calculated according to the formula: Among them, t acc T represents the cumulative time a pedestrian spends within the monitored area. Y λ1 represents the time series of pedestrians entering the buffer zone in history, λ2 represents the spatial weight, which is set according to the pedestrian in region z, λ3 represents the historical memory weight, and λ3 represents the time decay coefficient, which is used to weight and decay past behaviors entering the buffer zone. All of these are set based on historical experience.
[0144] Calculate the gradient vector of the potential field:
[0145] Calculate the instantaneous pedestrian motion vector based on the coordinate differences of several consecutive frames: Where Δt represents the preset time interval;
[0146] According to the risk scoring formula: The risk score S is calculated; where η1 represents the position term coefficient, η2 represents the acceleration term coefficient, and |||| represents the modulus.
[0147] For example, assume the fundamental potential energy intensity of each region is: k R =100, k Y =50,k G=10;
[0148] In the formula for calculating the potential energy intensity of the buffer zone, β = 0.5;
[0149] The coefficients in the dynamic potential energy field calculation formula are: λ1 = 0.6, λ2 = 0.4, λ3 = 0.1;
[0150] The coefficients in the risk scoring formula are: η1 = 0.3, η2 = 0.2;
[0151] The train is scheduled to arrive at the station at 14:00. The current time t = 13:57, then t train -t = 180s, at which point the cumulative time t after detecting a pedestrian entering the monitored area is reached. acc =150s, with a cumulative total of 3 entries into the buffer zone, and the time series is T. Y ={100,120,140};
[0152] The weighted sum of the historical terms is then calculated:
[0153] The pedestrian is currently in the buffer zone, so the potential energy intensity at this time is: k z (t)=k Y ·(1+β·N(t))=50×(1+0.5×3)=125;
[0154] Dynamic potential energy field
[0155] Assume pedestrian coordinates (x t ,y t )=(5,3), the potential energy gradient is (Example value; actual values require differentiation calculations.)
[0156] Δt = 0.5s, (x t+Δt ,y t+Δt )=(5.5,3.2);(x t+2Δt ,y t+2Δt = (6.1, 3.5);
[0157] acceleration
[0158] Risk score
[0159] Finally, in the tiered early warning module: pedestrians are given tiered early warnings and voice alerts based on risk scores, including:
[0160] Set a first-level threshold A and a second-level threshold B, where A > B;
[0161] When the risk score S∈(0,B], a green safety signal is sent to the staff.
[0162] When the risk score S∈(B,A], a yellow warning signal is sent to the staff.
[0163] When the risk score S∈[A,+∞), a red danger signal is sent to the staff and a broadcast voice alarm is issued.
[0164] Assuming thresholds A=5 and B=2 are set, the scoring result in the above example ∈(2,5]. The system will send a yellow warning signal to the staff's visualization device. The staff will quickly locate the pedestrian and provide prompts and warnings.
[0165] A second aspect of the present invention provides a railway pedestrian detection method based on an improved YOLOv7 algorithm, comprising:
[0166] S1, based on the train's GPS positioning and train timetable, the warning zone is dynamically divided into several zones; these zones include danger zones, buffer zones, and safety zones.
[0167] S2, using a high-resolution camera with a pedestrian target detection model to identify pedestrians in the monitored area and obtain pedestrian coordinates; the pedestrian target detection model is built based on the improved YOLOv7 algorithm;
[0168] S3 calculates the dynamic potential energy field and potential energy gradient of pedestrians in real time based on pedestrian coordinates and several areas, and obtains a risk score based on the potential energy gradient and pedestrian motion vector.
[0169] S4 provides graded warnings and voice alerts to pedestrians based on risk scores.
[0170] Some of the data in the above formula are calculated by removing dimensions and taking their numerical values. The formula is the closest to the real situation obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.
[0171] Working principle of the invention:
[0172] First, dynamic zone division is performed based on information such as the train's real-time location, speed, and arrival time. Before the train enters the station, the radius of the danger zone is dynamically adjusted; upon entering the station, fixed zones are divided along the track centerline, thus forming a virtual electronic fence encompassing three levels of zones: danger, buffer, and safety.
[0173] Next, pedestrian locations are identified in real time using a high-resolution camera based on an improved YOLOv7-FasterNet model. By leveraging a lightweight network and attention mechanism, the detection accuracy and efficiency in complex scenes are improved.
[0174] Then, a dynamic potential energy field is constructed by combining the pedestrian's current location, the train's approach time, and historical boundary crossing behaviors to quantify the risk. The risk score is calculated by comprehensively considering the pedestrian's direction of movement (risk increases when approaching the danger zone), potential energy gradient (the steepness of the danger boundary), and acceleration (the degree of speed change).
[0175] Finally, based on the risk score, three levels of warnings are triggered: green (safe), yellow (warning), and red (danger with voice alarm), to achieve differentiated safety responses and form a closed-loop protection system from detection to warning, thereby improving the intelligence and timeliness of railway pedestrian safety protection.
[0176] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A railway pedestrian warning system based on an improved YOLOv7 algorithm, characterized in that, include: The area division module is used to dynamically divide the warning area based on the train's GPS location and train timetable, resulting in several areas. The aforementioned areas include hazardous areas, buffer areas, and safe areas; Data processing module: used to identify pedestrians in the monitored area using a high-resolution camera with a deployed pedestrian detection model, and obtain pedestrian coordinates; wherein, the pedestrian detection model is built based on an improved YOLOv7 algorithm, and specifically includes: The backbone feature extraction network of the original YOLOv7 network was replaced with the lightweight network FasterNet; The upsampling operator of the original YOLOv7 network was replaced with the lightweight upsampling operator CARAFE; Introducing the SimAM attention mechanism enhances the model's ability to perceive people in complex backgrounds near intercity railways; In the post-processing part, the robustness of the model in dense locations near intercity railways is further enhanced based on the EIOU loss function, resulting in the YOLOv7-FasterNet network model. Risk scoring module: used to calculate the pedestrian dynamic potential energy field and potential energy gradient in real time based on pedestrian coordinates and several areas, and obtain the risk score based on the potential energy gradient and pedestrian motion vector; The process of calculating the pedestrian's dynamic potential energy field in real time based on the pedestrian's coordinates and several regions includes: The pedestrian target detection model is used to obtain the pedestrian's coordinates at time t in real time. ; According to the formula The dynamic potential energy field of the pedestrian was calculated. ;in, This indicates the cumulative time a pedestrian spends within the monitored area. This represents the dynamic potential energy intensity in region z at time t. This represents the time series of pedestrians entering the buffer zone throughout history. This indicates the specific points in time during which pedestrians entered the buffer zone. This represents the spatial weight, which is set based on the pedestrian's position in region z. Indicates the weight of historical memory. Indicates the time decay coefficient; This represents the dynamic potential energy intensity of a pedestrian when they enter the buffer zone at historical time τ. The dynamic potential energy intensity The methods of obtaining it include: Mark the danger zone as a set. Mark the buffer region as a set Mark the safe zone as a set ; ;in, , , These represent the basic potential energy in the danger zone, buffer zone, and safe zone, respectively. , Represents the time sensitivity coefficient. Indicates the time the train arrives at the platform. Indicates that pedestrians are entering the buffer zone The cumulative number of times, This represents the cumulative effect coefficient of the buffer zone; The risk score obtained based on the potential energy gradient and pedestrian motion vector includes: According to the formula The pedestrian motion vector is calculated; where, Indicates a preset time interval; According to the risk scoring formula: The risk score S is calculated; where, Indicates the coefficient of the position term. Represents the coefficient of the acceleration term. Indicates the modulus length. The potential gradient is expressed by the formula: ; Tiered early warning module: used to provide tiered early warnings and voice alerts to pedestrians based on risk scores.
2. The railway pedestrian warning system based on the improved YOLOv7 algorithm according to claim 1, characterized in that, The dynamic division of warning zones based on train GPS positioning and train timetables includes: A1. Obtain the train's latitude and longitude coordinates in real time based on the train's GPS positioning. ,speed And obtain the estimated arrival time of the next train from the train timetable. ; A2, according to the formula The dynamic expansion radius R is calculated; where, Indicates the base radius. This represents the current time, and k represents an adjustment coefficient used to control the expansion speed. A3, Before the train enters the platform, the restricted area is divided as follows: The area with a radius of R centered on the train is designated as the danger zone; The ring-shaped area extending D distance from the danger zone is designated as a buffer zone. The remaining areas are designated as safe zones; A4, When a train enters the platform, the restricted area is divided as follows: The area a distance 'a' on both sides of the railway track centerline is designated as a danger zone. The area at a distance of (a, b) on both sides of the railway track centerline is divided into a buffer zone; The remaining areas are designated as safe zones; A5 maps the coordinates of danger zones, buffer zones, and safe zones to the field of view of a high-resolution camera, generating a virtual electronic fence.
3. The railway pedestrian warning system based on the improved YOLOv7 algorithm according to claim 1, characterized in that, The pedestrian target detection model is built based on the improved YOLOv7 algorithm and includes: B1, an improvement on the original YOLOv7 network: The backbone feature extraction network of the original YOLOv7 network was replaced with the FasterNet network, and the upsampling operator of the original YOLOv7 network was replaced with the CARAFE upsampling operator. By introducing the SimAM attention mechanism and the EIOU loss function, the YOLOv7-FasterNet network model is obtained; B2. Retrieve video footage from cameras near the intercity railway station platform, use a frame-by-second extraction method to obtain pedestrian image data, and manually filter the images to obtain an image set; wherein, the frame-by-second extraction method means that a video frame is extracted every second. B3, use annotation tools to annotate the image set to obtain the original dataset containing image and pedestrian labels; B4. After augmenting the original dataset according to a preset ratio, it is divided into training set, validation set and test set according to a preset split ratio. B5. Pre-train the YOLOv7-FasterNet network using the ImageNet dataset, save the pre-trained weights, and obtain the pre-trained model; B6. The parameters of the pre-trained model are fine-tuned using the training set, validation set, and test set to obtain a pedestrian target detection model that meets the preset accuracy requirements.
4. The railway pedestrian warning system based on the improved YOLOv7 algorithm according to claim 3, characterized in that, The deployment process of the pedestrian target detection model includes: After configuring the deep learning environment on the NVIDIA Jetson Xavier NX, the pedestrian target detection model was deployed on the NVIDIA Jetson Xavier NX. TensorRT is used to accelerate inference of the deployed pedestrian target detection model, thereby improving the real-time processing performance of the pedestrian target detection model.
5. The railway pedestrian warning system based on the improved YOLOv7 algorithm according to claim 1, characterized in that, The time sensitivity coefficient is determined according to the formula... Calculated.
6. The railway pedestrian warning system based on the improved YOLOv7 algorithm according to claim 1, characterized in that, The method of providing graded early warnings and voice alerts to pedestrians based on risk scores includes: Set a first-level threshold A and a second-level threshold B, where A > B; When the risk score S∈(0,B], a green safety signal is sent to the staff. When the risk score S∈(B,A], a yellow warning signal is sent to the staff. When the risk score S∈[A,+∞), a red danger signal is sent to the staff and a broadcast voice alarm is issued.
7. A railway pedestrian warning method based on an improved YOLOv7 algorithm, applied to the railway pedestrian warning system based on an improved YOLOv7 algorithm as described in any one of claims 1-6, characterized in that, include: S1. Based on the train's GPS positioning and train timetable, the warning zone is dynamically divided into several zones; wherein the several zones include danger zones, buffer zones and safety zones. S2, using a high-resolution camera with a pedestrian target detection model to identify pedestrians in the monitored area and obtain pedestrian coordinates; wherein, the pedestrian target detection model is built based on the improved YOLOv7 algorithm; S3 calculates the dynamic potential energy field and potential energy gradient of pedestrians in real time based on pedestrian coordinates and several areas, and obtains a risk score based on the potential energy gradient and pedestrian motion vector. S4 provides graded warnings and voice alerts to pedestrians based on risk scores.
Citation Information
Patent Citations
Data aggregation method with attribute correlation
CN102118819A
Construction alert area monitoring and early warning system and method based on image recognition
CN112216049A