Rail transit video intelligent analysis method, medium and system

By constructing an adaptive resolution image hierarchical structure and hierarchical processing strategy, dynamically optimize computing resource allocation, combined with improved frame difference method and optical flow estimation technology, and applying the track scene attention network model, the problem of large amount of calculation and low efficiency in high-definition video processing of rail transit video surveillance system is solved, real-time and accurate abnormal behavior recognition and early warning.

CN120259946AActive Publication Date: 2025-07-04QINGDAO HENGXUN IND & TRADE CO LTD

Patent Information

Application Number
CN202510422245.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-04
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing rail transit video surveillance system has a large amount of calculation and low efficiency when processing high-definition videos, making it difficult to accurately identify abnormal behaviors in real time under limited computing resources, especially in complex and changeable platform environments.

Method used

The adaptive resolution image hierarchical structure is constructed, the minimum spanning tree algorithm is used to dynamically determine the optimal resolution level, combined with improved frame difference method and optical flow estimation technology to detect moving targets, the Hungarian algorithm is used to track targets, the orbital scene attention network model is introduced for abnormal target recognition, and the risk level is quantified through the abnormal behavior risk assessment function, and the layered processing strategy is used to optimize computing resource allocation.

Benefits of technology

It significantly reduces the computational complexity, improves processing efficiency and identification accuracy, can realize real-time abnormal behavior detection and early warning under limited resources, enhances the system's recognition ability in complex environments, and provides a scientific basis for security management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259946A_ABST
    Figure CN120259946A_ABST
Patent Text Reader

Abstract

The invention provides a rail transit video intelligent analysis method, medium and system, and belongs to the technical field of rail transit, and the method comprises the steps: constructing a self-adaptive resolution image hierarchical structure, and dynamically determining an optimal resolution hierarchy based on a minimum spanning tree algorithm; constructing an abnormal target movement matrix by combining an improved frame difference method with an optical flow estimation technology; applying a Hungary algorithm to realize target tracking; constructing an abnormal target change index to evaluate a behavior abnormal degree; dividing monitoring stages according to the running state of the vehicle to realize scene self-adaption; an orbit scene attention network model is introduced to identify abnormal behaviors; quantifying a risk level through an abnormal behavior risk assessment function; implementing a hierarchical processing strategy to enable a low-resolution hierarchy to be responsible for global rapid screening and a high-resolution hierarchy to be responsible for fine analysis of key areas; the risk trend is estimated in combination with the rail transit abnormal event prediction model, and the technical problems of large calculation amount and low efficiency of high-definition video processing are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of rail transit technology, and in particular, relates to a rail transit video intelligent analysis method, medium and system. Background Art

[0002] Rail transit platform security monitoring is a vital component of urban public transportation systems. Traditional rail transit video surveillance technology mainly relies on fixed-position cameras to collect high-definition video streams and conduct real-time analysis and processing through computer vision algorithms. Existing technologies usually use direct processing of full-resolution video frames, and use frame difference methods, background subtraction or basic target tracking algorithms to detect abnormal situations in the picture. These methods can meet basic security monitoring needs under standard video resolution and ideal environmental conditions.

[0003] However, with the development of high-definition camera technology, the video resolution collected by rail transit monitoring systems continues to increase, and the computing resources required to directly process high-resolution video data are growing exponentially. The traditional full-frame processing method significantly increases the computing burden when facing 4K or even 8K video, resulting in system response delays and low processing efficiency. Especially in edge computing environments, limited computing resources make it difficult to support real-time analysis of high-resolution videos, which increases the delay in abnormal behavior detection and reduces the effectiveness of the early warning system.

[0004] What is more serious is that the rail transit platform environment is complex and changeable, with dense traffic and unstable lighting conditions. Directly applying computationally intensive algorithms to process high-definition video is not only inefficient, but also difficult to balance the requirements of real-time and accuracy. How to efficiently process high-resolution video and achieve real-time and accurate identification of abnormal behavior under limited computing resources has become a technical problem that needs to be solved urgently in the rail transit safety monitoring system. Summary of the invention

[0005] In view of this, the present invention provides a rail transit video intelligent analysis method, medium and system, which can solve the technical problems in the prior art that the rail transit video monitoring system has large computational complexity and low efficiency when processing high-definition video, and it is difficult to achieve real-time and accurate identification of abnormal behavior under limited computing resources.

[0006] The present invention is implemented as follows: In the first aspect of the present invention, a method for intelligent analysis of rail transit videos is provided, including: constructing an adaptive resolution image hierarchical structure, and dynamically determining the optimal number of resolution levels based on the minimum spanning tree algorithm; calculating the difference between adjacent video frames using an improved frame difference method, and constructing an abnormal target movement matrix; based on the abnormal target movement matrix, calculating an abnormal target movement displacement matrix, and applying the Hungarian algorithm to solve the multi-target matching problem; constructing an abnormal target change index according to the abnormal target movement displacement matrix; dividing the platform monitoring area into four time periods according to the operating state of the rail transit vehicle; introducing a rail scene attention network model for abnormal target recognition; using an abnormal behavior risk assessment function to quantify the risk level of abnormal behavior; performing parallel calculations on images of different resolution levels based on a hierarchical processing strategy; and dividing the warning levels according to the detection results and the abnormal target change index according to a preset risk level.

[0007] Among them, the adaptive resolution image hierarchical structure generates a resolution pyramid from the original high-definition video frame through a recursive binary downsampling method, and the resolution of each level is determined by a resolution level evaluation function.

[0008] Among them, the resolution level evaluation function is used to dynamically determine the optimal number of resolution division levels and the resolution parameters of each level. The inputs include the original resolution of the video, the calculation resource limit parameter, the target detection accuracy requirement parameter, the real-time requirement parameter, and the scene complexity score, and the outputs are the optimal number of resolution levels and the specific downsampling ratio of each level.

[0009] Among them, the improved frame difference method combines the optical flow estimation technology to improve the motion detection accuracy, and screens out the regions with significant changes through an adaptive threshold.

[0010] Among them, the abnormal target movement matrix refers to a two-dimensional array generated by calculating the pixel difference between adjacent video frames, which represents the position information of the moving object in the picture, and the value of each element represents the change degree of the pixel point at the corresponding position; the abnormal target movement displacement matrix refers to the amount of position change of the detected abnormal target in consecutive multi-frame videos, including two components of horizontal displacement and vertical displacement, and is used to analyze the target movement trajectory.

[0011] Among them, the abnormal target change index refers to a weighted calculated value that comprehensively considers the target size change, speed change, and direction change. The higher the value, the more abnormal the target behavior. The calculation formula is the target size change rate multiplied by the weight coefficient plus the moving speed change rate multiplied by the weight coefficient plus the direction change rate multiplied by the weight coefficient.

[0012] Among them, the rail scenario attention network model integrates a spatio-temporal attention mechanism and a region proposal network to achieve accurate recognition of human behavior patterns in the platform environment; the specific structure of the rail scenario attention network model is a spatio-temporal dual attention network based on an improved Transformer architecture. The backbone network uses a deep residual structure to extract image features. The temporal attention module captures the temporal features of target behaviors, and the spatial attention module focuses on key areas of the platform. The integrated region proposal network improves the detection ability of small targets. The number of attention heads in the multi-head attention mechanism is dynamically adjusted according to the number of resolution levels output by the resolution level evaluation function, and the dimension of the attention matrix is associated with the abnormal target change index threshold.

[0013] Among them, the platform monitoring area is divided into four time periods: the vehicle approaching stage, the vehicle docking stage, the vehicle leaving stage, and the platform idle stage, and an application scenario adaptive analysis strategy is applied.

[0014] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions run on a computer, they are used to execute the above-mentioned rail transit video intelligent analysis method.

[0015] The third aspect of the present invention provides a rail transit video intelligent analysis system, which includes the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.

[0016] The present invention successfully solves the technical problems of large computational complexity and low efficiency in high-definition video processing by constructing an adaptive resolution image hierarchical structure and a hierarchical processing strategy. This method dynamically determines the optimal number of resolution levels based on the minimum spanning tree algorithm, and generates a resolution pyramid from the original high-definition video frames through a recursive binary downsampling method, realizing the reasonable allocation and utilization of computing resources.

[0017] Compared with the traditional full-resolution processing method, the hierarchical processing strategy adopted by the present invention significantly reduces the computational complexity. The low-resolution levels are responsible for global rapid screening, and the high-resolution levels focus on fine analysis of key areas, greatly improving the processing efficiency while ensuring the detection accuracy. Especially when processing high-definition videos of 4K and above, the consumption of computing resources is effectively controlled, the system response speed is significantly accelerated, and real-time detection and early warning of abnormal behaviors are realized.

[0018] The present invention also enhances the recognition ability of the system in complex environments by introducing a scene adaptive analysis strategy and an orbital scene attention network model. The analysis strategy is adjusted according to different operation stages of rail transit vehicles, and combined with a spatio-temporal dual attention mechanism, the accuracy of abnormal behavior recognition is improved under the condition of limited computing resources. Through multi-parameter comprehensive risk assessment, the system can timely detect potential safety hazards and predict their development trends, providing a scientific basis for safety management. This system design that takes into account both efficiency and accuracy effectively solves the technical problems of large computational volume and low efficiency in high-definition video processing, and improves the overall performance of the rail transit safety monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0021] As Figure 1 shown, it is a flowchart of an intelligent analysis method for rail transit videos provided in the first aspect of the present invention, and this method includes the following steps:

[0022] S01. Construct an adaptive resolution image hierarchical structure, dynamically determine the optimal number of resolution levels based on the minimum spanning tree algorithm, generate a resolution pyramid for the original high-definition video frame through a recursive binary downsampling method, and determine the resolution of each level through a resolution level evaluation function;

[0023] S02. Calculate the difference between adjacent video frames using an improved frame difference method, construct an abnormal target movement matrix, improve the accuracy of motion detection by combining optical flow estimation technology, and screen out regions with significant changes through an adaptive threshold;

[0024] S03. Based on the abnormal target movement matrix, calculate the abnormal target movement displacement matrix, and apply the Hungarian algorithm to solve the multi-target matching problem in target tracking to ensure accurate association of the position changes of the same target in consecutive frames in a complex scene;

[0025] S04. According to the abnormal target movement displacement matrix, construct an abnormal target change index, comprehensively evaluate the abnormal degree of the target behavior through a multi-parameter weighted calculation function, and introduce spatio-temporal context information to enhance the accuracy of abnormal detection;

[0026] S05. According to the operation state of the rail transit vehicle, divide the platform monitoring area into four time periods: the vehicle entering the station stage, the vehicle docking stage, the vehicle leaving the station stage, and the platform idle stage, and apply the scene adaptive analysis strategy;

[0027] S06. Introduce a pre-trained orbital scene attention network model for abnormal target recognition. The orbital scene attention network model integrates a spatio-temporal attention mechanism and a region proposal network to achieve accurate recognition of the behavior patterns of people in the platform environment;

[0028] S07. Adopt an abnormal behavior risk assessment function to comprehensively calculate the distance parameter between the current abnormal target and the orbital safety boundary, the target movement speed parameter, the target behavior pattern similarity parameter, the platform congestion parameter, and the vehicle operation stage parameter, and quantify the abnormal behavior risk level to output a normalized risk score and the corresponding warning level;

[0029] S08. Based on a hierarchical processing strategy, perform parallel calculations on images at different resolution levels. The low-resolution level is responsible for global rapid screening, and the high-resolution level is responsible for fine analysis of key areas to optimize the allocation of computing resources;

[0030] S09. According to the detection results and the abnormal target change index, divide the warning levels according to the preset risk levels, and estimate the development trend of potential risks through the rail transit abnormal event prediction model to achieve early risk intervention.

[0031] Among them, the abnormal target movement matrix refers to a two-dimensional array generated by calculating the pixel difference between adjacent frames of a video, which represents the position information of moving objects in the picture, and each element value represents the change degree of the pixel point at the corresponding position.

[0032] Among them, the abnormal target movement displacement matrix refers to the amount of position change of the detected abnormal target in multiple consecutive frames of video, including two components: horizontal displacement and vertical displacement, which are used to analyze the target movement trajectory.

[0033] Among them, the abnormal target change index refers to a weighted calculated value that comprehensively considers the target size change, speed change, and direction change. The higher the value, the more abnormal the target behavior. The calculation formula is the target size change rate multiplied by the weight coefficient plus the movement speed change rate multiplied by the weight coefficient plus the direction change rate multiplied by the weight coefficient.

[0034] Among them, the resolution level evaluation function is used to dynamically determine the optimal number of resolution division levels and the resolution parameters of each level. The inputs include the original video resolution, the computing resource limit parameter, the target detection accuracy requirement parameter, the real-time requirement parameter, and the scene complexity score, and the outputs are the optimal number of resolution levels and the specific downsampling ratio of each level.

[0035] Among them, the abnormal behavior risk assessment function is used to quantitatively calculate the risk level of the detected target behavior in the platform. The inputs include the distance parameter between the target and the track safety boundary, the target movement speed parameter, the target behavior pattern similarity parameter, the platform congestion parameter, and the vehicle operation stage parameter. The outputs are the normalized risk score and the corresponding warning level.

[0036] Among them, the specific structure of the track scene attention network model is a spatio-temporal dual attention network based on the improved Transformer architecture. The backbone network uses a deep residual structure to extract image features. The time attention module captures the temporal features of the target behavior. The space attention module focuses on the key areas of the platform. The integrated region proposal network improves the small target detection ability. The number of attention heads in the multi-head attention mechanism is dynamically adjusted according to the number of resolution levels output by the resolution level evaluation function. The dimension of the attention matrix is associated with the abnormal target change index threshold.

[0037] Among them, the steps for establishing the training dataset in the training process of the track scene attention network model specifically include collecting video data of multiple urban rail transit platforms, classifying and annotating them according to the platform congestion, lighting conditions, and platform layout types, constructing a balanced dataset containing normal behavior samples and abnormal behavior samples, performing data augmentation processing for the unique scenarios of rail transit, guiding the annotation of high-risk behavior patterns through expert knowledge, generating spatio-temporal behavior trajectory maps as auxiliary training signals, and establishing a behavior pattern knowledge base to support few-shot learning.

[0038] Among them, the steps for training the track scene attention network model specifically include adopting a multi-stage training strategy. First, train on a large-scale general behavior recognition dataset to obtain the basic feature expression ability. Then, perform domain adaptation fine-tuning on the rail transit scene dataset. Next, introduce a contrast learning mechanism to improve the sensitivity of the model to abnormal behaviors. Adopt a curriculum learning strategy to gradually increase the difficulty of training samples. Extract a lightweight model from a complex model through knowledge distillation technology. Finally, perform end-to-end fine-tuning to optimize the model performance.

[0039] Among them, the track safety boundary distance parameter refers to the shortest distance value between the abnormal target and the boundary of the dangerous area of the rail transit, which is used to evaluate the possibility of the target entering the dangerous area.

[0040] Among them, the target movement speed parameter refers to the ratio of the target displacement calculated from the abnormal target movement displacement matrix to the time, which represents the speed of the target movement.

[0041] Among them, the target behavior pattern similarity parameter refers to the calculation result of the similarity between the current detected target behavior characteristics and the high-risk behavior patterns stored in the track scene attention network model.

[0042] Among them, the platform congestion parameter refers to the ratio of the target density in the platform area to the designed carrying capacity of the platform, reflecting the current density of the crowd on the platform.

[0043] Among them, the vehicle operation stage parameter refers to the numerical representation of the current operation state of the rail transit vehicle, including four discrete state values: entering the station, stopping, leaving the station, and being idle.

[0044] Among them, the rail transit abnormal event prediction model refers to a time series prediction model trained based on historical abnormal event data, used to estimate the development trend of abnormal events according to the current abnormal target change index and the normalized risk score output by the abnormal behavior risk assessment function.

[0045] The following describes the specific implementation manners of the above steps in detail. The specific implementation manner of step S01 is to construct an adaptive resolution image hierarchical structure. First, dynamically evaluate the original high-definition video resolution through the minimum spanning tree algorithm, and calculate the optimal number of resolution levels. Specifically, when implementing, regard the original video frame as a graph node, and the calculation cost and accuracy loss between adjacent resolutions as edge weights, and use the Kruskal algorithm to construct a minimum spanning tree. The depth of the tree is the optimal number of levels, usually set to 3 - 5 levels. Then generate a resolution pyramid through the recursive binary downsampling method. After applying the Gaussian filter to each level for smoothing processing, perform 2×2 regional pixel averaging downsampling to generate an image sequence with gradually decreasing resolutions. Finally, determine the specific resolution parameters of each level through the resolution level evaluation function. The inputs of this function include the original video resolution (such as 1920×1080), the calculation resource limit parameter (such as the upper limit of the number of pixels that can be processed per second 10 8 ), the target detection accuracy requirement parameter (such as the proportion of the minimum detectable target size in the picture 0.5%), the real-time requirement parameter (such as the maximum allowable processing delay 30 milliseconds), and the scene complexity score (a normalized value from 1 to 10), and output the optimal number of resolution levels and the specific downsampling ratio of each level. The function of this step is to create a multi-resolution processing structure and lay a foundation for subsequent efficient analysis.

[0046] The specific implementation manner of step S02 is to detect abnormal moving targets by using the improved frame difference method combined with the optical flow estimation technique. Specifically, when implementing, first perform Gaussian filter preprocessing on adjacent video frames I t and I t-1 to reduce noise interference. Then calculate the inter-frame difference D t (x, y) = |I t (x, y) - I t-1 (x, y)| to generate a difference image. Then apply the adaptive threshold T a = μ D + k·σ D to perform binary processing on the difference image, where μD is the mean value of the difference image, and σ D is the standard deviation, k is the adjustment coefficient, and the recommended value is 2.5 - 3.5. At the same time, the Lucas-Kanade optical flow estimation algorithm is used to calculate the pixel point motion vector field, and the motion direction and speed information are extracted. The frame difference result is fused with the optical flow information, and through weighted combination M t (x, y) = w1·B t (x, y) + w2·|V t (x, y)| to construct the abnormal target movement matrix, where B t (x, y) is the binarized frame difference result, and V t (x, y) is the magnitude of the optical flow vector, and w1 and w2 are weight coefficients, and the recommended values are 0.6 and 0.4 respectively. Finally, morphological operations are applied to the movement matrix for optimization. The opening operation is used to eliminate noise, and the closing operation is used to fill the internal holes of the target. The role of this step is to accurately extract the moving areas in the video frame and provide a basis for abnormal target detection.

[0047] The specific implementation of step S03 is to calculate the abnormal target movement displacement matrix based on the abnormal target movement matrix and perform multi-target tracking. When specifically implemented, first, the connected component analysis algorithm is applied to the abnormal target movement matrix to extract and label each independent moving target, and the characteristic parameters such as the centroid coordinates, area, and bounding box of each target are calculated. Then, a target feature description vector is established, including position, size, color histogram, and texture features. Next, a target matching cost matrix is constructed between consecutive video frames. The matrix element C ij represents the matching cost between the i-th current target and the j-th historical target, and the calculation formula is C ij = w p ·d pos + w s ·d size + w a ·d app , where d pos is the position distance, d size is the size difference, d app is the appearance similarity, and w p , w s , w a are weight coefficients, and the recommended values are 0.5, 0.3, and 0.2 respectively. The matching problem is transformed into an assignment problem, and the Hungarian algorithm is applied to solve it to obtain the optimal matching result. Based on the matching result, the displacement vector of each target between consecutive frames is calculated to form the abnormal target movement displacement matrix, and the displacement amounts in the horizontal and vertical directions are recorded. The role of this step is to achieve the temporal association of the target and provide trajectory data for subsequent behavior analysis.

[0048] The specific implementation of step S04 is to construct an abnormal target change index based on the displacement matrix of the abnormal target. When specifically implemented, first calculate the target size change rate R size =(A t -A t-1 ) / A t-1 , where A t and A t-1 are the target areas in the current frame and the previous frame respectively. Then calculate the moving speed change rate R speed =(V t -V t-1 ) / V t-1 , where V t and V t-1 are the target speed amplitudes. Next, calculate the direction change rate R dir =θ t -θ t-1 / π, where θ t and θ t-1 are the target motion direction angles. Combining the above parameters, use the multi-parameter weighted calculation function I change =w1·|R size |+w2·|R speed |+w3·|R dir | to calculate the abnormal target change index, where w1, w2, and w3 are weight coefficients, and the recommended values are 0.3, 0.4, and 0.3 respectively. Introduce spatio-temporal context information for adjustment, construct a spatio-temporal memory module to record the target historical behavior patterns, calculate the deviation degree between the current behavior and the historical patterns, and correct the change index. Set the change index threshold T change =0.4, and if it exceeds the threshold, it is marked as a potential abnormal behavior. The function of this step is to quantify the abnormal degree of the target behavior and realize the preliminary detection of abnormal behaviors.

[0049] The specific implementation of step S05 is to divide the monitoring stage according to the operating state of rail transit vehicles and apply a scenario-adaptive analysis strategy. When specifically implemented, first, a vehicle state recognition module is established to determine the current operating state of the vehicle by analyzing video content or receiving signals from an external dispatching system. Then, the platform monitoring area is divided into four time periods: the vehicle approaching stage (when the train is about to enter the platform), the vehicle docking stage (when the train is docked at the platform and passengers are allowed to get on and off), the vehicle leaving stage (when the train is about to leave the platform), and the platform idle stage (when there is no train on the platform). Next, different analysis strategies are designed for each stage: in the approaching stage, the edge area of the platform is mainly monitored to detect whether the waiting crowd outside the yellow line has crossed the line, and the focus is set on the yellow line of the platform and the edge area of the platform; in the docking stage, the boarding and alighting areas are mainly monitored to identify abnormal situations such as passengers staying at the door and falling, and the focus is set on the area around the door; in the leaving stage, the gap between the person and the vehicle is mainly monitored to prevent passengers from chasing the train, and the focus is set on the gap between the edge of the platform and the train; in the idle stage, the entire platform area is comprehensively monitored to detect abnormal stay, suspicious behavior, etc., and the focus covers the entire platform area. Through the scenario-adaptive analysis strategy, the detection parameters and alarm thresholds are dynamically adjusted to improve the pertinence and accuracy of the system. The function of this step is to achieve an accurate match between the monitoring strategy and the actual scenario.

[0050] The specific implementation of step S06 is to introduce a pre-trained rail scene attention network model for abnormal target recognition. When specifically implemented, first, the pre-trained rail scene attention network model is loaded, and this model adopts a spatio-temporal dual attention network structure based on the Transformer architecture. Then, the detected abnormal target area is cropped and adjusted to the standard input size (such as 224×224 pixels) and input into the model for feature extraction and behavior recognition. The model captures temporal features through the temporal attention module, processes consecutive multi-frame (usually 16 - 32 frames) target images, and extracts action patterns. At the same time, the spatial attention module focuses on the key areas inside the target, such as human key points or characteristic parts. The integrated region proposal network improves the detection ability of small targets, generates candidate regions and classifies them. The model output includes the prediction result of the behavior category and the confidence score, and the behavior categories include normal behaviors (such as standing, walking, waiting for the train) and abnormal behaviors (such as climbing over the guardrail, staying at the edge of the track, falling onto the platform, etc.). The confidence threshold for behavior recognition is set to 0.75, and if it exceeds the threshold, it is confirmed as the corresponding behavior category. The function of this step is to accurately identify the behavior patterns of people in the platform environment and provide support for abnormal behavior recognition.

[0051] The specific implementation of step S07 is to use an abnormal behavior risk assessment function to quantitatively calculate the behavior risk level of the detected targets in the platform. When specifically implemented, first, calculate the distance parameter D between the target and the rail safety boundary track, measure the vertical distance from the centroid of the target to the safety boundary of the platform, and normalize it to a value between 0 and 1. The smaller the distance, the higher the risk. Then calculate the target movement speed parameter V obj , extract displacement data from the abnormal target movement displacement matrix, calculate the amplitude of the velocity vector and normalize it. Then calculate the target behavior pattern similarity parameter S pattern , compare the current target behavior characteristics with the predefined high-risk behavior pattern library, and use cosine similarity to calculate the degree of similarity. Then calculate the platform congestion parameter C platform , count the ratio of the number of targets in the unit area to the preset safety capacity. Finally, combine the vehicle operation stage parameter P phase (inbound is 0.9, docking is 0.7, outbound is 0.8, idle is 0.5), use the risk assessment function R risk = w1·(1 - D track ) + w2·V obj + w3·S pattern + w4·C platform + w5·P phase Calculate the normalized risk score, where w1 to w5 are weight coefficients, and the recommended values are 0.35, 0.25, 0.2, 0.1, and 0.1 respectively. Set the risk level thresholds: 0 - 0.3 is low risk (green warning), 0.3 - 0.6 is medium risk (yellow warning), 0.6 - 0.8 is high risk (orange warning), 0.8 - 1 is extremely high risk (red warning). The purpose of this step is to achieve the quantitative assessment and level classification of abnormal behavior risks.

[0052] The specific implementation of step S08 is to perform parallel computing on images at different resolution levels based on a hierarchical processing strategy. Specifically, when implementing, first, according to the resolution pyramid constructed in step S01, distribute images at different resolution levels to parallel processing units. Then, perform global fast screening at the low resolution level (such as 1 / 4 or 1 / 8 of the original resolution), and apply a lightweight object detection algorithm (such as improved YOLOv3 or SSD) for rough detection to quickly locate potential abnormal areas. At the same time, perform preliminary analysis on the identified key areas at the medium resolution level (such as 1 / 2 of the original resolution) to extract more detailed target features. Finally, perform fine analysis on the areas determined to require key attention at the high resolution level (original resolution), and apply a high-precision but computationally complex algorithm (such as improved Faster R-CNN) for accurate identification. Adopt a result fusion strategy to integrate the processing results of each level through a weighted voting mechanism to improve the detection reliability. Dynamically adjust the proportion of computing resource allocation for each level, and adaptively allocate computing resources according to the current scene complexity and the urgency of abnormal events to ensure that key areas obtain sufficient processing capabilities. The purpose of this step is to optimize the computing resource allocation and balance the requirements of real-time performance and accuracy.

[0053] The specific implementation of step S09 is to divide the warning level according to the detection result and the abnormal target change index and conduct risk prediction. When specifically implemented, first, summarize the foregoing analysis results, including the abnormal target change index and the abnormal behavior risk score. Then, divide the warning level according to the preset risk threshold to generate a warning signal corresponding to the level. Next, call the rail transit abnormal event prediction model for trend prediction. This model adopts a long short-term memory network (LSTM) structure, and the inputs include the current abnormal target change index time series (usually data in the past 10 - 30 seconds) and the abnormal behavior risk score sequence, and the outputs are the predicted value and confidence interval of the risk score within the next 5 - 15 seconds. Judge the risk development trend according to the prediction result. If the predicted risk score shows an upward trend and exceeds the warning threshold, trigger the warning upgrade mechanism. At the same time, generate a risk warning report, including information such as the location of the abnormal target, the type of behavior, the risk level, and the prediction trend, and take different response measures according to the warning level: continue to monitor for low risk, notify the on-duty personnel for medium risk, start the emergency plan for high risk, and trigger automatic intervention measures (such as broadcast reminder or notify the train to slow down) for extremely high risk. The function of this step is to achieve timely warning and early intervention of risks and prevent problems before they occur.

[0054] The specific structure of the rail scene attention network model adopts a spatio-temporal dual attention network based on an improved Transformer architecture. In terms of structure design, the backbone network of the model uses a deep residual structure (ResNet50 or ResNet101) to extract image features, and uses 3×3, 5×5, 7×7 multi-scale convolutional filters to extract multi-scale spatial features in parallel to capture target features of different sizes. The time attention module is based on the self-attention mechanism, processes a continuous 16-frame input sequence, calculates the frame-interval relationship weight matrix, and strengthens the feature expression at critical moments. The module internally includes a position encoding layer, a multi-head self-attention layer, and a feed-forward network layer. The position encoding uses sine and cosine functions to generate the temporal position representation. The spatial attention module constructs a pixel-level attention map to highlight the key areas of the platform (such as the platform edge, escalator entrance, and door docking location). The module uses a non-local neural network structure to calculate the spatial position correlation. The region proposal network part uses anchor box scales of {32 2 , 64 2 , 128 2 , 256 2 , 512 2}, and aspect ratios of {0.5, 1.0, 2.0} to improve the small target detection ability. In the multi-head attention mechanism, the number of attention heads is dynamically adjusted according to the number of resolution levels output by the resolution level evaluation function, usually set to 4 - 8 heads, and each head independently learns different feature subspace representations. The output features of each module are integrated through a feature fusion network to generate a target behavior feature vector, and the behavior classifier uses a fully connected layer structure to output the probability distribution of the behavior category.

[0055] The specific implementation of the steps for establishing the training dataset of the orbital scene attention network model includes links such as data collection and annotation, data classification and enhancement, expert knowledge-guided annotation, construction of auxiliary training signals, and establishment of a knowledge base. In the data collection and annotation stage, video data of multiple urban rail transit platforms are collected, covering different time periods (morning rush hour, evening rush hour, flat peak period), different weather conditions (sunny, rainy, foggy), and different lighting environments (daytime, dusk, night). The sampling rate is 25 frames per second, the resolution is 1920×1080, and the continuous acquisition duration of a single platform is not less than 72 hours. In the data classification and annotation stage, classification and annotation are carried out according to the platform congestion level (low, medium, high levels), lighting conditions (standard lighting, low light lighting, strong light lighting), and platform layout types (side platform, island platform, compound platform), and a dataset containing 90% normal behavior samples and 10% abnormal behavior samples is constructed. In the data enhancement processing stage, brightness adjustment (±20%), contrast transformation (±15%), random cropping (80% - 100% of the original size), horizontal flipping, rotation (±10°), etc. are performed for the unique scenes of rail transit to expand the dataset size to 5 times the original data. In the expert knowledge-guided annotation stage, rail transit safety experts are invited to annotate high-risk behavior patterns, including abnormal behaviors such as approaching the platform edge, running fast, climbing over the guardrail, and lying motionless, and a temporal behavior pattern description is established for each behavior. In the stage of constructing auxiliary training signals, a spatio-temporal behavior trajectory map is generated, recording the movement trajectory, speed change, and behavior state transition of the target in the platform, as an auxiliary supervision signal for model training. In the stage of establishing a knowledge base, a behavior pattern knowledge base is constructed, including typical abnormal behavior descriptions defined by rules and behavior pattern embedding vectors driven by data, supporting few-shot learning and zero-shot recognition capabilities.

[0056] The second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored, and when the program instructions run on a computer, they are used to execute the above-mentioned method for intelligent analysis of rail transit videos.

[0057] The third aspect of the present invention provides a rail transit video intelligent analysis system, including the above-mentioned computer-readable storage medium. The system can be any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is set inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is set inside the system.

[0058] The following describes in detail the mathematical models or calculation processes involved in the present invention.

[0059] In step S02, the calculation process of detecting abnormal moving targets by combining the improved frame difference method with the optical flow estimation technique is specifically expressed as follows:

[0060] D t (x, y) = |I t (x, y) - I t-1 (x, y)|;

[0061] In the formula, D t (x, y) is the pixel value of the difference image; I t (x, y) is the pixel value of the current frame at the coordinate (x, y); I t-1 (x, y) is the pixel value of the previous frame at the same coordinate; |·| represents the absolute value operation.

[0062] The formula for the adaptive threshold is:

[0063] T a = μ D + k·σ D ;

[0064] In the formula, T a is the adaptive threshold; μ D is the mean value of the difference image D t ; σ D is the standard deviation of the difference image D t ; k is the adjustment coefficient, and its value range is 2.5 to 3.5.

[0065] The mean value μ D of the difference image is calculated by the following formula:

[0066]

[0067] In the formula, W and H are the width and height of the image respectively; represents the summation over the entire image area.

[0068] The standard deviation σ D is calculated by the following formula:

[0069]

[0070] In the formula, represents the square root operation.

[0071] The formula for binarization is:

[0072]

[0073] In the formula, B t (x, y) is the binarized image, where the value 1 indicates that motion is detected at this pixel point, and the value 0 indicates that no motion is detected at this pixel point.

[0074] The calculation of the optical flow vector field uses the Lucas-Kanade algorithm, and its basic principle is to solve the optical flow constraint equation:

[0075]

[0076] In the formula, and are the gradients of the image in the x and y directions respectively; is the gradient in the time direction; u and v are the optical flow velocity components in the horizontal and vertical directions respectively.

[0077] By assuming that the optical flow velocity is consistent within a small window area and applying the least squares method to solve the overdetermined system of equations, the optical flow vector field V t (x, y) is obtained:

[0078]

[0079] In the formula, V t (x, y) represents the amplitude of the optical flow vector at the coordinates (x, y); u(x, y) and v(x, y) are the optical flow velocity components at this point in the horizontal and vertical directions respectively.

[0080] Finally, the abnormal target movement matrix is calculated through weighted combination:

[0081] M t (x, y) = w1·B t (x, y) + w2·V t (x, y);

[0082] In the formula, M t (x, y) is the element of the abnormal target movement matrix; w1 and w2 are weight coefficients, and the recommended values are 0.6 and 0.4 respectively. The purpose of weighted combination is to fuse the robustness of the frame difference method and the accuracy of the optical flow method to improve the accuracy of motion detection.

[0083] In step S03, the calculation formula of the target matching cost matrix is:

[0084] C ij = w p ·d pos (i, j) + w s ·d size (i, j) + w a ·d app (i, j);

[0085] In the formula, C ij represents the matching cost between the i-th current target and the j-th historical target; d pos (i, j) is the position distance; d size (i, j) is the size difference; d app(i, j) is the appearance similarity; w p 、w s 、w a are weight coefficients, and the recommended values are 0.5, 0.3, and 0.2 respectively.

[0086] The position distance d pos (i, j) is calculated as follows:

[0087]

[0088] In the formula, (x i , y i ) and (x j , y j ) are the centroid coordinates of the current target i and the historical target j respectively; d max is the normalization factor, usually set to 0.2 times the length of the image diagonal.

[0089] The size difference d size (i, j) is calculated as follows:

[0090]

[0091] In the formula, A i and A j are the areas of the current target i and the historical target j respectively; min(·, ·) and max(·, ·) represent the operations of taking the minimum value and the maximum value respectively.

[0092] The appearance similarity d app (i, j) is calculated using histogram similarity:

[0093]

[0094] In the formula, H i and H j are the color histograms of the current target i and the historical target j respectively; n is the number of bins of the histogram, usually set to 64.

[0095] The Hungarian algorithm is used to solve the following optimization problem:

[0096]

[0097] In the formula, X ij is the assignment variable, which is 1 when the i-th current target is matched to the j-th historical target, and 0 otherwise; m and n are the numbers of targets in the current frame and the historical frame respectively.

[0098] Based on the matching result, the abnormal target movement displacement matrix D t is calculated as follows:

[0099]

[0100] In the formula, D t (i) is the displacement vector of the i-th target; Δx i and Δy i are the displacement amounts in the horizontal and vertical directions respectively; and are the centroid coordinates of the target in the current frame and the previous frame respectively.

[0101] In step S04, the calculation process of the abnormal target change index is specifically expressed as follows:

[0102] Calculation formula for the target size change rate:

[0103]

[0104] In the formula, R size (i) is the size change rate of the i-th target; and are the areas of the target in the current frame and the previous frame respectively.

[0105] Calculation formula for the moving speed change rate:

[0106]

[0107] In the formula, R speed (i) is the speed change rate of the i-th target; and are the speed amplitudes of the target in the current frame and the previous frame respectively, calculated through the displacement vector: where Δt is the time interval between adjacent frames.

[0108] Calculation formula for the direction change rate:

[0109]

[0110] In the formula, R dir (i) is the direction change rate of the i-th target; and are the motion direction angles of the target in the current frame and the previous frame respectively, calculated through the displacement vector: where arctan2 is the four-quadrant arctangent function.

[0111] Comprehensive calculation formula for the abnormal target change index:

[0112] I change (i) = w1·|R size (i)| + w2·|R speed (i)| + w3·|R dir (i)| + ∈;

[0113] In the formula, I change (i) is the abnormal change index of the i-th target; w1, w2, and w3 are weight coefficients, and the recommended values are 0.3, 0.4, and 0.3 respectively; |·| represents the absolute value operation; ∈ is the noise correction term, and its value range is 0.05 to 0.1, which is used to prevent the denominator from being zero and enhance the robustness of the algorithm.

[0114] The calculation formula of the abnormal change index after introducing spatio-temporal context correction:

[0115]

[0116] In the formula, is the corrected abnormal change index; D hist (i) is the deviation degree between the target behavior and the historical behavior pattern, and the range is 0 to 1; α is the adjustment coefficient, and the recommended value is 0.5.

[0117] The deviation degree D hist (i) is calculated by a recursive update method:

[0118]

[0119] In the formula, is the historical deviation value at the previous moment; D curr (i) is the difference between the current observation and the historical pattern; β is the smoothing coefficient, and the recommended value is 0.7.

[0120] In step S07, the calculation process of the abnormal behavior risk assessment function is specifically expressed as follows:

[0121] The calculation formula of the distance parameter between the target and the orbit safety boundary:

[0122]

[0123] In the formula, D track (i) is the orbit safety boundary distance parameter of the i-th target; d i is the vertical distance (in meters) from the target centroid to the platform safety boundary; d safe is the safety distance threshold, and the recommended value is 1.5 meters.

[0124] The calculation formula of the target moving speed parameter:

[0125]

[0126] In the formula, V obj (i) is the moving speed parameter of the i-th target; v i is the target speed amplitude (in meters per second); v max is the speed threshold, and the recommended value is 3 meters per second.

[0127] Calculation formula for the similarity parameter of the target behavior pattern:

[0128]

[0129] In the formula, S pattern (i) is the similarity parameter of the behavior pattern of the i-th target; f i is the current behavior feature vector of the target; is the feature vector of the j-th predefined high-risk behavior pattern; k is the number of patterns in the risk behavior pattern library; · represents the vector inner product operation; |·| represents the vector norm operation.

[0130] Calculation formula for the platform crowding degree parameter:

[0131]

[0132] In the formula, C platform is the platform crowding degree parameter; n obj is the number of targets in the current platform area; A platform is the effective area of the platform (square meters); ρ safe is the safety density threshold, and the recommended value is 1.5 people per square meter.

[0133] Discrete value settings for the vehicle operation stage parameter P phase are as follows:

[0134] In the approach stage: P phase = 0.9;

[0135] In the docking stage: P phase = 0.7;

[0136] In the departure stage: P phase = 0.8;

[0137] In the idle stage: P phase = 0.5;

[0138] Comprehensive calculation formula for the abnormal behavior risk assessment function:

[0139] R risk (i) = w1·(1 - D track (i)) + w2·V obj (i) + w3·S pattern (i) + w4·C platform + w5·P phase + δ;

[0140] In the formula, R risk(i) is the normalized risk score for the i-th target; w1 to w5 are weight coefficients, with recommended values of 0.35, 0.25, 0.2, 0.1, and 0.1 respectively; δ is a random perturbation term, with a value range of -0.05 to 0.05, used to prevent over-certainty in risk assessment.

[0141] The abnormal behavior risk assessment function adopts a weighted linear combination form, considering multiple factors: the track distance parameter uses a reciprocal relationship (1 - D track ) because the smaller the distance, the higher the risk, which conforms to the characteristics of rail transit safety; the speed parameter uses a linear positive correlation because high-speed movement usually means higher risk in the platform environment; the cosine similarity is used to calculate the similarity of behavior patterns to capture the pattern matching of behavior characteristics in the high-dimensional space; the crowding degree parameter and the vehicle operation stage parameter, as environmental factors, affect the overall risk assessment through a linear combination method.

[0142] In the rail transit abnormal event prediction model, the risk trend prediction uses an LSTM network, and the prediction equation is as follows:

[0143]

[0144] In the formula, is the predicted risk score at the future time t + Δt; R t-kτ (i) and are the risk score at the past time t - kτ and the corrected abnormal change index respectively; τ is the sampling time interval; T is the length of the historical sequence; f LSTM is the LSTM network model function.

[0145] Prediction confidence interval calculation formula:

[0146]

[0147] In the formula, R lower (i) and R upper (i) are the lower bound and upper bound of the predicted risk score respectively; z α / 2 is the critical value of the normal distribution, and when the confidence level is 95%, the value is approximately 1.96; σ R is the standard deviation of the predicted risk score, estimated through historical prediction errors.

[0148] Optionally, the calculation process of the resolution level evaluation function is specifically expressed as follows:

[0149]

[0150] In the formula, L opt is the number of optimal resolution levels; L max is the maximum allowable number of levels, usually set to 8; C comp(L) is the computing resource consumption function; E acc (L) is the detection accuracy loss function; T real (L) is the real-time performance loss function; S complex is the scene complexity score; w1, w2, w3, w4 are weight coefficients, and the recommended values are 0.3, 0.3, 0.3, 0.1 respectively.

[0151] Optionally, the calculation formula of the computing resource consumption function:

[0152]

[0153] In the formula, W0 and H0 are the width and height (pixels) of the original video frame respectively; 4 -i represents the resolution reduction factor of the i-th layer; c i is the unit pixel computing complexity coefficient of the i-th layer; C max is the computing resource upper limit parameter.

[0154] Optionally, the calculation formula of the detection accuracy loss function:

[0155]

[0156] In the formula, s min is the minimum detectable target size (pixel area); 0.005·W0·H0 is the target detection accuracy requirement parameter, indicating the pixel area corresponding to the minimum detectable target accounting for 0.5% of the screen area.

[0157] Optionally, the calculation formula of the real-time performance loss function:

[0158]

[0159] In the formula, t i is the processing time required for the i-th layer (milliseconds); T max is the maximum allowable processing delay parameter (milliseconds).

[0160] Optionally, the calculation formula for determining the specific downsampling ratio of each layer:

[0161]

[0162] In the formula, r i is the downsampling ratio of the i-th layer; r0 = 1 represents the original resolution; r min is the minimum resolution ratio, usually set to 0.125; represents the ceiling operation.

[0163] Optionally, the actual resolution calculation formula of each layer:

[0164]

[0165] Wherein, W i and H i are respectively the width and height of the i-th layer image; represents the floor operation.

[0166] The design of the resolution level evaluation function follows the principle of balancing computational efficiency and detection accuracy, and solves the optimal layer configuration through a multi-objective optimization method. The function adopts the form of weighted summation, considering various factors such as computational resource limitations, target detection accuracy requirements, real-time requirements, and scene complexity, and realizes the adaptive optimization of the resolution level.

[0167] Specifically, the principle of the present invention is: The core technical principle of the present invention to solve the problems of large computational amount and low efficiency in high-definition video processing lies in adopting the idea of "divide and conquer", and optimizing the computational resource allocation through a multi-resolution level collaborative analysis framework. First, the invention introduces an adaptive resolution image hierarchical structure, and dynamically determines the optimal number of resolution levels based on the minimum spanning tree algorithm, avoiding the problems of resource waste or insufficient accuracy caused by fixed levels. The resolution level evaluation function comprehensively considers factors such as the original resolution of the video, computational resource limitations, and target detection accuracy requirements, and scientifically determines the specific downsampling ratio for each layer, providing a mathematical basis for the system.

[0168] In terms of computational resource allocation, the present invention adopts a hierarchical processing strategy to achieve parallel computing, and improves the processing efficiency through task decomposition. The low-resolution level uses an algorithm with a small computational amount to perform global rapid screening, and can locate potential abnormal areas in a short time; the high-resolution level focuses on the fine analysis of key areas, only processes valuable image parts, and avoids the full-scale calculation of the entire high-definition picture, significantly reducing the computational complexity. This resource allocation strategy enables the system to process high-resolution videos under limited computational conditions, achieving the balance between computational efficiency and detection accuracy.

[0169] Another technical innovation point of the present invention is to introduce an improved target detection and tracking algorithm, which improves the motion detection accuracy through optical flow estimation technology, applies the Hungarian algorithm to solve the multi-target matching problem, and constructs an abnormal target movement matrix and a displacement matrix. These optimized algorithms improve the detection accuracy while reducing the computational amount. Especially the track scene attention network model, which is a spatio-temporal dual attention network based on an improved Transformer architecture, can intelligently focus on key areas and behavior features, avoiding redundant calculations of irrelevant information.

[0170] In the risk assessment stage, a quantitative analysis of abnormal behavior risks is achieved through a multi-parameter weighted calculation function. Based on the particularity of the platform environment, the system designs an evaluation model including key parameters such as the distance between the target and the track safety boundary and the moving speed, and completes the risk level determination at a relatively low computational cost. Meanwhile, the rail transit abnormal event prediction model can estimate the risk development trend based on the current data, providing a basis for early intervention. This multi-level and lightweight analysis framework fundamentally solves the technical problems of large computational volume and low efficiency in high-definition video processing, achieving efficient and reliable intelligent analysis of rail transit videos.

[0171] A specific Embodiment 1 of the present invention is provided below. The specific implementation manners of each step in this Embodiment 1 are described in detail as follows.

[0172] The specific implementation manner of step S01 is to construct an adaptive resolution image hierarchical structure. First, the dynamic evaluation of the original high-definition video resolution is performed through the minimum spanning tree algorithm, and the optimal number of resolution levels is calculated. In specific implementation, the original video frames are regarded as graph nodes, and the computational cost and accuracy loss between adjacent resolutions are used as edge weights. The Kruskal algorithm is used to construct the minimum spanning tree, and the depth of the tree is the optimal number of levels, usually set to 3 - 5 levels. The optimal number of resolution levels is determined by the following optimization formula: In the formula, L opt is the optimal number of resolution levels; L max is the maximum allowable number of levels, usually set to 8; C comp (L) is the computational resource consumption function; E acc (L) is the detection accuracy loss function; T real (L) is the real-time performance loss function; S complex is the scene complexity score, a normalized value in the range of 1 - 10; w1, w2, w3, w4 are weight coefficients, and the recommended values are 0.3, 0.3, 0.3, 0.1 respectively. The computational resource consumption function is calculated by the formula In the formula, W0 and H0 are the width and height of the original video frame respectively; 4 -i represents the resolution reduction factor of the i-th layer; c i is the unit pixel computational complexity coefficient of the i-th layer; C max is the computational resource upper limit parameter. The detection accuracy loss function is In the formula, s min is the minimum detectable target size; 0.005·W0·H0 is the target detection accuracy requirement parameter. The real-time performance loss function is In the formula, t i is the processing time required for the i-th layer; T maxis the maximum allowable processing delay parameter. Then, a resolution pyramid is generated through a recursive binary downsampling method. After applying a Gaussian filter to each level for smoothing, 2×2 regional pixel averaging downsampling is performed to generate an image sequence with gradually decreasing resolution. The specific downsampling ratio for each layer is determined by the formula where r i is the downsampling ratio of the i-th layer; r0 = 1 represents the original resolution; r min is the minimum resolution ratio, usually set to 0.125. The actual resolution of each layer is calculated as where W i and H i are the width and height of the i-th layer image respectively. The purpose of this step is to create a multi-resolution processing structure, laying the foundation for subsequent efficient analysis.

[0173] The specific implementation of step S02 is to detect abnormal moving targets by using an improved frame difference method combined with optical flow estimation technology. In specific implementation, first, adjacent video frames I t and I t-1 are preprocessed by Gaussian filtering to reduce noise interference. Then, the inter-frame difference D t (x, y) = |I t (x, y) - I t-1 (x, y)| is calculated to generate a difference image, where D t (x, y) is the pixel value of the difference image; I t (x, y) is the pixel value of the current frame at the coordinate (x, y); I t-1 (x, y) is the pixel value of the previous frame at the same coordinate; |·| represents the absolute value operation. Then, an adaptive threshold T a = μ D + k·σ D is used to perform binary processing on the difference image, where T a is the adaptive threshold; μ D is the mean value of the difference image D t ; σ D is the standard deviation of the difference image D t ; k is an adjustment coefficient, and its value range is 2.5 - 3.5. The mean value μ D of the difference image is calculated by the formula , where W and H are the width and height of the image respectively. The standard deviation σ D is calculated by the formula . The binary processing formula is where B t (x, y) is the binary image. At the same time, the Lucas-Kanade optical flow estimation algorithm is used to calculate the motion vector field of pixel points to solve the optical flow constraint equation In the formula, and are the gradients of the image in the x and y directions respectively; is the gradient in the time direction; u and v are the optical flow velocity components in the horizontal and vertical directions respectively. The magnitude of the optical flow vector is calculated by the formula . The frame difference result is fused with the optical flow information, and an abnormal target movement matrix is constructed through weighted combination M t (x, y) = w1·B t (x, y) + w2·V t (x, y), where M t (x, y) is an element of the abnormal target movement matrix; w1 and w2 are weight coefficients, and the recommended values are 0.6 and 0.4 respectively. Finally, morphological operations are applied to the movement matrix for optimization. The opening operation is used to eliminate noise, and the closing operation is used to fill the internal holes of the target. The purpose of this step is to accurately extract the moving regions in the video frames and provide a basis for abnormal target detection.

[0174] The specific implementation of step S03 is to calculate the abnormal target movement displacement matrix based on the abnormal target movement matrix and perform multi-target tracking. Specifically, when implementing, first apply the connected component analysis algorithm to the abnormal target movement matrix, extract and label each independent moving target, and calculate the characteristic parameters such as the centroid coordinates, area, and bounding box of each target. Then establish a target feature description vector, including position, size, color histogram, and texture features. Next, construct a target matching cost matrix between consecutive video frames. The calculation formula for the matrix elements is C ij = w p ·d pos (i, j) + w s ·d size (i, j) + w a ·d app (i, j), where C ij represents the matching cost between the i-th current target and the j-th historical target; d pos (i, j) is the position distance; d size (i, j) is the size difference; d app (i, j) is the appearance similarity; w p , w s , w a are weight coefficients, and the recommended values are 0.5, 0.3, and 0.2 respectively. The position distance calculation formula is In the formula, (x i , y i ) and (x j , y j ) are the centroid coordinates of the current target i and the historical target j respectively; d maxis the normalization factor, usually set to 0.2 times the length of the image diagonal. The size difference calculation formula is In the formula, A i and A j are the areas of the current target i and the historical target j respectively. The appearance similarity is calculated using histogram similarity: In the formula, H i and H j are the color histograms of the current target i and the historical target j respectively; n is the number of bins of the histogram, usually set to 64. The matching problem is transformed into an assignment problem, and the Hungarian algorithm is applied to solve it to obtain the optimal matching result and solve the optimization problem Satisfy the constraint conditions In the formula, X ij is the assignment variable. Based on the matching result, the displacement vector of each target between consecutive frames is calculated to form an abnormal target movement displacement matrix In the formula, D t (i) is the displacement vector of the i-th target; Δx i and Δy i are the displacement amounts in the horizontal and vertical directions respectively; and are the centroid coordinates of the target in the current frame and the previous frame respectively. The purpose of this step is to achieve the temporal association of the target and provide trajectory data for subsequent behavior analysis.

[0175] The specific implementation of step S04 is to construct an abnormal target change index according to the abnormal target movement displacement matrix. When specifically implemented, first calculate the target size change rate In the formula, R size (i) is the size change rate of the i-th target; and are the areas of the target in the current frame and the previous frame respectively. Then calculate the moving speed change rate In the formula, R speed (i) is the speed change rate of the i-th target; and are the speed amplitudes of the target in the current frame and the previous frame respectively, calculated through the displacement vector: where Δt is the time interval between adjacent frames. Then calculate the direction change rate In the formula, R dir (i) is the direction change rate of the i-th target; and are the motion direction angles of the target in the current frame and the previous frame respectively, calculated through the displacement vector: where arctan2 is the four-quadrant arctangent function. Combining the above parameters, use the multi-parameter weighted calculation function I change(i) = w1·|R size (i)| + w2·|R speed (i)| + w3·|R dir (i)| + ∈ Calculate the abnormal target change index, where change (i) is the abnormal change index of the i-th target; w1, w2, and w3 are weight coefficients, and the recommended values are 0.3, 0.4, and 0.3 respectively; |·| represents the absolute value operation; ∈ is the noise correction term, and the value range is 0.05 - 0.1. Introduce spatio-temporal context information for adjustment and calculate the corrected abnormal change index Where is the corrected abnormal change index; D hist (i) is the deviation degree between the target behavior and the historical behavior pattern, and the range is 0 - 1; α is the adjustment coefficient, and the recommended value is 0.5. The deviation degree is calculated by the recursive update method: Where is the historical deviation value at the previous moment; D curr (i) is the difference between the current observation and the historical pattern; β is the smoothing coefficient, and the recommended value is 0.7. Set the change index threshold T change = 0.4, and if it exceeds the threshold, it is marked as a potential abnormal behavior. The function of this step is to quantify the abnormal degree of the target behavior and realize the preliminary detection of abnormal behavior.

[0176] The specific implementation manners of steps S05 - S06 are the same as those described above and will not be elaborated here.

[0177] The specific implementation manner of step S07 is to use the abnormal behavior risk assessment function to quantitatively calculate the risk level of the detected target behavior in the platform. Specifically, when implementing, first calculate the distance parameter between the target and the track safety boundary Where D track (i) is the track safety boundary distance parameter of the i-th target; d i is the vertical distance (in meters) from the centroid of the target to the platform safety boundary; d safe is the safety distance threshold, and the recommended value is 1.5 meters. Then calculate the target movement speed parameter Where V obj (i) is the movement speed parameter of the i-th target; v i is the amplitude of the target speed (in meters per second); v max is the speed threshold, and the recommended value is 3 meters per second. Then calculate the target behavior pattern similarity parameter Where S pattern (i) is the behavior pattern similarity parameter of the i-th target; f i is the current behavior feature vector of the target; is the feature vector of the j-th predefined high-risk behavior pattern; k is the number of patterns in the risk behavior pattern library. Then calculate the platform congestion parameter In the formula, C platform is the platform congestion parameter; n obj is the number of targets in the current platform area; A platform is the effective area of the platform (square meters); ρ safe is the safety density threshold, and the recommended value is 1.5 people per square meter. Finally, combined with the vehicle operation stage parameter P phase (0.9 for inbound, 0.7 for docking, 0.8 for outbound, 0.5 for idle), use the risk assessment function R risk (i) = w1·(1 - D track (i)) + w2·V obj (i) + w3·S pattern (i) + w4·C platform + w5·P phase + δ to calculate the normalized risk score. In the formula, R risk (i) is the normalized risk score of the i-th target; w1 to w5 are weight coefficients, and the recommended values are 0.35, 0.25, 0.2, 0.1, 0.1 respectively; δ is a random perturbation term, and the value range is -0.05 to 0.05. Set the risk level threshold: 0 - 0.3 is low risk (green warning), 0.3 - 0.6 is medium risk (yellow warning), 0.6 - 0.8 is high risk (orange warning), 0.8 - 1 is extremely high risk (red warning). The role of this step is to realize the quantitative evaluation and level division of abnormal behavior risks.

[0178] The specific implementation manner of step S08 is the same as the foregoing, and will not be elaborated here.

[0179] The specific implementation manner of step S09 is to divide the warning level and perform risk prediction according to the detection result and the abnormal target change index. Specifically, first summarize the foregoing analysis results, including the abnormal target change index and the abnormal behavior risk score. Then divide the warning level according to the preset risk threshold to generate a warning signal of the corresponding level. Then call the rail transit abnormal event prediction model for trend prediction. This model adopts a long short-term memory network structure, and the input includes the current abnormal target change index time series and the abnormal behavior risk score series, and the output is the predicted value of the future risk score. The prediction equation is In the formula, is the predicted risk score at the future time t + Δt; R t-kτ (i) and are the risk scores and the corrected abnormal change index at the past time t - kτ respectively; τ is the sampling time interval; T is the length of the historical sequence; f LSTM is the LSTM network model function. The prediction confidence interval calculation formula is In the formula, R lower (i) and R upper (i) are respectively the lower bound and the upper bound of the predicted risk score; z α / 2 is the critical value of the normal distribution, which is approximately 1.96 when the confidence level is 95%; σ R is the standard deviation of the predicted risk score, which is estimated by the historical prediction error. According to the prediction result, the risk development trend is judged. If the predicted risk score shows an upward trend and exceeds the warning threshold, the warning upgrade mechanism is triggered. At the same time, a risk warning report is generated, including information such as the abnormal target location, behavior type, risk level, prediction trend, etc., and different response measures are taken according to the warning level: continue to monitor for low risk, notify the on-duty personnel for medium risk, activate the emergency plan for high risk, and trigger automatic intervention measures (such as broadcast reminder or notify the train to slow down) for extremely high risk. The function of this step is to realize the timely warning and early intervention of risks and prevent problems before they occur.

[0180] To better understand and implement the present invention, Example 2 of a specific application scenario of the present invention is provided below: Researchers conducted an implementation test of the rail transit video intelligent analysis method at a certain urban rail transit station. The station has an island platform structure, with an average daily passenger flow of about 85,000 person-times. The platform is severely crowded during the morning and evening rush hours, and there are certain safety risks. Three key monitoring points on the platform were selected for the implementation test, covering the inbound area, the middle platform area, and the outbound area respectively. Three high-definition network cameras (with a resolution of 1920×1080 pixels and a frame rate of 25 frames per second) were used and connected to an edge computing server (configured with an Intel i7-9700 processor, 16GB RAM, and an NVIDIA RTX 2080 graphics card) for real-time analysis and processing.

[0181] First, an adaptive resolution image hierarchical structure is constructed according to step S01. Through the resolution level evaluation function, the optimal number of resolution levels is calculated to be 4 layers, and the resolution parameters of each layer are shown in Table 1:

[0182] Table 1 Adaptive Resolution Hierarchical Parameter Table

[0183] Hierarchical number Resolution (pixels) Downsampling ratio Type of processing task Proportion of computing resource allocation 0 1920×1080 1.0 Fine analysis 35% 1 960×540 0.5 Object tracking 25% 2 480×270 0.25 Motion detection 20% 3 240×135 0.125 Global screening 20%

[0184] The improved frame difference method combined with the optical flow estimation technique is used for abnormal moving target detection. The adaptive threshold adjustment coefficient k is set to 3.2, and the weight coefficients w1 and w2 are set to 0.65 and 0.35 respectively. The comparison of the motion detection results of the system in multiple different scenarios is shown in Table 2:

[0185] Table 2 Comparison Table of Motion Detection Performance in Different Scenarios

[0186] Scene type Detection accuracy (AP) Recall rate (%) False positive rate (%) Average processing time (ms) Low passenger flow scene 0.92 93.5 3.2 16.7 Medium passenger flow scene 0.89 91.2 5.8 18.3 High passenger flow scene 0.85 88.7 7.3 21.5 Scene with changing light 0.83 87.4 8.1 19.2 Rainy day scene 0.80 85.1 9.6 22.7

[0187] Apply the Hungarian algorithm to solve the multi - target tracking problem. The weight coefficients \(w\) of position, size, and appearance in the target matching cost matrix p , \(w\) s , \(w\) a are set to 0.6, 0.25, and 0.15 respectively. The system tracking performance during the test is shown in Table 3:

[0188] Table 3 System multi - target tracking performance table

[0189] Evaluation index Value Description MOTA 82.7% Accuracy of multi-object tracking MOTP 79.3% Precision of multi-object tracking IDS 43 Number of identity switches (in 24 hours) FP 127 Number of false detections (in 24 hours) FN 86 Number of missed detections (in 24 hours) Tracking speed 23.6ms Average processing time per frame

[0190] When constructing the abnormal target change index, the weight coefficients \(w_1\), \(w_2\), \(w_3\) of the size change rate, speed change rate, and direction change rate are set to 0.25, 0.45, and 0.3 respectively, the noise correction term \(\epsilon\) is set to 0.07, and the change index threshold \(T\) change is set to 0.42. The abnormal target detection results at different times of a day during the test are shown in Table 4:

[0191] Table 4 Statistical table of abnormal target detection results at different times

[0192]

[0193]

[0194] For different operating stages of the platform, the system dynamically adjusts the monitoring strategy. The safety distance threshold \(d\) during the vehicle approaching stage safe is set to 1.6 meters, 1.2 meters during the vehicle docking stage, 1.8 meters during the departure stage, and 2.0 meters during the idle stage. The proportion of the area of the concerned area and the detection sensitivity parameters in each stage are shown in Table 5:

[0195] Table 5 Monitoring parameter table for different operating stages

[0196] Operation stage Proportion of area of interest (%) Detection sensitivity Risk factor Type of target to be prioritized Train approaching the station 35 0.85 0.9 People loitering near the platform edge Train docking 50 0.75 0.7 People staying near the door Train leaving the station 30 0.82 0.8 People chasing the train Platform is idle 100 0.65 0.5 Abnormally staying people

[0197] In the risk assessment stage, the system sets the weight coefficients \(w_1\) to \(w_5\) to 0.4, 0.25, 0.2, 0.08, and 0.07 respectively, and the range of the random perturbation term \(\delta\) is from - 0.04 to 0.04. The main abnormal behavior types captured during the test and their risk scores are shown in Table 6:

[0198] Table 6 Statistical table of abnormal behavior risk scores

[0199] Behavior type Occurrence times Average risk score Early warning level Response measures Approaching the platform edge 43 0.68 Orange Broadcast reminder Running inside the platform 126 0.55 Yellow Monitoring attention Abnormal stay 37 0.42 Yellow Monitoring attention Climbing over the guardrail 2 0.88 Red Personnel intervention Falling incident 8 0.73 Orange Personnel assistance Left items 23 0.35 Yellow Monitoring attention Door blocking 15 0.65 Orange Broadcast reminder

[0200] Adopt a hierarchical processing strategy to perform parallel computing on images at different resolution levels. The processing performance of the system at each resolution level is shown in Table 7:

[0201] Table 7 Processing Performance Table for Each Resolution Level

[0202] Resolution level Processing time per frame (ms) Detection accuracy (mAP) Processing content Speedup ratio Original resolution 42.3 0.91 Fine recognition 1.0 1 / 2 resolution 12.7 0.83 Object tracking 3.3 1 / 4 resolution 5.8 0.75 Anomaly detection 7.3 1 / 8 resolution 2.1 0.63 Global screening 20.1

[0203] The system uses an LSTM network model to predict the potential risk trend. The length T of the input historical sequence of the model is set to 20, the sampling time interval τ is 1 second, and the confidence level is 95%. The performance of the system risk prediction accuracy at different prediction durations is shown in Table 8:

[0204] Table 8 Risk Prediction Accuracy Table

[0205]

[0206]

[0207] Compared with traditional rail transit video analysis methods, the method of the present invention shows obvious advantages in the implementation process. Traditional methods mainly rely on motion detection with fixed thresholds and simple rule judgments, and cannot adapt to complex and changeable platform environments. Especially when the crowd is dense, the accuracy rate drops significantly. Traditional methods usually adopt full-resolution processing or fixed downsampling ratios, with low utilization rate of computing resources, and cannot implement differential analysis strategies for different regions. Most traditional systems adopt independent object detection and tracking modules, lack effective trajectory analysis and behavior recognition capabilities, and cannot give early warnings of potential risks.

[0208] The present invention optimizes the computing resource allocation through an adaptive resolution hierarchical structure. The system processing delay is reduced from an average of 47.6 milliseconds of traditional methods to 23.9 milliseconds, with a reduction amplitude of 49.8%, while maintaining a high detection accuracy. The improved frame difference method combined with the optical flow estimation technology significantly improves the robustness of motion detection, and the detection accuracy rate in complex lighting and crowded scenarios is increased by 18.5% compared with traditional methods. The abnormal target change index with multi-parameter weighting realizes more accurate behavior anomaly evaluation, and the false alarm rate is reduced from 12.7% of traditional systems to 7.3%. After introducing the rail scene attention network model, the system has made remarkable progress in identifying complex behavior patterns. Especially when identifying high-risk behaviors such as "approaching the platform edge", the recognition accuracy rate is increased from 72.6% of traditional systems to 89.3%. In addition, the risk prediction model can give early warnings of potential risk events 3 to 15 seconds in advance, providing a valuable time window for safety intervention and effectively preventing the occurrence of safety accidents.

[0209] It should be noted that the detailed explanations of the variables involved in the present invention are shown in Tables 9 and 10 below.

[0210] Table 9 Variable Explanation Table (The First Part)

[0211]

[0212] Table 10 Variable Explanation Table (Second Part)

[0213]

[0214] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention.

Claims

1. An intelligent video analysis method for rail transit, characterized in that, Including: Construct an adaptive resolution image hierarchical structure and dynamically determine the optimal number of resolution levels based on the minimum spanning tree algorithm; Use the improved frame difference method to calculate the difference between adjacent video frames and construct an abnormal target movement matrix; Based on the abnormal target movement matrix, calculate the abnormal target movement displacement matrix, and apply the Hungarian algorithm to solve the multi-target matching problem; according to the abnormal target movement displacement matrix, construct an abnormal target change index; according to the running state of the rail transit vehicle, divide the platform monitoring area into four time periods; introduce an orbital scene attention network model for abnormal target recognition; use an abnormal behavior risk assessment function to quantify the risk level of abnormal behavior; Based on the hierarchical processing strategy, perform parallel calculations on images of different resolution levels; according to the detection results and the abnormal target change index, divide the warning levels according to the preset risk levels.

2. The rail transit video intelligent analysis method according to claim 1, wherein The adaptive resolution image hierarchical structure generates a resolution pyramid from the original high-definition video frame through a recursive binary downsampling method, and the resolution of each level is determined by a resolution level evaluation function.

3. The rail transit video intelligent analysis method according to claim 2, wherein The resolution level evaluation function is used to dynamically determine the optimal number of resolution division levels and the resolution parameters of each layer. The inputs include the original video resolution, computing resource limit parameters, target detection accuracy requirement parameters, real-time requirement parameters, and scene complexity scores, and the outputs are the optimal number of resolution levels and the specific downsampling ratio of each layer.

4. The rail transit video intelligent analysis method according to claim 3, wherein The improved frame difference method combines the optical flow estimation technique to improve the motion detection accuracy and filters out the regions with significant changes through an adaptive threshold.

5. The rail transit video intelligent analysis method according to claim 4, wherein The abnormal target movement matrix refers to a two-dimensional array generated by calculating the pixel difference between adjacent video frames, representing the position information of the moving objects in the picture, and each element value represents the change degree of the corresponding position pixel point; the abnormal target movement displacement matrix refers to recording the position change amount of the detected abnormal target in multiple consecutive video frames, including two components of horizontal displacement and vertical displacement, and is used to analyze the target movement trajectory.

6. The rail transit video intelligent analysis method according to claim 5, characterized in that The abnormal target change index refers to a weighted calculated value that comprehensively considers the target size change, speed change, and direction change. The higher the value, the more abnormal the target behavior. The calculation formula is the target size change rate multiplied by the weight coefficient plus the moving speed change rate multiplied by the weight coefficient plus the direction change rate multiplied by the weight coefficient.

7. The rail transit video intelligent analysis method according to claim 6, characterized in that The orbital scene attention network model integrates the spatio-temporal attention mechanism and the region proposal network to achieve accurate recognition of the behavior patterns of people in the platform environment; the specific structure of the orbital scene attention network model is a spatio-temporal dual attention network based on the improved Transformer architecture. The backbone network uses a deep residual structure to extract image features. The time attention module captures the temporal features of target behavior, and the spatial attention module focuses on the key areas of the platform. The integrated region proposal network improves the small target detection ability. The number of attention heads in the multi-head attention mechanism is dynamically adjusted according to the number of resolution levels output by the resolution level evaluation function, and the dimension of the attention matrix is associated with the threshold of the abnormal target change index.

8. The rail transit video intelligent analysis method according to claim 7, wherein The platform monitoring area is divided into four time periods: the vehicle approaching stage, the vehicle docking stage, the vehicle leaving stage, and the platform idle stage, and an application scenario adaptive analysis strategy is applied.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when running on a computer, are used to execute the method for intelligent analysis of rail transit videos according to any one of claims 1-8.

10. An intelligent video analysis system for rail transit, characterized in that, It includes the computer-readable storage medium according to claim 9. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged inside the system, and a microprocessor for executing the program instructions stored in the computer-readable storage medium is arranged inside the system.

Citation Information

Patent Citations

  • Fire-fighting equipment detection and evaluation system and method based on video processing and deep learning

    CN118095867A

  • Unmanned aerial vehicle routing inspection line adaptive obstacle detection method and system based on monocular camera

    CN119672577A

  • Monitoring information analysis method and system based on artificial intelligence

    CN119693838A

  • AI video low-altitude target identification and real-time tracking method based on deep learning

    CN119723421A

  • Imaging systems and methods for immersive surveillance

    US20120169842A1

Cited By

  • Traffic monitoring video rapid target extraction method for edge device

    CN120953871A

  • Jewel case processing equipment monitoring method and monitoring system

    CN120953915A

  • Information processing method and device for psychological crisis screening

    CN120998534A

  • Information processing methods and devices for psychological crisis screening

    CN120998534B

  • Traffic event report generation method, electronic device, and storage medium

    CN122510842A