A method of monitoring rail intrusion

CN119007098BActive Publication Date: 2026-09-25BEIJING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410965295.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-09-25
Estimated Expiration
2044-07-18

AI Technical Summary

Technical Problem

这些方法的缺点包括:检测精度较低,应用场景单一

Benefits of technology

[0046]由上述本发明的实施例提供的技术方案可以看出,本发明实施例针对现有方式对大规模标注数据的高度依赖,采用弱监督和无监督方式,减少数据标注的复杂度,提高算法模型的适用范围。针对现有方式无法实时处理与资源无法高效利用的困境,本发明通过精心设计的模型轻量化策略,大幅度削减了模型参数规模与计算负担,极大提升了在计算资源有限的移动部署场景下的应用可行性与效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119007098B_ABST
    Figure CN119007098B_ABST
Patent Text Reader

Abstract

The application provides a track intrusion monitoring method. The method comprises the following steps: collecting monitoring video data of a track by a monitoring device installed beside the track, pre-processing the monitoring video data to obtain input video data; performing weakly supervised anomaly feature detection on the input video data to obtain anomaly data; performing anomaly target positioning on the anomaly data to obtain a track intrusion detection result. The embodiment of the application adopts a weakly supervised and unsupervised manner to reduce the complexity of data labeling and improve the application range of the algorithm model, aiming at the high dependence of the existing manner on large-scale labeled data. Through the lightweight strategy of the carefully designed model, the model parameter size and the calculation burden are reduced, and the application feasibility and efficiency in the mobile deployment scene with limited computing resources are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of track monitoring technology, and in particular to a method for monitoring track intrusion. Background Technology

[0002] With the rapid development of China's high-speed rail network, as of early 2024, the total length of railways nationwide exceeded 150,000 kilometers, and the operating length of high-speed railways surpassed 40,000 kilometers, highlighting its global leadership position. This vast transportation network plays a core role in promoting socio-economic development and regional integration, but its expansion also places higher demands on safety performance, making the balance between transportation efficiency and safety an urgent issue.

[0003] High-speed railway systems are complex, with diverse and multi-layered safety considerations, among which environmental factors are particularly critical. Extreme weather and geological disasters pose threats to the natural environment, while foreign object intrusions from the external environment (such as falling objects from railway bridges or landslides) are challenging to monitor due to their suddenness. Therefore, developing a real-time and efficient foreign object intrusion monitoring system, coupled with a rapid response mechanism, is crucial for enhancing the railway's ability to respond to environmental and human-related safety risks.

[0004] my country's railway monitoring system has demonstrated strong effectiveness, leveraging its wide coverage and continuous operation to effectively monitor security threats along the railway lines. However, faced with the exponential growth of monitoring video data, the traditional model relying on manual review suffers from inefficiency and is prone to missed detections, especially given the current surge in the number of railway monitoring nodes.

[0005] The integration of machine vision and deep learning technologies offers an innovative solution. Through automated and efficient detection and in-depth data analysis capabilities, it not only improves monitoring efficiency but also enhances adaptability to complex environments and security levels. Currently, existing vision-based track intrusion detection methods include traditional vision-based methods and deep learning-based methods.

[0006] Traditional vision-based orbital intrusion detection technologies encompass a variety of methods, including but not limited to motion difference or background modeling methods, weighted classification methods combining scene priors and motion patterns, and methods combining feature description and multi-scale analysis. These methods have drawbacks such as low detection accuracy and limited application scenarios.

[0007] The field of foreign object intrusion detection technology based on visual deep learning has shown a diversified development trend in research both domestically and internationally. The core branches of this field can currently be summarized into two main categories: 2D image-based detection technology and 3D vision-based detection technology. For 2D image-based detection, based on the different characteristics of the processed data, it is further subdivided into two directions: static image analysis and video sequence analysis. In static image processing, research focuses on applying image segmentation or object detection algorithms, such as Faster R-CNN or the YOLO series, to accurately depict and locate the morphology and position of foreign objects on the track. In video sequence analysis, research employs techniques such as optical flow, background subtraction, and video anomaly detection, effectively capturing transient foreign object intrusion events occurring within the track area through in-depth analysis of continuous video frames. On the other hand, the 3D vision-based detection field encompasses two main parts: point cloud processing technology and stereo vision technology. In point cloud processing, researchers combine 3D LiDAR and other three-dimensional sensing devices to collect point cloud data, and utilize deep learning architectures specifically designed for three-dimensional point clouds, such as PointNet or PointCNN, to achieve high-precision localization of foreign objects in the three-dimensional space of the track. Stereo vision technology uses binocular or multi-view camera systems to build stereo vision models, calculates parallax to generate depth images, and then combines deep learning technology to accurately identify foreign object intrusion phenomena embedded in stereo images.

[0008] Existing vision-based deep learning-based methods for detecting foreign object intrusions on tracks have several drawbacks: the diversity and morphological variations of foreign object types require models with high generalization capabilities, but comprehensive coverage is difficult in practice; deep learning models consume a lot of resources, increasing hardware costs and potentially limiting real-time processing performance; the lack of sufficient and accurate labeled data limits the training effectiveness of the models; furthermore, balancing detection accuracy with reducing false positives and false negatives, such as misidentification of natural objects or non-threatening objects, remains a challenge in practice. These issues collectively constitute the main obstacles to the implementation of this technology. Summary of the Invention

[0009] Embodiments of the present invention provide a method for monitoring track intrusions, so as to effectively monitor track intrusions.

[0010] To achieve the above objectives, the present invention adopts the following technical solution.

[0011] A method for monitoring track intrusion includes:

[0012] The monitoring video data of the track is collected by the monitoring equipment installed next to the track, and the monitoring video data is preprocessed to obtain the input video data;

[0013] Weakly supervised anomaly detection is performed on the input video data to obtain abnormal data;

[0014] The abnormal data is used to locate the abnormal target and obtain the track intrusion detection results.

[0015] Preferably, the step of collecting track monitoring video data through monitoring equipment installed beside the track, preprocessing the monitoring video data, and obtaining input video data includes:

[0016] Track intrusion videos are collected by surveillance cameras installed beside the track. Half of the collected surveillance video data images are positive samples and half are negative samples. The positive samples are video frames under normal conditions and the negative samples are video frames under abnormal conditions. The selected surveillance video data is preprocessed, including image masking and cropping and far-point information magnification. The preprocessed surveillance video data is used as input video data.

[0017] Preferably, the preprocessing of the selected surveillance video data includes image masking and cropping, and magnification of far-point information, comprising:

[0018] The track spacing is set to 1435mm, and the track area detection range is set to 2500mm from the center of the track to both sides. The track lines are obtained through the track detection algorithm, and the center line of the track lines is drawn. The track lines on both sides are translated to the outside by a ratio of 2500 / (1435 / 2) to form boundary lines. A mask is drawn based on the boundary lines. Other local areas are cropped according to the region of interest (ROI). The local image of the far point of the track is cropped separately and merged with the complete image that has not been cropped into a set of complementary image samples to obtain the final input image video.

[0019] Preferably, the step of performing weakly supervised anomaly detection on the input video data to obtain anomaly data includes:

[0020] The X3D model is pre-trained using Kinetics-400. The trained X3D model is used to extract features from the input video data to obtain feature vectors. After the feature vectors are enhanced with Non-Local and multi-scale temporal attention mechanisms, optimized feature vectors are obtained. The optimized feature vectors are scored by a classifier. The K scores with the highest absolute values ​​are selected from the scores of normal and abnormal video frames respectively. Based on the frame index corresponding to the K highest scores, the corresponding K feature vectors are extracted.

[0021] Based on the selected K scores and their corresponding feature vectors, the magnitude fraction loss function, magnitude feature loss function, and time smoothing loss function are calculated. The calculation method for the magnitude fraction loss function is as follows:

[0022]

[0023] Where x represents a normal or abnormal video frame, y represents the label of the video frame, and f represents the model operation;

[0024] Calculate the L2 norm of the scores of K video frames from the normal and abnormal classes, and then sum them:

[0025]

[0026] Where N represents the number of video frames, Ω k (X) is a set containing K video frames, ||f(x) i )|| 2 Let L2 norm represent the score of the i-th video frame;

[0027] The difference between the sum of the feature norms of the K normal video frames and the sum of the feature norms of the K abnormal video frames is shown in the following formula:

[0028] d(X + X - ) = g k (X + )-g k (X - )

[0029] Among them, g k (X + ) represents the L2 norm sum of the normal video frame fractions, g k (X - ) represents the sum of the L2 norms of the abnormal video frame scores. This function indicates the degree of distinction between normal and abnormal video frames.

[0030] The amplitude characteristic loss function is calculated as follows:

[0031] l s =max(0,md(X) i X j ))

[0032] m is the threshold defined to ensure that the loss function is a positive number; d(X) i X j That is, the function mentioned above that represents the degree of distinction between normal and abnormal video frames;

[0033] By adding a temporal smoothing loss function, the scores between adjacent normal frames and abnormal frames in a video become closer after multiple training sessions. The calculation method of the temporal smoothing loss function is as follows.

[0034]

[0035] Where λ is a pre-set hyperparameter, i represents the index order of the current frame within the video, and f(x) i f(x) represents the score of the i-th video frame. i+1 () represents the score of the (i+1)th video frame, and N represents the number of video frames;

[0036] The specific formula for the total loss function is shown below:

[0037]

[0038] Video frames whose total loss function value exceeds a set threshold are judged as abnormal video frames, while video frames whose total loss function value does not exceed the set threshold are judged as normal video frames.

[0039] Preferably, the step of locating abnormal targets in the abnormal data and obtaining track intrusion detection results includes:

[0040] A multi-scale orbit anomaly localization model architecture based on a parameter sharing mechanism is constructed. This multi-scale orbit anomaly localization model architecture divides the encoder and decoder structure into three progressive layers. In the first two layers, two downsampling / upsampling units are configured in each layer. In the third layer, a conventional upsampling / downsampling module and a layer of upsampling / downsampling modules that are dynamically enabled as needed are integrated.

[0041] The video frames identified as abnormal are input into the multi-scale orbital anomaly localization model architecture. The encoder and decoder structure in the multi-scale orbital anomaly localization model architecture contains five consecutive downsampling or upsampling units. Each sampling module integrates a spatial location encoding mechanism. It performs downsampling operations in the encoding stage and upsampling operations in the decoding stage to capture spatial information from local to global. Point-based convolutional or deconvolutional layers enhance the expressiveness of feature mapping. Batch normalization steps are applied after all convolution operations. Activation functions are used in the core calculation links of each module. The number of sampling layers is configured according to the input image of different sizes.

[0042] An effective channel attention (ECA) mechanism is set in the multi-scale orbital anomaly localization model architecture. The multi-scale VAE architecture contains latent variables at multiple scales, each with a corresponding mean and variance. The total KL divergence can be defined as follows:

[0043]

[0044] Here, z l Let q(z) represent the latent variable at the l-th scale level, where L is the total number of scale levels. l |x, z<l) is a given observation data x and a low-level latent variable z <lAt that time, the approximate posterior distribution of the high-level latent variable z, while p(z) l |z <l ) is the corresponding prior distribution;

[0045] The multi-scale orbital anomaly localization model architecture receives anomalous video frames selected by anomaly detection algorithms as input, uses a deep neural network architecture to predict the background image that constitutes the video background, and calculates the difference map between the anomalous video frames and the background image by performing pixel-level comparison, i.e., the foreground mask. The foreground mask highlights anomalous activity areas that are expected to change significantly relative to the background. Activation at different positions on the foreground mask corresponds to multiple independent anomalous targets in the video.

[0046] As can be seen from the technical solutions provided by the embodiments of the present invention above, the embodiments of the present invention address the high dependence of existing methods on large-scale labeled data by adopting weakly supervised and unsupervised methods to reduce the complexity of data labeling and improve the applicability of the algorithm model. Addressing the limitations of existing methods in real-time processing and efficient resource utilization, the present invention, through a carefully designed model lightweighting strategy, significantly reduces the scale of model parameters and computational burden, greatly improving the feasibility and efficiency of application in mobile deployment scenarios with limited computing resources.

[0047] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 A flowchart illustrating a track intrusion monitoring method provided in an embodiment of the present invention;

[0050] Figure 2 This invention provides a sample of track intrusion detection data.

[0051] Figure 3 An example of input data composition provided in an embodiment of the present invention;

[0052] Figure 4 This invention provides an example of image masking and cropping.

[0053] Figure 5This invention provides a method for far-point cropping and global preservation of an image (the far-point image is the original image in the upper left corner cropped to a certain size, and the global image is the original image scaled down to the same size as the far-point image).

[0054] Figure 6 This invention provides a weakly supervised orbital video anomaly detection model framework.

[0055] Figure 7 A Non-Local module structure diagram provided in an embodiment of the present invention;

[0056] Figure 8 This invention provides a multi-scale temporal domain attention module.

[0057] Figure 9 This invention provides an architecture diagram of a multi-scale foreground segmentation algorithm based on parameter sharing.

[0058] Figure 10 An improved downsampling block and upsampling block are provided for embodiments of the present invention;

[0059] Figure 11 This invention provides a demonstration of track anomaly detection results.

[0060] Figure 12 This is a sample of foreground segmentation results provided in an embodiment of the present invention. Detailed Implementation

[0061] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0062] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.

[0063] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.

[0064] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.

[0065] This invention provides a method for monitoring track intrusion, which is based on weakly supervised video anomaly detection and unsupervised foreground segmentation techniques, aiming to overcome the limitations of existing methods. Addressing the trade-off between accuracy and breadth in anomaly detection, this invention combines video anomaly detection and foreground segmentation to solve the track intrusion detection problem, improving the ability to identify unseen or atypical intrusion behaviors and enhancing system reliability and practicality.

[0066] The processing flowchart of a track intrusion monitoring method provided in this embodiment of the invention is as follows: Figure 1 As shown, the processing steps include the following:

[0067] Step S10: Data collection.

[0068] Track intrusion videos are captured by surveillance cameras installed alongside the track. The captured video data images are as follows: Figure 2 As shown, the video resolution is 2560*1440 and the frame rate is 25 frames per second.

[0069] Step S20: Data input settings.

[0070] First, the input settings for the model are as follows: In video anomaly detection based on weak supervision, half of the positive samples (video frames under normal conditions) and half of the negative samples (video frames under abnormal conditions) are selected from the collected surveillance video data images, and the selected surveillance video data is used as the input video data.

[0071] The above design strategy aims to ensure that the model is fully exposed to anomalous cases during the learning process to accurately identify various abnormal phenomena, while ensuring sufficient representativeness of normal behavior to prevent the model from misidentifying normal situations as anomalous. An example of input video data is shown below. Figure 3 As shown.

[0072] Step S30: Preprocess the input video data.

[0073] The video preprocessing in this embodiment includes two processes: image masking and cropping, and far-point information magnification. The specific principles and processes of the two preprocessing processes are described in detail below.

[0074] The first step is the image masking and cropping process. In this embodiment, the area outside the track boundaries is masked to reduce the impact of non-detection regions on the results during model detection. Specifically, the track spacing is known to be 1435mm, and the typical track region detection range is 2500mm from the track center to both sides. First, the track lines are obtained using a track detection algorithm. Then, the center line of the track lines is drawn. Using a ratio of 2500 / (1435 / 2), the two side track lines are shifted to the outside to form boundary lines. A mask is then drawn based on these boundary lines. Finally, based on the ROI (region of interest), other local areas are cropped to obtain the final input image / video. It is worth noting that by implementing this operation, the anomaly features, which were originally tightly coupled with location, are transformed into location-independent anomaly representations. This allows the model to adapt to anomaly detection tasks in various scenarios, even when there are no obvious location clues.

[0075] Figure 4 This is a schematic diagram showing the data after two operations: masking and cropping. The original image is 2560*1440. After masking and cropping, the image becomes 1790*660. A large number of useless areas are removed, which helps to extract more useful information from feature extraction. The smaller input image can also improve the running speed of the model to some extent.

[0076] Secondly, there is the process of amplifying the far-point information. This invention proposes a strategy that involves cropping a local image of the far-point region of the orbit separately and merging it with the complete image that has not been cropped into a set of complementary image samples. The former highlights the detailed features of distant targets in the current frame, while the latter reflects the overall features of the scene.

[0077] Its deliberate preservation of far-point information is essentially equivalent to constructing an additional, abstract scene perspective. Despite the pruning operation, logically speaking, this preprocessing step does not weaken the model's detection accuracy. On the contrary, considering and fusing far-point details with global scene information for anomaly detection significantly improves the efficiency of detecting anomalous events.

[0078] Figure 5 This is a schematic diagram of the image after far-point cropping and global preservation. The original image is a 1770*660 image, which is transformed into two 512*189 images after processing.

[0079] Step S40: Weakly supervised video anomaly feature detection.

[0080] The weakly supervised video anomaly detection technology employs a phased processing strategy for videos of foreign object intrusions into orbits. First, the system performs binary classification on the pre-processed video data, identifying it as "normal" or "abnormal." When the video scene is determined to meet the expected "normal" criteria, the detection process terminates without further processing, thus optimizing computational efficiency. Conversely, if the video is marked as "abnormal," the next stage of analysis—foreground segmentation and anomaly target localization—is triggered.

[0081] Figure 6 This is a schematic diagram of the overall framework of a weakly supervised learning learning detection model for foreign object intrusion on tracks based on video surveillance, provided by an embodiment of the present invention. It includes video input settings, a video preprocessing module, a feature extractor pre-trained using a large video dataset, an attention enhancement module (including a non-local attention module for establishing global dependencies and a multi-scale temporal domain attention module for establishing local temporal dependencies), a linear classifier module, a TOP-K selection module, and loss functions (including amplitude fraction loss function, amplitude feature loss function, and temporal smoothing term loss function).

[0082] This invention uses the Xception3D (X3D) model for video feature extraction. X3D is a deep neural network model for video understanding, proposed by Facebook AI. It is designed to improve the performance of tasks such as video classification, action recognition, and video generation. The core idea of ​​the X3D model is to design an efficient three-dimensional convolutional structure to handle the spatiotemporal characteristics of video data.

[0083] This invention utilizes the Kinetics-400 (K400) dataset, a large-scale action recognition benchmark dataset, to pre-train the X3D model. This model is then transferred and applied to a weakly supervised learning learning model for detecting foreign object intrusion on tracks based on video surveillance. The K400 dataset covers 400 diverse action categories, such as sports activities, daily life behaviors, and performing arts. Each category includes approximately 400 independent video clips, contributing a total of about 240,000 unique video samples with a cumulative duration exceeding 320 hours. Pre-training the X3D model on this dataset aims to extract more universal video feature representations. This helps improve the model's transfer learning efficiency in various video understanding tasks, significantly reducing the time and computational resources required for training from scratch for a specific task, and enhancing the model's adaptability to different visual scenes and target tasks.

[0084] Non-Local attention is a neural network architecture used to model long-range dependencies. It was first proposed by Facebook AI, and its structure is as follows: Figure 7As shown, this network architecture aims to address the limitations of traditional convolutional neural networks in handling long-range dependencies. Its core idea is to introduce non-local operations, enabling the network to establish connections between any two locations, thereby capturing global contextual information. This operation is based on a key observation: rich correlations exist between pixels in an image.

[0085] While non-local attention mechanisms are highly effective at strengthening global dependencies between features, temporal local correlations remain equally crucial in anomaly detection tasks. Therefore, this invention introduces a pyramid structure scheme, employing one-dimensional dilated convolution to deeply mine the multi-scale expressive characteristics of video segments along the temporal dimension. Notably, this dilated convolution mechanism aims to capture local temporal correlations across different receptive fields; its intuitive representation can be found in [reference needed]. Figure 8 The diagram shows a multi-scale temporal attention module.

[0086] The feature extractor extracts features from the preprocessed video data to obtain feature vectors. These feature vectors are then enhanced with non-local and multi-scale temporal attention mechanisms to obtain optimized feature vectors. These optimized feature vectors are then passed through a classifier to generate scores for the corresponding video frames. Further, the K highest absolute scores are selected from the scores of both normal and abnormal video frames, and the amplitude score loss is calculated based on these scores. Subsequently, based on the frame indices corresponding to these K highest scores, the corresponding feature vectors are extracted to calculate the amplitude feature loss for the K most significant features.

[0087] Based on Top-K selection, this embodiment of the invention performs loss calculations on the selected K scores and their corresponding features. The prototype of the score loss function is the cross-entropy loss function, but the input is only the scores of K normal and abnormal video frames, as shown below.

[0088]

[0089] Where x represents a normal or abnormal video frame, y represents the label of the video frame, and represents the abstract representation of the weakly supervised video anomaly detection model operation. The weakly supervised video anomaly detection model is as follows: Figure 6 As shown.

[0090] To enhance the discriminative power of the loss function in distinguishing between normal and abnormal videos, the TOP-K selection method is used to filter the features with the greatest differences. Each feature is typically a vector, and the L2 norm can be used to convert it into a positive number, which represents the score corresponding to that feature. The formula below calculates the L2 norm of the scores from K video frames and then sums them.

[0091]

[0092] Next, the sum of the feature norms of the K normal video frames in the normal video is subtracted from the sum of the feature norms of the K abnormal video frames, as shown in the formula below.

[0093] d(X + X - ) = g k (X + )-g k (X - )

[0094] Finally, the feature magnitude loss function is defined as follows, where m is the threshold defined to ensure that the loss function is a positive number.

[0095] l s =max(0,md(X) i X j ))

[0096] Furthermore, normal and abnormal frames in a video often appear consecutively, meaning that several adjacent frames of a normal or abnormal frame are also normal or abnormal frames. Therefore, this embodiment of the invention adds a temporal smoothness constraint loss function (SCL) so that after multiple training iterations, the scores between adjacent normal and abnormal frames in a video are relatively close.

[0097]

[0098] Where λ is a pre-set hyperparameter that limits the value of the temporal smoothing loss function to a reasonable range, and i represents the index order of the current frame in the video.

[0099] Ultimately, the loss function consists of three parts: amplitude fraction loss function, amplitude feature loss function, and time smoothing loss function, as shown in the following formula.

[0100]

[0101] Video frames whose total loss function value exceeds a set threshold are judged as abnormal video frames, while video frames whose total loss function value does not exceed the set threshold are judged as normal video frames.

[0102] Step S50: Foreground segmentation and abnormal target localization.

[0103] Following weakly supervised video anomaly detection, foreground segmentation and anomalous target localization are performed, operating based on video frame image data identified as "anomalies" in the previous steps. The core input to this process is the sequence of video images initially screened as anomalous. The aim is to produce two key outputs through in-depth analysis of these images: first, information about the background components of the corresponding images; and second, a mask image used to identify prominent anomalous objects within the video frames. This mask image is crucial for accurately locating the geometric boundaries of anomalous targets. This step aims to precisely locate the intrusion position within the video frame, extract key anomalous region information, and provide accurate information for timely intervention and response. This hierarchical processing mechanism ensures rapid response to potential threats while avoiding unnecessary resource waste.

[0104] The activation conditions for this subsequent processing step are clearly defined: the foreground segmentation and abnormal target localization process is only activated when the weakly supervised video anomaly localization step determines that the video content violates the normal pattern and confirms its abnormal attributes, thus achieving precise spatial definition of the abnormal entity's location. Conversely, if the initial analysis concludes that the video content conforms to the norm (i.e., the "normal" category), this step is skipped directly, avoiding unnecessary resource consumption and ensuring the efficient and accurate operation of the entire detection system. This strategy not only reflects the rational allocation of computing resources but also profoundly demonstrates the hierarchical processing logic from macroscopic anomaly identification to microscopic target precise localization.

[0105] Figure 9 This paper presents the architecture of a multi-scale orbital anomaly localization model based on a parameter-sharing mechanism. The model's unique feature lies in its division of the encoder and decoder structure into three progressive layers. The first two layers each contain two downsampling / upsampling units, which extract and generate low-dimensional variance and mean feature vectors to construct the latent representation space. The third layer integrates a conventional upsampling / downsampling module and a dynamically activated upsampling / downsampling module to generate higher-dimensional latent variables. These latent variables at different levels capture image features at different scales, from fine-grained to coarse-grained.

[0106] By employing this multi-layered variational autoencoder architecture, the model effectively fuses latent image feature information across various scales. Notably, although the model essentially exhibits characteristics similar to three independent VAEs (Variational Autoencoders) operating in parallel, thanks to the parameter-sharing design, its overall parameter count is close to that of a single top-level complexity VAE, thus achieving efficient parameter reuse between low- and high-level VAE structures.

[0107] Both the encoder and decoder structures contain five consecutive downsampling or upsampling units, and the design of this series of units is as follows: Figure 10 As shown, each process module integrates multiple key components. Specifically, each sampling module integrates a spatial location encoding mechanism to capture spatial information from local to global; then, depthwise separable convolution (which is a downsampling operation in the encoding stage and a corresponding upsampling operation in the decoding stage) is used to achieve efficient feature extraction and reconstruction; the subsequent pointwise convolutional or deconvolutional layers further enhance the expressiveness of the feature mapping; at the same time, all convolutional operations are supplemented with batch normalization steps to stabilize the training process and optimize model convergence; finally, the core computational links of each module use activation functions to improve the effect of nonlinear transformation and the learning ability of the network. In addition, in order to enhance the model's distinguishing attention to the features of each convolutional channel and strengthen its robustness to illumination noise, this embodiment of the invention further introduces the ECA (Efficient Channel Attention) mechanism into the network structure.

[0108] A multi-scale VAE architecture contains latent variables at multiple scales (or levels), each with its own mean and variance, both of which affect the KL divergence loss. In the objective function of a variational autoencoder (VAE), the KL divergence is a metric used to measure the difference between the approximate posterior distribution learned by the model and the prior distribution. For a multi-scale hierarchical VAE, the total KL divergence can be defined as follows:

[0109]

[0110] Here, z l Let q(z) represent the latent variable at the l-th scale level, where L is the total number of scale levels. l |x, z<l) is a given observation data x and a low-level latent variable z <l At that time, the approximate posterior distribution of the high-level latent variable z, while p(z) l |z <l ) is the corresponding prior distribution.

[0111] The multi-scale orbital anomaly localization model architecture receives anomalous video frames selected by anomaly detection algorithms as input. It then employs a deep neural network architecture to predict static image features constituting the video background. This prediction process involves reconstructing background image information to capture normal visual patterns within the scene. By comparing the input anomalous video frames with the background image generated by the model at the pixel level, a difference map, or foreground mask, can be calculated. This mask highlights areas that exhibit significant expected changes relative to the background, often associated with potential anomalous activity. Notably, activations at different locations on the mask correspond to multiple independent anomalous targets in the video, enabling parallel identification and localization of various anomaly types.

[0112] To enhance the model's adaptability and generalization ability to input images of different sizes, this embodiment of the invention adopts a flexible strategy in model design: configuring a corresponding number of sampling layers based on the input image size. Specifically, for smaller images (e.g., less than or equal to 500x500 pixels), considering the efficient use of computing resources and the prevention of overfitting, a 5-layer sampling block is selected, and the number of convolutional channels is configured as [3, 64, 160, 160, 32, 16]. This channel allocation ensures that the model fully captures image features with a limited number of parameters, thereby achieving efficient and accurate feature extraction.

[0113] For larger images, which contain more spatial information and potential details, an additional sampling block was added, bringing the total to six layers. The corresponding number of convolutional channels was set to [3, 64, 160, 160, 160, 32, 16]. This design allows the model to maintain low computational complexity when processing large images, while also enabling deeper mining and integration of image features, further enhancing the model's performance and robustness in complex scenes.

[0114] In summary, by flexibly utilizing a multi-layer sampling block structure and combining it with finely adjusted convolution channel configurations for images of different sizes, the model can effectively enhance its adaptability and generalization performance to input images of various sizes while ensuring learning efficiency, thus demonstrating better recognition and processing effects in practical applications.

[0115] In summary, compared to current state-of-the-art methods, this invention innovatively integrates weakly supervised and unsupervised learning strategies to address the intelligent identification and precise localization of abnormal targets in orbital intrusion situations. This significantly simplifies the dataset construction and annotation process, effectively alleviating the complexity of annotation. Most importantly, this invention demonstrates superior generalization ability, autonomously discovering potential threats in unlabeled abnormal instances, greatly expanding the detection scope and applicability.

[0116] Furthermore, this invention, through a meticulously designed model lightweighting strategy, significantly reduces the size of model parameters and the computational burden, greatly improving the feasibility and efficiency of application in mobile deployment scenarios with limited computing resources. This optimization not only promotes efficient resource utilization but also paves the way for the miniaturization and portable deployment of real-time monitoring systems.

[0117] In terms of performance indicators, this invention achieves a dual improvement in detection accuracy and false negative rate, ensuring high precision and reliability in monitoring track intrusion events, and contributing a more accurate, efficient and adaptable solution to the field of rail transit safety monitoring.

[0118] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0119] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0120] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0121] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for monitoring track intrusion, characterized in that, include: The monitoring video data of the track is collected by the monitoring equipment installed next to the track, and the monitoring video data is preprocessed to obtain the input video data; Weakly supervised anomaly detection is performed on the input video data to obtain abnormal data; The abnormal data is used to locate the abnormal target and obtain the track intrusion detection results; The process of performing weakly supervised anomaly detection on the input video data to obtain anomaly data includes: The X3D model was pre-trained using Kinetics-400. The trained X3D model was then used to extract features from the input video data, resulting in feature vectors. These feature vectors were then enhanced with Non-Local and multi-scale temporal attention mechanisms to obtain optimized feature vectors. A classifier was used to score these optimized feature vectors, and the highest absolute values ​​were selected from the scores of normal and abnormal video frames. Each score, based on Extract the corresponding frame index of the highest score. 1 eigenvector; According to the selected The magnitude fraction loss function, magnitude feature loss function, and time smoothing loss function are calculated for each score and its corresponding feature vector. The calculation method for the magnitude fraction loss function is as follows: in This represents a video frame that is either normal or abnormal. The tag indicating the video frame is located in the video. Represents model operations; For normal and abnormal classes Calculate the L2 norm of each video frame score and sum them: in, Indicates the number of video frames. It includes A collection of video frames. Indicates the first The L2 norm of each video frame score; Normal class The feature norm of each video frame and anomalies The difference between the feature norms of each video frame is calculated using the formula below: in, This represents the sum of the L2 norms of the normal video frame fractions. This function represents the sum of the L2 norms of the abnormal video frame scores, indicating the degree of distinction between normal and abnormal video frames. The amplitude characteristic loss function is calculated as follows: To achieve the defined threshold, the loss function must be a positive number; That is, the function mentioned above that represents the degree of distinction between normal and abnormal video frames; By adding a temporal smoothing loss function, the scores between adjacent video frames in the same video remain relatively smooth after multiple training sessions. The calculation method of the temporal smoothing loss function is as follows; in, These are pre-set hyperparameters. This indicates the index order of the current frame within the video. Indicates the first Score per video frame Indicates the first Score per video frame Indicates the number of video frames; The specific formula for the total loss function is shown below: Video frames whose total loss function value exceeds a set threshold are judged as abnormal video frames, and video frames whose total loss function value does not exceed the set threshold are judged as normal video frames. The process of locating abnormal targets in the abnormal data and obtaining track intrusion detection results includes: A multi-scale orbit anomaly localization model architecture based on a parameter sharing mechanism is constructed. This multi-scale orbit anomaly localization model architecture divides the encoder and decoder structure into three progressive layers. In the first two layers, two downsampling / upsampling units are configured in each layer. In the third layer, a conventional upsampling / downsampling module and a layer of upsampling / downsampling modules that are dynamically enabled as needed are integrated. The video frames identified as abnormal are input into the multi-scale orbital anomaly localization model architecture. The encoder and decoder structure in the multi-scale orbital anomaly localization model architecture contains five consecutive downsampling or upsampling units. Each sampling module integrates a spatial location encoding mechanism. It performs downsampling operations in the encoding stage and upsampling operations in the decoding stage to capture spatial information from local to global. Point-based convolutional or deconvolutional layers enhance the expressiveness of feature mapping. All convolutional operations are supplemented with batch normalization steps. The core calculation links of each module use activation functions, and the corresponding number of sampling layers are configured according to the input images of different sizes. An effective channel attention (ECA) mechanism is set in the multi-scale orbital anomaly localization model architecture. The multi-scale VAE architecture contains latent variables at multiple scales, each with a corresponding mean and variance. The total KL divergence can be defined as follows: here, Indicates the first Latent variables at each scale level, It is the total number of scale levels. Given observation data and low-level latent variables At that time, high-level latent variables The approximate posterior distribution, while It is the corresponding prior distribution; The multi-scale orbital anomaly localization model architecture receives anomalous video frames selected by anomaly detection algorithms as input, uses a deep neural network architecture to predict the video background image, and calculates the difference map between the anomalous video frames and the background image by performing pixel-level comparison between the anomalous video frames and the background image, i.e., the foreground mask. The foreground mask highlights anomalous activity areas that are expected to change significantly relative to the background. The activation regions at different positions on the foreground mask correspond to multiple independent anomalous targets in the video.

2. The method according to claim 1, characterized in that, The process of collecting track monitoring video data using monitoring equipment installed beside the track, preprocessing the monitoring video data, and obtaining input video data includes: Track intrusion videos are collected by surveillance cameras installed beside the track. Half of the collected surveillance video data images are positive samples and half are negative samples. The positive samples are video frames under normal conditions and the negative samples are video frames under abnormal conditions. The selected surveillance video data is preprocessed, including image masking and cropping and far-point information magnification. The preprocessed surveillance video data is used as input video data.

3. The method according to claim 2, characterized in that, The selected surveillance video data is preprocessed, including image masking and cropping, and far-point information magnification, comprising: The track spacing is set to 1435mm, and the track area detection range is set to 2500mm from the track center to both sides. The track line is obtained using a track detection algorithm, and the center line of the track line is drawn. The proportions are used to shift the track lines on both sides to the outside to form boundary lines. A mask is drawn based on the boundary lines. Other local areas are cropped according to the region of interest (ROI). The local image of the far point of the track is cropped separately and merged with the complete image that has not been cropped into a set of complementary image samples to obtain the final input image video.

Citation Information

Patent Citations

  • Weak supervision video violence detection method based on time domain enhancement and comparative learning

    CN118015507A

  • Hyperspectral abnormal target detection method based on multi-scale analysis and variational auto-encoder

    CN118279747A