Rail transit environment foreign matter intrusion detection method and system based on visual model
Through the visual model fusion of multi-level similarity weighted fusion of image and text features, and combined with the dynamic threshold model, the blind spots and false alarm problems of foreign object detection in rail transit environment are solved, and efficient and accurate foreign object recognition is achieved.
Patent Information
- Application Number
- CN202510750141.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The existing rail transit environmental foreign object detection methods have problems such as large detection blind spots, poor environmental adaptability and high false alarm rates, especially in complex and changeable railway scenarios, which are difficult to effectively identify foreign object types.
Using a detection method based on visual model, the coarse-grained global features and fine-grained local features are extracted through the image encoder, combined with the text features generated by the text encoder, multi-level similarity is calculated and weighted fusion is weighted. Dynamic threshold judgment is used for the hybrid Gaussian model to realize foreign object intrusion detection.
It improves the comprehensiveness and accuracy of foreign object detection in complex railway scenarios, reduces the risk of misjudgment, adapts to light changes and weather interference, reduces computing resource consumption, and improves the robustness and real-timeness of the detection system.
Smart Images

Figure CN120259986A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of rail transit, and particularly to a method and system for detecting foreign object intrusion in the rail transit environment based on a vision model. Background Art
[0002] In recent years, the high-speed railway network in China has developed rapidly, and the operating mileage has continued to grow, putting forward higher requirements for the safety monitoring of the line environment. Foreign object intrusion along the railway (such as gravel, branches, lost tools, etc.) may trigger major safety accidents, and there is an urgent need for efficient and reliable detection means. Traditional detection methods mainly rely on fixedly deployed sensors or manual inspections, and there are problems such as large detection blind spots and poor environmental adaptability. With the development of deep learning technology, vision-based detection methods have gradually become the mainstream, but still face many challenges: supervised learning methods rely on a large amount of labeled data and are difficult to cover complex and variable foreign object types; unsupervised methods are based on the reconstruction error of autoencoders and have a high false alarm rate in dynamic lighting and complex background environments; existing large vision models lack optimization for railway scenarios and are difficult to balance the association between global environmental information and local foreign object features. In addition, existing technologies mostly use fixed thresholds to judge abnormalities and cannot meet the dynamic change requirements of train operation scenarios. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides a method and system for detecting foreign object intrusion in the rail transit environment based on a vision model.
[0004] In the first aspect, the present invention provides a method for detecting foreign object intrusion in the rail transit environment based on a vision model, including the following steps: S1. Preprocess the original images collected by the on-vehicle cameras of high-speed railways, identify and extract the suspicious regions containing potential foreign objects, and obtain the original complete images and at least one sub-image of the suspicious region; S2. Input the original complete images and the sub-images of the suspicious region into an image encoder to extract coarse-grained global feature vectors and fine-grained local feature vectors respectively; S3. Generate corresponding text prompt information for the original complete images and the sub-images of the suspicious region, and extract text feature vectors through a text encoder; S4. Calculate the multi-level similarity between the coarse-grained global feature vectors and the text feature vectors of the complete images, and the multi-level similarity between the fine-grained local feature vectors and the text feature vectors of the suspicious regions; S5. Weight and fuse the two types of similarities according to a preset weight to obtain a comprehensive anomaly score; S6. Based on a dynamic threshold model, perform distribution modeling on the comprehensive anomaly score to determine whether there are intruding foreign objects in the images.
[0005] Optionally, the multi-level similarity calculation in S4 includes: Extract the first-level feature vector, second-level feature vector, and third-level feature vector of the complete image, as well as the corresponding-level feature vectors of the suspicious region sub-graph, respectively, at the 6th, 9th, and 12th layers of the image encoder; Calculate the similarity Sim between the feature vectors of each level of the complete image and the text feature vector of the complete image Di , and the formula is: ; Among them, represents the image feature vector output by the kth layer of the image encoder, where k ∈ {6, 9, 12} represents the network depth level, represents the text feature vector of the complete image at the corresponding level, which is generated from the text prompt information containing the description of the railway scene; Calculate the similarity Sim between the feature vectors of each level of the suspicious region sub-graph and the text feature vector of the suspicious region Di j , and the formula is: ; Among them, represents the kth layer image feature vector of the suspicious region sub-graph, represents the text feature vector of the suspicious region at the corresponding level, which is generated from the text prompt information containing the description of the foreign object feature.
[0006] Optionally, the weighted fusion in S5 satisfies: ; Among them, An_score represents the comprehensive anomaly score, represents the complete image anomaly score, defined as ; represents the suspicious region sub-graph anomaly score, defined as , and α and β are preset weight coefficients, satisfying α + β = 1.
[0007] Optionally, the dynamic threshold model in S6 is a mixture Gaussian model, and its probability density function is: ; Among them, x is the input comprehensive anomaly score vector, μ is the mean vector, Σ is the covariance matrix, n is the dimension of x, T represents the transpose symbol, and π is the pi.
[0008] Optionally, the mixture Gaussian model optimizes the parameters through the expectation maximization algorithm, and iteratively calculates the posterior probability γ(z i ) belonging to the jth Gaussian distribution for the sample x ij as: ; where K is the number of Gaussian distributions, and π j is the mixing coefficient of the j-th distribution, and satisfies , represents the probability density value of the sample point x i under the j-th Gaussian distribution, represents the probability density value of the sample point x i under the k-th Gaussian distribution.
[0009] Optionally, the model parameter update satisfies: ; ; ; where N is the total number of training samples, μ j is the updated mean vector of the j-th Gaussian distribution, π j is the updated mixing coefficient, Σ j is the updated covariance matrix, and T represents the transpose symbol.
[0010] Optionally, the image encoder adopts a hierarchical feature extraction structure, and the image feature vector I fk output by its k-th layer satisfies: ; where represents the k-th layer network of the encoder, is the image feature vector output by the previous layer.
[0011] Optionally, the generation of the suspicious region subgraph satisfies: ; where S is the image segmentation model, and D i represents the input original image, is the output j-th suspicious region subgraph.
[0012] In a second aspect, the present invention also provides a rail transit environment foreign object intrusion detection system based on a vision model, including: A preprocessing module for preprocessing the original image collected by a high-speed railway vehicle-mounted camera, identifying and extracting suspicious regions containing potential foreign objects, and obtaining the original complete image and at least one suspicious region subgraph; A first extraction module for inputting the original complete image and the suspicious region subgraph into an image encoder to respectively extract a coarse-grained global feature vector and a fine-grained local feature vector; A second extraction module is used to generate corresponding text prompt information for the original complete image and the suspicious area sub-image respectively, and extract text feature vectors through a text encoder; A calculation module, used for calculating the multi-level similarity between the coarse-grained global feature vector and the complete image text feature vector, and the multi-level similarity between the fine-grained local feature vector and the suspicious area text feature vector; The fusion module is used to perform weighted fusion of the two types of similarities according to the preset weights to obtain a comprehensive anomaly score; The determination module is used to perform distribution modeling on the comprehensive abnormality score based on a dynamic threshold model to determine whether there is an intrusive foreign body in the image.
[0013] The present invention has the following technical effects: The present invention significantly improves the comprehensiveness and accuracy of foreign body detection in complex railway scenes by fusing the multimodal association of coarse-grained global feature vectors and fine-grained local feature vectors, combined with the semantic information generated by text prompt information. Based on the multi-level feature vectors extracted by the image encoder at different network depths, the similarity matching is performed with the corresponding text feature vectors respectively, so as to enhance the synergy of global environmental anomaly perception and local foreign body detail capture, and avoid the omission of small foreign bodies or large-scale anomalies in a single feature layer. The complete image anomaly score and the suspicious area sub-graph anomaly score are weighted and fused with preset weights to optimize the sensitivity of foreign body detection at different scales, taking into account the overall environmental stability of the track and the coordinated discrimination of mutation characteristics in local areas. The comprehensive anomaly score is dynamically distributed modeled by a mixed Gaussian model, and the judgment threshold is adaptively adjusted according to real-time data, which effectively adapts to complex environmental fluctuations such as lighting changes and weather interference, and reduces the risk of misjudgment caused by traditional fixed thresholds. The pre-screening mechanism of the suspicious area sub-graph is combined with the image segmentation model to reduce the interference feature extraction of non-critical areas and improve the pertinence of local foreign body analysis. The scene description feature vector and foreign object description feature vector generated by the text encoder provide semantic guidance for the image content and strengthen the differentiated representation capabilities of normal scenes and abnormal areas. In an unsupervised framework, it relies only on normal sample training and mines potential abnormal patterns through cross-modal feature comparison to solve the actual detection problem of scarce foreign object samples and unknown types in railway scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0015] Figure 1Schematic flowchart of the method for detecting foreign object intrusion in the rail transit environment based on a visual model provided by an embodiment of the present invention; Figure 2 Schematic structural diagram of the system for detecting foreign object intrusion in the rail transit environment based on a visual model provided by an embodiment of the present invention. Detailed implementation manners
[0016] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope protected by the present invention.
[0017] Figure 1 Schematic flowchart of the method for detecting foreign object intrusion in the rail transit environment based on a visual model provided by an embodiment of the present invention, including the following steps: S1. Preprocess the original images collected by the on-vehicle cameras of high-speed railways, identify and extract the suspicious regions containing potential foreign objects, and obtain the original complete images and at least one sub-image of the suspicious region; For the original images collected by the on-vehicle cameras of high-speed railways, in the preprocessing stage, an image segmentation model is used to perform a preliminary analysis on the images, and the suspicious regions containing potential foreign objects are identified and extracted. The image segmentation model is based on the deep learning method, locates the abnormal candidate regions through multi-level feature fusion and attention mechanism, and generates the original complete images and multiple sub-images of the suspicious regions. This process uses a sliding window and a region clustering algorithm to screen the candidate boxes that meet the size and position characteristics of the foreign objects, reducing the computational redundancy of subsequent processing. The preprocessed image data is divided into a complete scene view and a local detail view, respectively representing the overall state of the track environment and the microscopic characteristics of the suspected foreign objects.
[0018] S2. Input the original complete images and the sub-images of the suspicious regions into an image encoder, and extract the coarse-grained global feature vectors and the fine-grained local feature vectors respectively; The original complete images and the sub-images of the suspicious regions are input into the image encoder in parallel for feature vector extraction. The image encoder adopts a hierarchical feature extraction structure. The underlying network captures basic visual elements such as lines and textures, the middle layer network aggregates spatial context information, and the deep layer network extracts global semantic features. The original complete images are processed by the encoder and output coarse-grained global feature vectors, representing macroscopic scene information such as the track direction and the layout of the catenary; the sub-images of the suspicious regions are output by the same encoder and output fine-grained local feature vectors, focusing on the details such as the edge shape and the material reflection characteristics of the foreign objects. During the feature vector extraction process, the network retains the cross-level detail information through residual connections, avoiding the loss of local information caused by excessive abstraction of the deep features.
[0019] S3. Generate corresponding text prompt information for the original complete image and the sub - graph of the suspicious area, and extract text feature vectors through a text encoder; The text prompt generation module dynamically constructs semantic descriptions according to the image content. The complete image corresponds to a regular - state text template for the railway scene, such as "straight railway tracks without occlusion" or "catenary equipment operating normally"; the sub - graph of the suspicious area corresponds to a foreign - object feature description template, such as "foreign object hanging above the track" or "scattered obstacles on the rail surface". The text prompt is converted into an embedding vector through a learnable encoder and shares a high - dimensional semantic space with the image features. The text encoder adopts a sequence modeling method based on the attention mechanism to capture the semantic associations between keywords and generate text feature vectors aligned with the dimension of the image feature vectors. The text encoder uses CoOP (Context Optimization for Prompting) to optimize text feature generation through learnable prompt vectors.
[0020] S4. Calculate the multi - level similarity between the coarse - grained global feature vector and the text feature vector of the complete image, and the multi - level similarity between the fine - grained local feature vector and the text feature vector of the suspicious area; The multi - level similarity calculation module measures the matching degree between the image feature vector and the text feature vector through matrix operations. The cosine similarity is calculated between the coarse - grained global feature vector and the scene - description text vector, reflecting the overall consistency between the current scene and the normal railway environment; the similarity is calculated between the fine - grained local feature vector and the foreign - object description text vector to quantify the degree of deviation of the suspicious area from the normal state. The similarity calculation results are mapped to abnormal probability values after normalization. The abnormal probability of the complete image reflects the large - scale environmental abnormal risk, and the abnormal probability of the local sub - graph characterizes the possibility of foreign objects existing in a specific area.
[0021] S5. Weight - fuse the two types of similarities according to the preset weights to obtain a comprehensive abnormal score; S6. Based on a dynamic threshold model, perform distribution modeling on the comprehensive abnormal score to determine whether there are intrusive foreign objects in the image.
[0022] The weight - fusion module presets weight parameters according to the characteristics of the railway monitoring scene. The abnormal score of the complete image focuses on the detection of the integrity of the track structure, and the abnormal score of the sub - graph of the suspicious area focuses on the identification of the risk of foreign - object intrusion. The fused comprehensive abnormal score integrates the multi - scale detection results and avoids the one - sidedness of judgment at a single feature level. The dynamic threshold model constructs a probability density function based on the abnormal score distribution of historical normal samples, updates the model parameters through an online learning mechanism, and adapts to score fluctuations caused by environmental factors such as day - night illumination changes and seasonal vegetation differences in real - time. When the comprehensive abnormal score exceeds the confidence interval of the normal distribution, an alarm signal for foreign - object intrusion is triggered, and the positioning information of the suspicious area is synchronously output for manual review.
[0023] Through the collaborative analysis of image and text modalities, this method strengthens the distinguishability between normal scenarios and abnormal states, and overcomes the false alarm problem of traditional visual detection methods under the interference of complex backgrounds. The hierarchical feature extraction structure takes into account both the global environmental stability assessment and the precise localization requirements of local foreign objects, and the dynamic threshold mechanism improves the robustness of the system under different climate conditions and operating periods. The pre-screening strategy for suspicious areas effectively reduces the consumption of computing resources, meets the real-time processing requirements of in-vehicle devices, and provides reliable technical support for the safe operation of high-speed railways.
[0024] In some embodiments, the multi-level similarity calculation in S4 includes: Extract the first-level feature vector, second-level feature vector, and third-level feature vector of the complete image, as well as the corresponding-level feature vectors of the suspicious area sub-graph, respectively, at the 6th, 9th, and 12th layers of the image encoder; Calculate the similarity Sim between the feature vectors of each level of the complete image and the text feature vector of the complete image Di , and the formula is: ; where represents the image feature vector output by the kth layer of the image encoder, where k ∈ {6, 9, 12} represents the network depth level, represents the text feature vector of the complete image at the corresponding level, generated from the text prompt information containing the description of the railway scene; Calculate the similarity Sim between the feature vectors of each level of the suspicious area sub-graph and the text feature vector of the suspicious area Di j , and the formula is: ; where represents the kth layer image feature vector of the suspicious area sub-graph, represents the text feature vector of the suspicious area at the corresponding level, generated from the text prompt information containing the description of the foreign object feature.
[0025] To make full use of the feature vectors of different levels extracted by the image encoder, this method extracts the image feature vectors at 3 different depths (6th, 9th, and 12th layers respectively). For ease of representation, the entire image encoder is divided into three parts, namely , and . For the input full image, its image feature vectors at different levels are represented as , and , then: , , ; Among them, D i is the i-th complete original image. For the sub-images where abnormal objects may exist, the image feature vectors at three different levels are respectively: , , ; Since the image encoder has strong feature extraction capabilities, for the original complete image and the sub-images of suspicious regions, the two can share the image encoder ( , and ).
[0026] For two different text prompt messages, they are respectively defined as and , and the corresponding generated text feature vectors are respectively and .
[0027] In the process of multi-level similarity calculation, the image encoder extracts image feature vectors for the complete image and the sub-images of suspicious regions at the network depths of the 6th, 9th, and 12th layers respectively. The designs of these three levels correspond to different stages of feature abstraction in the image encoder: the image feature vectors output by the 6th layer network mainly contain low-level geometric structure information such as track edges and catenary supports, the image feature vectors output by the 9th layer network fuse local texture and spatial context relationships (such as the correlation between the reflective characteristics of the rail surface and the surrounding vegetation), and the image feature vectors output by the 12th layer network represent the semantic abstraction information of the global scene (such as the layout of the track direction and the background skyline). The image feature vectors at each level are reduced in dimension and normalized through a fully connected layer to ensure dimension alignment with the text feature vectors.
[0028] For the complete image, its multi-level similarity calculation is achieved through the following formula: ; Among them, represents the image feature vectors output by the 6th, 9th, and 12th layers of the image encoder, It is the text feature vector corresponding to the complete image. This text feature vector is generated by a predefined railway scene description template, such as "straight and unobstructed railway tracks" or "normally operating catenary equipment", and is converted into a high-dimensional vector through a text encoder. The dot product operation in the formula measures the projection consistency of the image feature vector and the text feature vector in the semantic space, and the summation and averaging operations balance the contribution weights of feature vectors at different levels. For example, the feature vector at the 6th layer is sensitive to the details of the track edge, and the feature vector at the 12th layer is sensitive to the overall scene layout. The average calculation can avoid misjudgment caused by environmental interference (such as local reflection or shadow) at a single level.
[0029] For the sub-graph of the suspicious area, its similarity calculation formula is: ; where represents the image feature vector output by the sub-graph of the suspicious area at the same three network layers, is the text feature vector describing the foreign object feature (such as "hanging foreign object" or "scattered obstacle"). The dot product operation reflects the matching degree between the area of the sub-graph of the suspicious area and the semantic description of the foreign object, and the average calculation synthesizes the detection results at different levels. For example, the feature vector at the 6th layer captures the sharpness of the foreign object edge, the feature vector at the 9th layer analyzes the texture contrast between the surface of the foreign object and the surrounding environment, and the feature vector at the 12th layer evaluates the spatial incongruity between the foreign object and the overall scene. If the area of the sub-graph of the suspicious area is a bird or a fallen leaf in a normal environment, its deep features have a low matching degree with the text description, thereby reducing the false alarm probability.
[0030] The dot product operation in the formula essentially maps the image feature vector and the text feature vector to the same semantic space, and measures the similarity through the cosine value of the vector angle. If the image content is highly consistent with the text description, the vector directions tend to be the same, and the dot product result approaches the maximum value; if there are significant differences, the dot product value decreases. The averaging operation avoids the deviation caused by the different network depths of the single-level feature vector: the shallow feature vector (such as the 6th layer) has high resolution but is vulnerable to noise interference, and the deep feature vector (such as the 12th layer) is highly abstract but may ignore local details. For example, for a small foreign object, the high-resolution feature vector at the 6th layer can capture its contour, while the global feature vector at the 12th layer may weaken its response due to the too large receptive field. The average calculation can balance the contributions of the two.
[0031] In implementation, the image encoder and the text encoder are jointly optimized through contrastive learning. During the training phase, the similarity between the complete image of the normal sample and the scene description text is maximized, while the similarity between the sub-graph of the normal region and the foreign object description text is minimized. In the present invention, during the training phase, the present invention only requires the images of the normal samples, without the need for detailed annotation of abnormal objects or the introduction of abnormal samples. This unsupervised training method greatly reduces the cost of data collection and annotation, and also avoids the problem of performance degradation caused by inaccurate or missing annotation. This makes the present invention more efficient in practical applications. During the inference phase, the anomaly score is dynamically generated based on the degree of similarity deviation. For example, when there is a foreign object hanging on the track, the similarity between the coarse-grained global feature vector of the complete image and the scene description text decreases (reflecting the overall scene anomaly), while the similarity between the fine-grained local feature vector of the suspicious region sub-graph and the foreign object description text increases (confirming the presence of the foreign object). Through the multi-angle comparison of hierarchical features, the model can distinguish real foreign objects from environmental interference (such as temporary birds or dynamic shadows), improving the detection robustness in complex scenarios.
[0032] This method enhances the detection ability for foreign objects of different scales (such as large-area collapses and small gravels) by fusing multi-level image feature vectors and cross-modal text semantics. The dynamic weight allocation and hierarchical feature complementary mechanism effectively cope with complex conditions such as light changes and weather interference, avoiding the limitations of traditional single-level feature detection, and providing a high-precision and adaptive solution for railway environmental safety monitoring.
[0033] In some embodiments, the weighted fusion in S5 satisfies: ; where An_score represents the comprehensive anomaly score, represents the anomaly score of the complete image, defined as ; represents the anomaly score of the suspicious region sub-graph, defined as , and α and β are preset weight coefficients, satisfying α + β = 1.
[0034] In the anomaly score fusion stage, the anomaly score of the complete image and the anomaly score of the suspicious region sub-graph are weighted and calculated through preset weights. The anomaly score of the complete image is defined as the complement of the mean value of the multi-level similarity of the complete image, and the formula is expressed as: .
[0035] where, is the mean value of the similarity between the 6th, 9th, and 12th layer features of the complete image and the scene description text features, reflecting the consistency between the current scene and the normal railway environment.
[0036] The anomaly score of the suspicious region sub-graph is defined as the complement of the mean value of the similarity of the suspicious region sub-graph: Among them, is the average similarity between the multi-level features of the sub-graph and the text features of the foreign object description, quantifying the degree of deviation of the sub-graph region from the normal semantics. The weighted fusion formula integrates two types of anomaly scores (the complete image anomaly score and the sub-graph anomaly score of the suspicious region) through linear combination to obtain the comprehensive anomaly score An_score.
[0037] Among them, α and β are preset weight coefficients, satisfying α + β = 1. The physical meaning of this formula is to dynamically balance the sensitivity of global environmental anomalies and local foreign object detection by adjusting the ratio of α and β. For example, in the monitoring of straight sections of the track, if it is necessary to prioritize the identification of foreign objects scattered on the track surface, then increase β to enhance the contribution of the sub-graph anomaly score of the suspicious region; in curved or bridge sections, to prevent the risk of landslides or structural deformation, then increase α to strengthen the weight of the complete image anomaly score.
[0038] In the formula is calculated based on the degree of similarity attenuation between the complete image and the text description of the normal scene. When the overall environment of the track deviates (such as large-scale detachment of the catenary), significantly decreases, increases, triggering a global anomaly warning. reflects the local anomaly probability through the matching degree between the sub-graph of the suspicious region and the foreign object description. For example, when there is a hanging foreign object in the sub-graph region of the suspicious region, decreases, increases, indicating a local intrusion risk.
[0039] The settings of the weight coefficients α and β are optimized according to the actual monitoring scenario requirements. In the training stage, the optimal weight combination is determined through grid search and cross-validation. For example, in a multi-curved region with complex background, set α = 0.6 and β = 0.4 to suppress local false alarms caused by vegetation shaking; in the monitoring of straight tracks, use α = 0.4 and β = 0.6 to enhance the detection ability of tiny foreign objects on the track surface. The inference stage supports dynamic weight adjustment. For example, in rainy or foggy weather, temporarily increase α to compensate for the fuzzy local feature extraction caused by decreased visibility.
[0040] The normalization constraint (α + β = 1) in the formula ensures the stability of the value range of the comprehensive anomaly score, avoiding score scale drift caused by uneven weight distribution. For example, when α = 0.7 and β = 0.3, the comprehensive score is more inclined to reflect large-scale anomalies such as catenary fractures; when α = 0.3 and β = 0.7, it is more sensitive to tiny foreign objects such as track surface gravel and bolt detachment. The fused comprehensive anomaly score is modeled by a mixture of Gaussian models to dynamically distinguish the score distribution boundary between normal and abnormal states, improving the decision-making robustness in complex environments.
[0041] This method realizes the flexible adaptation of detection strategies through a weight-configurable mechanism. In tunnel monitoring, the global weight focuses on the assessment of the stability of the mountain structure, and the local weight monitors rockfalls on the track surface; in the elevated bridge section, the global weight focuses on the deformation trend of the bridge, and the local weight detects bolt loosening or corrosion. The dynamic adjustment ability of the weight coefficient enhances the adaptability of the system to different line characteristics and operating conditions, providing multi-dimensional and refined abnormal perception support for the safe operation of railways.
[0042] In some embodiments, the dynamic threshold model in S6 is a mixture Gaussian model, and its probability density function is: ; where x is the input comprehensive anomaly score vector, μ is the mean vector, Σ is the covariance matrix, n is the dimension of x, T represents the transpose symbol, and π is the pi. The input variable x of the mixture Gaussian model is essentially an extended form of the above-mentioned weighted comprehensive anomaly score (An_score). In the single-frame detection scenario, x is a scalar, that is, x = An_score, directly inheriting the linear fusion result of the above An_score; in the multi-frame time series analysis or multi-region joint monitoring scenario, x is extended to a vector, for example, x = [An_score t , An_score t−1 , An_score 轨道 , An_score 接触网 , where An_score t is the comprehensive anomaly score at the current moment, An_score t−1 is the comprehensive anomaly score at the previous moment, An_score 轨道 is the independent anomaly score of the track area, and An_score 接触网 is the independent anomaly score of the catenary area, and the output results of the above An_score are fused through the spatio-temporal dimension. The core difference between the two lies in the application granularity: the above An_score focuses on the single-frame real-time anomaly scoring, while x improves the decision-making robustness in complex scenarios through probability modeling, and both share the same anomaly scoring generation logic.
[0043] This formula describes the difference in the score distribution between normal and abnormal states through the linear superposition of multiple Gaussian distributions, where the anomaly scores of normal samples are concentrated in the low-value region, and the abnormal samples are distributed in the high-value region.
[0044] In the model training stage, the parameters of the mixture Gaussian model are initialized using the anomaly score data of historical normal samples. The expectation-maximization algorithm iteratively optimizes the mean vector, covariance matrix, and mixing coefficients to maximize the log-likelihood function of the sample data , where \(L\) is the likelihood function, \(K\) is the number of Gaussian distributions, and \(N\) is the total number of training samples. During the model training phase, only the image data of normal railway scenes is used to optimize the network parameters. By means of a contrastive learning strategy, the similarity between the coarse-grained global feature vector of the complete image and the feature vector of the scene description text is maximized ( ), while the similarity between the fine-grained local feature vector of the suspicious region sub-graph and the feature vector of the foreign object description text is minimized ( ). Specifically, the contrastive loss function is defined as: ; where \(\lambda_1\) and \(\lambda_2\) are balance coefficients, and the parameters of the image encoder and the text encoder are jointly optimized through backpropagation, forcing the model to align the semantics of the global image and the scene text in the normal scene and suppressing the false matching between the local region and the foreign object text.
[0045] In some embodiments, in the posterior probability calculation step, the probability \(\gamma(z i )\) that the sample \(x ij \) belongs to the \(j\)-th Gaussian distribution is determined by the following formula: ; where \(K\) is the number of Gaussian distributions, \(\pi j \) is the mixing coefficient of the \(j\)-th distribution, and satisfies , represents the probability density value of the sample point \(x i \) under the \(j\)-th Gaussian distribution, represents the probability density value of the sample point \(x i \) under the \(k\)-th Gaussian distribution. This formula calculates the attribution probability of each sample to each Gaussian distribution for subsequent parameter updates.
[0046] During the parameter update process, the mean vector \(\mu j \) is calculated by weighted average, and the weight is the posterior probability of the sample belonging to the \(j\)-th distribution: ; The covariance matrix \(\Sigma j is updated to the weighted covariance of the sample deviating from the mean: ; The mixing coefficient \(\pi j is updated to the proportion of samples in the corresponding distribution: ; where \(N\) is the total number of training samples, \(\mu j is the updated mean vector of the \(j\)-th Gaussian distribution, \(\pi j is the updated mixing coefficient (the same parameter as described above "\(\pi j is the mixing coefficient of the \(j\)-th distribution", that is, \(\pi j is updatable), \(\Sigmaj is the updated covariance matrix, and T represents the transpose symbol.
[0047] The iterative optimization process continues until the log-likelihood function converges to obtain stable Gaussian distribution parameters.
[0048] Finally, the conditional probability of each sample is obtained according to the optimized model parameters, and then its maximum value j is obtained, that is , where argmax represents finding the input parameter j that makes the function γ(z ij ) achieve the maximum value, that is, the i-th sample is assigned to the most likely Gaussian distribution, and λ i is the classification label, and its value is 1 or 2, representing that the i-th sample is determined to be normal or abnormal respectively. In the embodiment of the present invention, it can be considered that the obtained anomaly scores follow two distributions (the anomaly scores of normal images and abnormal images follow different distributions respectively). These anomaly scores are used as the input of the GMM model, and the model can learn parameters such as the means and variances of the two different distributions (the mean of the anomaly score distribution of normal images is less than the mean of the anomaly score distribution of abnormal images). Then, for each anomaly score, the GMM model can obtain the probability that it belongs to different distributions, and further obtain which category the anomaly score belongs to, that is, judge whether the corresponding image is normal or abnormal.
[0049] In the initialization stage of the Gaussian mixture model, the initial parameters are estimated using the comprehensive anomaly score data of historical normal samples, including the mean vector μ, the covariance matrix Σ, and the mixing coefficient π j . After the model training is completed, an online update mechanism is supported: the anomaly scores of the latest normal samples are collected through a sliding window, and the expectation-maximization algorithm is periodically executed to update the model parameters, dynamically adapting to the offset of the anomaly score distribution caused by environmental factors such as changes in light intensity and seasonal vegetation coverage. When it is detected that the comprehensive anomaly score exceeds the confidence interval of the current Gaussian distribution, a foreign object intrusion alarm is triggered.
[0050] In the inference stage, the real-time calculated comprehensive anomaly score is input into the trained Gaussian mixture model to calculate the probability that it belongs to each Gaussian distribution. If the probability of the sample in the normal distribution (the Gaussian component with a low mean and a small variance) is lower than the preset confidence threshold, it is determined to be abnormal. For example, when the visibility of the track environment decreases due to heavy rain, the overall anomaly score distribution of normal samples will shift to the right. The model dynamically expands the confidence interval of the normal distribution by adaptively adjusting the distribution parameters, avoiding false alarms caused by environmental interference.
[0051] This method distinguishes the dynamic boundary between normal and abnormal states through probability density modeling. For example, in the day-night alternation scenario, during the day when there is sufficient light, the normal scores are concentrated in the low-value area. At night, the activation of the supplementary lighting equipment may cause a slight right shift in the score distribution. The Gaussian mixture model captures such distribution changes through multi-component fitting, avoiding the imbalance in day-night detection sensitivity caused by fixed thresholds. The update mechanism of the covariance matrix can characterize the fluctuation characteristics of abnormal scores. For example, the impact of vegetation growth caused by seasonal changes on the coarse-grained global feature vector reduces the risk of misjudgment by expanding the covariance range of the normal distribution.
[0052] The multi-distribution characteristics of the Gaussian mixture model can simultaneously model multi-modal data in complex scenarios. For example, the normal scores in the straight track and curve areas may show different distributions, which are characterized by two Gaussian components by the model, avoiding missed detections caused by insufficient fitting of a single distribution. When making an abnormal determination, if the score deviates from all normal distributions, an alarm is triggered. This method overcomes the defect of the traditional single Gaussian model's insufficient modeling ability for asymmetric distribution data, and is especially suitable for the problem of score distribution diversity caused by line section differences and weather changes in railway scenarios.
[0053] In technical implementation, the number K of the Gaussian mixture model is set according to the distribution complexity of historical data. For the straight track scenario with stable background, K = 2 can cover the day-night difference; for complex lines with multiple curves and bridges, K = 3 or higher is used to fit the normal state distributions of different sections. The online update mechanism of the model supports regularly incorporating the latest normal sample data and dynamically correcting the distribution parameters to adapt to the gradual environmental changes (such as vegetation growth and equipment aging) during long-term operation, ensuring the continuous effectiveness of the detection system.
[0054] In some embodiments, the image encoder adopts a hierarchical feature extraction structure, and the image feature vector I output by its k-th layer fk satisfies: ; where represents the k-th layer network of the encoder, is the image feature vector output by the previous layer.
[0055] This formula discloses the hierarchical transfer mechanism of the feature vector in the encoder: the k-th layer network performs convolution, normalization, and non-linear activation operations on the input feature vector, gradually extracting higher-level semantic information. For example, the first layer network (shallow layer) extracts basic visual features such as rail edges and ballast textures; the sixth layer network (middle layer) aggregates the relative position relationship between the track and the catenary; the twelfth layer network (deep layer) captures the global semantics of the scene, such as the spatial layout of the track direction and the background skyline.
[0056] Each layer of the encoder's network module is composed of a residual connection structure. The input feature vector is added to the convolutional output feature vector through a skip connection, alleviating the problem of vanishing gradients in the training of deep networks. Taking the 6th layer of the network as an example, the input feature vector extracts local details through a 3×3 convolutional layer, and after being processed by batch normalization and the ReLU activation function, it is added to the original input feature vector, and the output is obtained. This design ensures that the shallow detail information is retained in the deep network, avoiding the loss of local feature vectors caused by the increase in network depth. For example, the edge information of tiny foreign objects on the rail surface is prominent in the shallow features and is transmitted to the deep network through residual connections to assist in considering local anomalies during global semantic analysis.
[0057] The output feature vectors at different levels have different resolutions and receptive fields. Shallow feature vectors (such as the 1st - 3rd layers) retain high spatial resolution and can accurately locate the edge contours of foreign objects; middle-level feature vectors (such as the 4th - 9th layers) reduce the resolution through downsampling to expand the receptive field to capture the context association between foreign objects and the surrounding environment; deep feature vectors (such as the 10th - 12th layers) are further abstracted to represent the global structural stability of the scene. For example, when detecting hanging foreign objects, the 6th layer feature vector combines the local structural information of the catenary support to judge the rationality of the hanging position of the foreign object; the 12th layer feature vector identifies large-scale anomalies caused by the overall offset of the track or landslides through global layout analysis.
[0058] The progressive relationship between feature vector levels is optimized through parameter sharing and weight linkage. During the training process, the encoder synchronously updates the network parameters of each layer through backpropagation, forcing the shallow network to focus on basic feature extraction and the deep network to focus on high-order semantic modeling. For example, in a normal scene, the deep network learns to encode patterns such as the straight track direction and the regular arrangement of the catenary into low-dimensional vectors; when track deformation occurs, the deep feature vector shows a significant deviation from the normal semantic description, triggering an increase in the overall image anomaly score.
[0059] This hierarchical structure supports multi-scale feature fusion, enhancing the detection robustness in complex scenarios. During the inference stage, the feature maps output by the 6th, 9th, and 12th layers are respectively calculated for similarity with the corresponding text descriptions, and the contributions of different levels are balanced through mean fusion. For example, when detecting ballast on the rail surface, the high-resolution feature vector of the 6th layer captures the edge of the ballast, the 9th layer feature vector analyzes the texture difference between the ballast and the roadbed, and the 12th layer feature vector evaluates whether the distribution of the ballast conforms to the normal state of the track. The three work together to reduce missed detections or false alarms caused by misjudgments at a single level.
[0060] In the technical implementation, the image encoder is based on the Vision Transformer architecture. It divides the input image into image patches of a fixed size and generates serialized token vectors through linear embedding and positional encoding. Each layer of the network consists of a multi-head self-attention mechanism and a feed-forward neural network, and calculates the global dependencies between image patches through self-attention weights. For example, the 6th layer of the network strengthens the spatial association between the railway tracks and the catenary through the self-attention mechanism, and suppresses the interference of irrelevant background areas; the 12th layer of the network models the global consistency between the overall trend of the track and the geographical environment. This architecture makes full use of the global modeling ability of Transformer, makes up for the deficiency of traditional convolutional networks in capturing long-range dependencies, and improves the detection accuracy of large-range environmental anomalies.
[0061] Through the design of hierarchical progression and residual connection, this method achieves a balance between local detail retention and global semantic extraction. Under complex conditions such as light changes and partial occlusion, the shallow feature vectors provide stable edge and texture information, and the deep feature vectors ensure the overall rationality judgment of the scene. The two work together to enhance the adaptability of the model to unknown foreign object types and environmental interference, providing reliable technical support for high-speed railway safety monitoring.
[0062] In some embodiments, the generation of the suspicious region sub-graph satisfies: ; where S is the image segmentation model, where D i represents the input original image, is the j-th suspicious region sub-graph of the output.
[0063] The image segmentation model adopts a pre-trained semantic segmentation architecture, extracts multi-scale features of the image through an encoder-decoder structure, and generates a pixel-level segmentation mask. The encoder part gradually downsamples through convolutional layers and residual modules, capturing the semantic information of regions such as railway tracks, catenaries, and background vegetation; the decoder part restores the spatial resolution through upsampling and skip connections, and outputs the boundaries of target regions with high precision.
[0064] The segmentation model generates candidate regions on the original image through a sliding window strategy. Each window region undergoes feature extraction and classifier judgment to calculate its probability score of belonging to the foreign object region. Windows with probability scores higher than the preset threshold are retained as candidate regions, and redundant candidate boxes with high overlap are removed through the non-maximum suppression algorithm. For example, for the scene of scattered gravel on the rail surface, the model identifies candidate regions through local texture contrast and edge irregularity; for the scene of hanging foreign objects, it locates the abnormal region based on the interruption of the structural continuity of the catenary support.
[0065] In the post - processing stage of candidate regions, the segmentation results are screened by combining prior knowledge. According to the typical features of railway foreign object intrusion (such as the size range of foreign objects, the aspect ratio threshold, and the relative position relationship with the track), the mis - detected regions that do not meet the conditions are filtered. For example, transient interference objects such as birds are excluded because of their too - small size or position deviation from the center line of the track, while the regions that meet the preset foreign - object characteristics are retained as suspicious sub - images. The screened sub - images are processed by cropping and normalization to ensure the size consistency of the input to the image encoder.
[0066] In the formula, the segmentation model S uses a segmentation network based on the attention mechanism. Its encoder outputs multi - level feature maps, and the decoder enhances the significant features of the foreign - object region through a channel attention module. For example, in the deep - layer feature map of the encoder, the spatial position of the foreign object hanging above the track is focused through spatial attention weights; in the shallow - layer feature map, the contrast between the foreign - object edge and the background is enhanced through channel attention. The multi - scale feature fusion module stitches together feature maps of different levels and generates the final segmentation mask through convolutional layers to accurately frame the suspicious regions.
[0067] The advantage of this method lies in the dual - screening mechanism that combines semantic segmentation and prior rules. The segmentation model learns the common features of foreign - object regions in railway scenes (such as irregular shapes and material differences from the background) through end - to - end training, while rule - based filtering eliminates mis - detections that do not conform to physical laws based on domain knowledge. For example, the model may misjudge the shadow of a ballast pile as a foreign object, but the aspect ratio and position constraints can identify it as part of the normal ballast structure, thus avoiding false alarms. The dynamic threshold adjustment mechanism adaptively adjusts the segmentation probability threshold according to the environmental light intensity, appropriately reducing the threshold under low - light conditions to compensate for the problem of fuzzy feature extraction.
[0068] In technical implementation, the segmentation model supports multi - task joint optimization, and simultaneously outputs the segmentation mask and confidence score of the foreign - object region. The confidence score reflects the degree of certainty of the model about the existence of a foreign object in the current region and is used for subsequent sub - image sorting and selective processing. For example, in real - time monitoring, only the top - five sub - images with the highest confidence scores are subjected to fine - grained feature extraction and abnormal score calculation, significantly reducing the consumption of computing resources. The confidence score and the abnormal score calculated by multi - level similarity calculation work together to further improve the efficiency and accuracy of the detection system.
[0069] This method significantly reduces the computational workload of subsequent image encoding and cross-modal matching through precise pre-screening of suspicious regions. For example, only a few high-confidence sub-images in the complete image are subjected to local feature extraction, avoiding indiscriminate processing of all regions of the entire image. At the same time, the attention mechanism of the segmentation model enhances the saliency of the foreign object region, making the local feature extraction more focused on key details and improving the discriminative ability of fine-grained similarity calculation. This design not only ensures the detection accuracy but also meets the stringent requirements of in-vehicle devices for real-time performance and computational efficiency, providing efficient and reliable technical support for the safe operation of high-speed railways.
[0070] For experimental research and comparison, the present invention collected and organized a real high-speed railway anomaly image dataset. First, videos of train operations were taken on high-speed railways in different sections. Then, 2000 images were extracted, among which 1600 images without abnormal objects were used as the training set; 400 images were used as the test set, including 200 images without abnormal objects and 200 images with abnormal objects, and the dataset was named RailwayAnomaly. When evaluating the detection performance of the model on this dataset, following the general image anomaly detection standard, precision and recall were used as key metrics. Precision represents the proportion of actual positive samples among the instances predicted as positive samples by the model. Recall indicates the proportion of actual positive samples correctly identified as positive samples by the model. To provide a comprehensive evaluation, we used the F1 score, which combines precision and recall into a single metric: 。
[0071] In addition, AUROC was used as another evaluation metric. It evaluates the overall performance of the model at different thresholds by calculating the area under the ROC curve. The ROC curve shows the relationship between the true positive rate (TPR) and the false positive rate (FPR) of the model at different threshold settings. The AUROC value ranges from 0 to 1, and a higher value indicates better classification performance of the model. The AUROC value of a perfect classifier is 1, while that of a classifier making random guesses is close to 0.5. Therefore, AUROC can intuitively reflect the classification effect of the model, facilitating quantitative comparison of model performance.
[0072] A comparative experiment was conducted on the RailwayAnomaly dataset between the present invention and other existing abnormal object detection methods, and the experimental results are shown in Table 1.
[0073] Table 1 - Comparison of Experimental Results of Different Models
[0074] Table 1 shows the results of different models on the RailwayAnomaly dataset. The method RailWay-VLM proposed in the present invention reaches the optimal level in both F1-score and AUROC score, which are 0.9267 and 0.9725 respectively. The reason why the present invention is superior to other methods is that although other methods have achieved good results in other fields, they cannot achieve good results in the railway operation environment. The detection of abnormal objects based on the railway environment is carried out in a real environment with complex backgrounds and drastic changes in weather and lighting. Therefore, methods such as PaDim cannot effectively complete the task of detecting abnormal objects in such a complex and changeable environment. The reason why RailWay-VLM can achieve better results is twofold. First, using the SAM segmentation network to extract suspicious abnormal regions reduces the difficulty of detection; second, the multimodal-based method proposed in the present invention not only improves the detection accuracy but also reduces the false alarm rate.
[0075] The present invention also verifies the effectiveness of each module proposed in RailWay-VLM. Raw represents using only the basic image encoder and the ordinary text encoder. Feat3 means that feature outputs are performed at 3 different depths in the image encoder and the similarity calculation is carried out with the text features. Prompt Learner means using a learnable text encoder. The experimental results are shown in Table 2.
[0076] Table 2 - Effects of Different Modules in the Model
[0077] From the experimental results in Table 2, it can be seen the improvement effect of the model performance when different modules are added to the basic model. The experiments in the above table show that both the Feat3 and Prompt Learner modules improve the model performance.
[0078] When preprocessing the image, it is necessary to find out the regions in the figure where abnormal objects may exist. In the specific implementation, either traditional image processing methods (Selective Search) can be used, or large model-based methods (SAM) can be used. The present invention compares these two methods in terms of the number of candidate regions and the performance on the railway anomaly dataset (RailwayAnomaly Dataset). The experimental results are shown in Table 3.
[0079] Table 3 - Comparison of Experimental Results of Different Candidate Region Selection Algorithms
[0080] As can be seen from Table 3, both the Selective Search-based method and the SAM-based method can detect the regions with anomalies in the image. Under the two metrics of F1-score and AUROC for the anomaly detection task, the SAM-based method performs better than Selective Search when the number of effective candidate regions is small. Selective Search performs segmentation based on low-level features and lacks the understanding of semantic information, so there are redundancies in the detected candidate regions. While the SAM-based model can learn more semantic information from multiple large-scale datasets, so it can more effectively detect the potentially anomalous regions and avoid detecting redundant information.
[0081] After obtaining the anomaly scores of the complete image and the sub-images of the suspicious regions that may contain anomalies, it is necessary to sum up these two parts of anomaly scores. In the present invention, different weights are tried for summation to obtain a higher anomaly detection accuracy. The experimental results are shown in Table 4.
[0082] Table 4 - Comparison of experimental results of summing two anomaly scores with different weights
[0083] As can be seen from the experimental results in Table 4, when the weight ratio of the two is 4:1, the model achieves the best effect, where the F1-score is 0.9152 and the AUROC is 0.9725.
[0084] Figure 2 The following is a schematic structural diagram of a rail transit environment foreign object intrusion detection system based on a vision model provided by an embodiment of the present invention, including: A preprocessing module 201, configured to preprocess the original images collected by high-speed railway vehicle-mounted cameras, identify and extract the suspicious regions containing potential foreign objects, and obtain the original complete images and at least one sub-image of the suspicious region; A first extraction module 202, configured to input the original complete images and the sub-images of the suspicious regions into an image encoder, and extract coarse-grained global features and fine-grained local features respectively; A second extraction module 203, configured to generate corresponding text prompt information for the original complete images and the sub-images of the suspicious regions respectively, and extract text feature vectors through a text encoder; A calculation module 204, configured to calculate the multi-level similarities between the coarse-grained global features and the text features of the complete images, and the multi-level similarities between the fine-grained local features and the text features of the suspicious regions; A fusion module 205, configured to perform weighted fusion on the two types of similarities according to preset weights to obtain a comprehensive anomaly score; A determination module 206, configured to perform distribution modeling on the comprehensive anomaly score based on a dynamic threshold model, and determine whether there is an intrusion foreign object in the image.
[0085] The system provided by the embodiments of the present invention has the same technical features as the above method embodiments, and thus can achieve the same technical effects, which will not be elaborated herein.
[0086] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting foreign object intrusion in the rail transit environment based on a vision model, characterized in that It includes the following steps: S1. Preprocess the original images collected by the on-vehicle cameras of high-speed railways, identify and extract the suspicious regions containing potential foreign objects, and obtain the original complete images and at least one sub-image of the suspicious region; S2. Input the original complete images and the sub-images of the suspicious region into an image encoder to extract the coarse-grained global feature vectors and the fine-grained local feature vectors respectively; S3. Generate corresponding text prompt information for the original complete images and the sub-images of the suspicious region respectively, and extract the text feature vectors through a text encoder; S4. Calculate the multi-level similarities between the coarse-grained global feature vectors and the text feature vectors of the complete images, and the multi-level similarities between the fine-grained local feature vectors and the text feature vectors of the suspicious region sub-images; S5. Perform weighted fusion on the two types of similarities according to the preset weights to obtain the comprehensive anomaly score; S6. Based on the dynamic threshold model, perform distribution modeling on the comprehensive anomaly score to determine whether there are intrusion foreign objects in the images.
2. The method according to claim 1, characterized in that The multi-level similarity calculation in S4 includes: Extract the first-level feature vectors, second-level feature vectors, and third-level feature vectors of the complete images, and the corresponding-level feature vectors of the sub-images of the suspicious region respectively at the 6th, 9th, and 12th layers of the image encoder; Calculate the similarity Sim between the feature vectors of each layer of the complete image and the text feature vector of the complete image Di , and the formula is as follows: ; Among them, represents the image feature vector output by the k-th layer of the image encoder, where k ∈ {6, 9, 12} represents the network depth level, represents the complete image-text feature vector of the corresponding level, which is generated by the text prompt information containing the description of the railway scene; Calculate the similarity Sim between the feature vectors of each level of the sub-graph in the suspicious area and the text feature vector of the suspicious area Di j , and the formula is: ; Among them, represents the k-th layer image feature vector of the suspicious region sub-graph, represents the suspicious region text feature vector of the corresponding level, which is generated from the text prompt information containing the foreign object feature description.
3. The method according to claim 2, characterized in that, The weighted fusion in S5 satisfies: ; Among them, An_score represents the comprehensive anomaly score, represents the complete image anomaly score, which is defined as ; represents the anomaly score of the suspicious region sub - graph, which is defined as , where α and β are preset weight coefficients, and α + β = 1.
4. The method according to claim 1, wherein The dynamic threshold model in S6 is a mixture Gaussian model, and its probability density function is: ; where x is the input comprehensive anomaly score vector, μ is the mean vector, Σ is the covariance matrix, n is the dimension of x, T represents the transpose symbol, and π is the pi.
5. The method according to claim 4, characterized in that The Gaussian mixture model optimizes parameters through the expectation-maximization algorithm and iteratively calculates the sample x i The posterior probability γ(z ij ) belonging to the j-th Gaussian distribution is as follows: ; where K is the number of Gaussian distributions, and π j is the mixing coefficient of the j-th distribution and satisfies , represents the probability density value of the sample point x i under the j-th Gaussian distribution, represents the probability density value of the sample point x i under the k-th Gaussian distribution.
6. The method according to claim 5, wherein The update of the model parameters satisfies: ; ; ; where N is the total number of training samples, μ j is the mean vector of the j-th Gaussian distribution after update, π j is the mixing coefficient after update, Σ j is the covariance matrix after update, and T represents the transpose symbol.
7. The method according to claim 2, characterized in that, The image encoder adopts a hierarchical feature extraction structure, and the image feature vector I output by its k-th layer fk satisfies: ; Among them, represents the k-th layer network of the encoder, is the image feature vector output by the previous layer.
8. The method according to claim 1, wherein The generation of the sub-images of the suspicious region satisfies: ; Among them, S is the image segmentation model, and D i represents the input original image, is the j-th suspicious region sub-graph of the output.
9. A foreign object intrusion detection system for rail transit environment based on a vision model, characterized in that, It includes: A preprocessing module for preprocessing the original images collected by the on-vehicle cameras of high-speed railways, identifying and extracting the suspicious regions containing potential foreign objects, and obtaining the original complete images and at least one sub-image of the suspicious region; A first extraction module for inputting the original complete images and the sub-images of the suspicious region into an image encoder to extract the coarse-grained global feature vectors and the fine-grained local feature vectors respectively; A second extraction module for generating corresponding text prompt information for the original complete images and the sub-images of the suspicious region respectively, and extracting the text feature vectors through a text encoder; A calculation module for calculating the multi-level similarities between the coarse-grained global feature vectors and the text feature vectors of the complete images, and the multi-level similarities between the fine-grained local feature vectors and the text feature vectors of the suspicious region sub-images; A fusion module for performing weighted fusion on the two types of similarities according to the preset weights to obtain the comprehensive anomaly score; A determination module for performing distribution modeling on the comprehensive anomaly score based on the dynamic threshold model to determine whether there are intrusion foreign objects in the images.
Citation Information
Patent Citations
Cross-modal retrieval method and system based on multi-granularity feature fusion
CN115391625A
Model training method, video classification method, device and equipment
CN116563669A
Palm vein recognition method, device and equipment and storage medium
CN116758600A
Dynamic identification method for abnormal water quality
CN119249332A
Pedestrian re-identification method combined with text guidance
CN119274207A
Cited By
Industrial scene adaptive continuous learning foreign matter detection method
CN120931900A
Method and device for detecting surface defects of inner cavity of cavity and electronic equipment
CN121167573A
Text content review system based on artificial intelligence
CN121279318A
Urban governance scene-oriented abnormal event handling method and device
CN121459242A
Method for sensing track state based on multi-modal monitoring data fusion of ballastless track
CN121527721A