Subway key component fault detection method and system based on AI visual large model
By introducing a large-scale pre-trained visual model and an adaptive feature enhancement mechanism, an end-to-end intelligent detection system was constructed, which solved the problems of accuracy, generalization ability and real-time performance in the fault detection of key subway components. It achieved high-precision and low-latency fault identification, and improved detection efficiency and reliability.
Patent Information
- Application Number
- CN202511923667.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-01-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing fault detection technologies for key subway components suffer from systemic bottlenecks in terms of detection accuracy, generalization ability, multimodal understanding, real-time performance, and state evolution modeling. These limitations make it difficult to achieve high-precision, robust, and adaptive fault identification and localization, and fail to meet the requirements of high-end rail transit equipment for highly reliable and robust online monitoring.
By introducing a large-scale pre-trained visual model and combining an adaptive feature enhancement mechanism with a spatiotemporal consistency verification architecture, an end-to-end intelligent detection system is constructed. Images are acquired through a multi-view high-definition industrial camera array, and the illumination compensation algorithm based on Retinex theory and the deep semantic parsing of the Transformer architecture are combined. Long short-term memory networks are used to capture the changing trends of components, and a fault discrimination decision unit is combined to achieve high-precision, low-latency fault identification.
It significantly improves the ability to identify minute, hidden, and atypical defects, increasing the detection accuracy by 19.7 percentage points, reducing the false negative rate to below 0.8%, and decreasing the false alarm rate by 42%. It achieves fully automated, non-contact online detection, increasing detection efficiency by more than 12 times and reducing deployment costs by 28%.
Smart Images

Figure CN121364081A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence and computer vision, and particularly relates to a subway key component fault detection method and system based on an AI vision large model. BACKGROUND
[0002] With the deep integration of artificial intelligence and the rail transit industry, the safety and intelligent level of subway operation have increasingly become important indicators for measuring the quality of urban public transportation services. A subway system contains a large number of key mechanical and electrical components, such as bogies, brake discs, pantographs, and track fasteners, and the running state of these components is directly related to the safety and efficiency of train operation.
[0003] In the current field of subway equipment state monitoring, traditional fault detection methods mainly rely on manual inspection and rule-based image recognition technology. However, these methods have many limitations in practical applications. First, the detection accuracy is limited by environmental interference and feature extraction capability, making it difficult to deal with small defects or gradual damage under complex working conditions. Second, the model generalization ability is weak, and repeated training and parameter tuning are required for different lines, vehicle types, or lighting conditions, resulting in high deployment costs and poor adaptability. Third, existing methods mostly use small-scale neural networks or traditional machine learning models, which lack sufficient semantic understanding of multi-dimensional visual information (such as visible light, infrared, and depth images), making it difficult to achieve cross-modal feature alignment and context awareness. Fourth, the system response is not real-time enough, and there is processing delay when dealing with massive video stream data, making it difficult to meet online detection requirements. Finally, there is a lack of modeling mechanism for component degradation trends, and historical state evolution is not incorporated into the decision-making process, leading to delayed warnings or frequent false alarms.
[0004] To address the systematic bottlenecks in existing subway key component fault detection technologies in terms of detection accuracy, generalization ability, multi-modal understanding, real-time performance, and state evolution modeling, a new intelligent detection system is needed that can deeply integrate the cognitive ability of AI vision large models and the prior knowledge of industrial detection scenarios, to achieve high-precision, strong-robust, and self-adaptive fault recognition and positioning, and to comprehensively improve the automation and intelligent level of subway operation and maintenance. SUMMARY
[0005] The purpose of the present application is to make up for the shortcomings of the prior art, and provide a subway key component fault detection method and system based on AI visual large model, which can effectively solve the problems in the background art. In current subway operation, key mechanical structural components such as bogies, brake discs and couplings are in a long-term high-load running state, and early faults such as small cracks, deformation or looseness are difficult to be found in time through conventional manual inspection. The traditional rule-based image recognition method is sensitive to complex background interference, and the feature extraction capability is limited, which cannot adapt to the precise identification needs of multi-scale, multi-morphology and low-contrast defects. At the same time, the existing automatic detection system relies on fixed threshold judgment mechanism, and lacks dynamic modeling ability for the evolution trend of component running state, resulting in high false alarm rate and high false alarm rate, which is difficult to meet the technical requirements of high-end rail transit equipment for high reliability and high robustness online monitoring. The present application introduces a large-scale pre-trained visual large model, combines an adaptive feature enhancement mechanism and a spatiotemporal consistency verification architecture, and constructs an end-to-end intelligent detection system to realize high-precision, strong generalization and low-delay identification of surface and structural abnormalities of subway key components, and significantly improve the timeliness and accuracy of fault warning.
[0006] To achieve the above purpose, the present application provides the following technical scheme: on the one hand, a subway key component fault detection system based on AI visual large model, the system comprises the following components: a data acquisition unit for synchronously acquiring original visual image sequences of key components through a multi-view high-definition industrial camera array deployed on both sides and the bottom of the track during the process of subway vehicle entering the station or returning to the warehouse, and recording the corresponding time stamp and position information; an image preprocessing module connected to the data acquisition unit for performing distortion correction, illumination normalization, noise suppression and interested region cropping operation on the received original image to generate standardized input data; a visual large model inference engine receiving standardized images output from the image preprocessing module, using a large-scale convolutional neural network based on Transformer architecture to perform deep semantic analysis on the image and extract multi-level and multi-granularity spatial feature representation; a time series state modeling module coupled with the visual large model inference engine, using a long short-term memory network to dynamically aggregate feature vectors of the same component region between consecutive frames to capture the trend pattern of component appearance change; a fault discrimination decision unit integrating a classification head and an abnormal score mechanism, outputting component health state label and confidence score based on fused spatial-temporal features; an alarm and feedback interface for pushing the determination result to the operation and maintenance management platform in real time, and supporting automatic optimization of discrimination threshold parameters according to historical cases; Preferably, the visual large model inference engine adopts a ViT-Adapter structure pre-trained on a massive industrial defect dataset, embeds a pluggable adaptation layer in the backbone network, and enables the model to quickly adapt to the local texture characteristics of specific subway components while maintaining general representation capabilities; Furthermore, the adaptation layer contains a lightweight convolution projection module and a channel attention gate unit to enhance key area feature responses and suppress irrelevant background activations; In addition, an adaptive illumination compensation algorithm based on the Retinex theory is introduced into the image preprocessing module, which effectively alleviates the imaging unevenness problem caused by car bottom shadows and reflective metal surfaces by estimating the illumination component and reconstructing the reflection component. Preferably, the time series state modeling module uses a bidirectional LSTM structure to encode the feature sequence within a sliding time window, with each time step input being a high-level semantic feature vector extracted at different time points for the same physical location, and the output being a state hidden variable containing context-dependent relationships; Furthermore, the sliding window length is set to 8 to 16 frames to ensure coverage of at least one complete wheelset rotation period to eliminate artifacts caused by periodic posture changes; In addition, the fault discrimination decision unit is configured with a dual-path output mechanism, one path performs a multi-class classification task to identify specific fault types, and the other path calculates the deviation of the current sample from the normal class cluster center based on the Mahalanobis distance to generate a continuous abnormality index. Preferably, the industrial camera array in the data acquisition unit is arranged in a stereo distribution manner, covering multiple observation angles such as top view, side view, and elevation view, and at least one camera is equipped with a polarizing filter to reduce the influence of metal component specular reflection; Furthermore, the camera trigger signals are precisely synchronized by the train positioning system to ensure that the image acquisition time strictly corresponds to the car position, achieving spatial coordinate mapping consistency; In addition, the system has built-in edge computing nodes deployed in the local cabinet at the station, which are responsible for real-time processing tasks from image reception to preliminary inference, with an end-to-end response delay controlled within 300 milliseconds; Preferably, the visual large model inference engine adopts a curriculum learning strategy during the fine-tuning stage, first uses a large-scale public defect dataset for preliminary transfer learning, and then gradually introduces labeled fine subway component-specific data for hierarchical progressive training according to fault difficulty; Furthermore, gradient clipping and weight decay constraints are applied during training to prevent the model from overfitting to a small number of high-frequency samples; In addition, the fault discrimination decision unit supports an online incremental update mechanism, and new samples labeled by maintenance personnel are automatically added to the negative sample library after review for periodic retraining of the classifier to expand the recognition boundary; Preferably, the image preprocessing module implements a ROI extraction method based on superpixel segmentation, divides the image into uniform and compact local regions using the SLIC algorithm, screens out superpixel blocks located within the contour range of key components combined with prior knowledge, and only performs subsequent high-cost model reasoning on the region; Furthermore, the ROI extraction process introduces a dynamic mask update mechanism, preloads a three-dimensional projection template according to the vehicle model and component geometric parameters, and adaptively adjusts the mask boundary according to the actual imaging results; In addition, the system has cross-line deployment compatibility, decouples the data input interface and the core algorithm module, and supports plug-and-play access of different manufacturers and different resolution imaging devices; In another aspect, a subway key component fault detection method based on an AI vision large model, the method comprising the following steps: S110, synchronously collecting multi-angle visible light images of subway key components by a trackside multi-view industrial camera array when a train passes at low speed, and adding timestamp and carriage number information; S120, performing distortion correction, light equalization and noise filtering processing on the collected original images to generate a standard image set with uniform size and color space; S130, using a pre-trained vision large model to perform forward reasoning on the standard images, extracting deep convolution feature maps and pooling them into global feature vectors; S140, arranging the feature vectors extracted from the same component in consecutive time periods in chronological order, and inputting them into a bidirectional time encoding network to obtain state representations with context awareness; S150, performing classification discrimination and abnormal score calculation based on the fused state representations, and outputting component health state categories and risk levels; S160, when the determination result exceeds a preset safety threshold, generating a structured alarm message and uploading it to a central operation and maintenance platform, and recording the original images and intermediate features for subsequent traceability analysis; Furthermore, the vision large model in S130 is trained in the field adaptation mode of freezing the backbone network, fine-tuning the top classification head and the adapter module, and the training data contains labeled samples of not less than 5 typical fault modes, and the number of samples of each class is not less than 2000; In addition, the hidden layer dimension of the bidirectional time encoding network in S140 is set to 512, and the dropout ratio is set to 0.3, to balance the expression ability and the generalization performance; Furthermore, the abnormal score calculation in S150 adopts a probability modeling method based on kernel density estimation to evaluate the likelihood of the current feature under the normal sample distribution, and a potential abnormality is determined if it is lower than the 5th percentile; In addition, the structured alarm message in S160 contains fields such as fault location, suspected type, confidence value, occurrence time and associated image storage path, which supports direct connection with the asset management system; Compared with the prior art, the present application has the following advantages: By introducing a large-scale pre-trained visual large model, the recognition ability of small, hidden and atypical defects is significantly improved, and the average detection accuracy is improved by 19.7 percentage points compared with the traditional CNN method, and the missed detection rate is reduced to below 0.8%; A space-time joint modeling paradigm is constructed, and the feature evolution rule between consecutive frames is used to enhance the discrimination stability and effectively suppress the single frame misjudgment propagation, so that the overall false alarm rate of the system is reduced by 42%; An automatic and non-contact online detection process is realized, which replaces manual visual inspection, and a single vehicle scanning can be completed in 45 seconds, and the detection efficiency is improved by more than 12 times; A modular decoupling design and edge computing deployment architecture are adopted, which has good scalability and cross-line reuse capability, and the deployment cost is reduced by 28% compared with similar systems, and the maintenance convenience is greatly improved. BRIEF DESCRIPTION OF DRAWINGS
[0007] Figure 1 is the overall technical scheme architecture schematic diagram of the AI visual large model subway key component fault detection method and system proposed by the present application; Figure 2 is the core principle framework schematic diagram of the adaptive feature enhancement and space-time consistency verification fusion mechanism in the present application. DETAILED DESCRIPTION
[0008] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the specific embodiments, structures, features and effects according to the present application are described in detail as follows in combination with the drawings and preferred embodiments.
[0009] Embodiment one Please refer to Figure 1 and Figure 2 In the daily operation and maintenance scene of a certain urban rail transit A line vehicle depot, the subway train needs to return to the warehouse for routine state inspection after completing the operation task every day. The traditional manual inspection method depends on the visual observation of the technical personnel to observe whether the key mechanical components such as bogie, brake disc and coupling exist cracks, wear or looseness, which is limited by the short operation time window, visual fatigue accumulation and uneven illumination under the complex vehicle bottom environment, and the micro defect missed detection rate is high. In order to improve the detection reliability and automation level, an AI visual large model subway key component fault detection system is deployed in the present embodiment, which realizes non-contact, high-precision and full-process automatic online monitoring.
[0010] The system first performs image acquisition operation through a data acquisition unit. The data acquisition unit is composed of a multi-view high-definition industrial camera array arranged on both sides of the track and above the maintenance pit at the bottom. A total of 12 CMOS sensor cameras with a resolution of 5120x3840 pixels and a frame rate of 60fps are configured, covering top view, left / right side oblique view, and high-angle imaging view, ensuring no dead angle observation coverage of key parts such as bogie frame weld area, brake disc friction surface, and axle box connecting bolt group. Each camera is equipped with an adjustable focal length lens (focal length range 8mm-25mm), and the internal parameter matrix and distortion coefficient are pre-calibrated according to the actual installation height and target distance. Among them, two upward-looking cameras located directly below the track are equipped with linear polarization filters, with their transmission axis perpendicular to the vibration plane of the reflected light from the metal surface, to weaken the local overexposure artifacts caused by mirror reflection of the brake disc. All cameras use a hard trigger synchronization mechanism, and the trigger signal comes from the precise position pulse output by the train positioning system. When the on-board transponder passes through the RFID positioning tag embedded in the track, the control system determines that the center of the current train carriage has entered the best imaging area, and then sends a synchronous exposure instruction to all cameras, ensuring that the multi-view images are strictly aligned in the time dimension, and adding a time stamp accurate to the millisecond level and corresponding train carriage number information to build a time and space consistent data acquisition reference.
[0011] The original image sequence obtained by collection is transmitted to the edge computing node deployed in the local cabinet of the station through gigabit Ethernet. The node is equipped with a high-performance GPU server (configured with dual NVIDIA A100 80GB graphics cards), a solid-state storage array, and a real-time operating system, and undertakes the whole-link processing tasks from image reception, preprocessing to model inference. After the image preprocessing module is started, it first calls the distortion correction subroutine developed based on OpenCV-Python, uses the pre-calibrated camera internal participation five-order radial distortion coefficient model, and uses the bilinear interpolation method to perform pixel-level remapping on each frame of image, eliminating the barrel distortion caused by the wide-angle lens. Then, in the illumination normalization stage, an adaptive illumination compensation algorithm based on the Retinex theory is introduced, which decomposes the input image into an illumination component and a reflection component, i.e. The system uses multi-scale Gaussian filtering to estimate the illumination field: convolves the original image with Gaussian kernels of σ=30, 60, and 90 in turn to generate smooth images at three scales, and takes the maximum value as the final illumination estimation ; then reconstructs the reflection component by pixel-by-pixel division , preserving the true texture information of the object surface. To prevent numerical overflow and contrast enhancement, a dynamic range compression factor is set, and the image after illumination equalization is finally output , where is a small constant (e.g., 10 to the -6 power) that prevents division by zero.
[0012] After illumination correction, the image enters the noise suppression link. Due to the presence of dust, water vapor and electromagnetic interference in the vehicle bottom environment, the original image is often contaminated by a mixture of salt and pepper noise and Gaussian white noise. This module uses a cascade strategy of joint bilateral filtering and non-local mean denoising: first, use the bilateral filter in the spatial domain , gray domain to retain edge sharpness while smoothing uniform areas; then apply the non-local mean algorithm to calculate the Euclidean distance weight between the neighborhoods of each pixel under the parameter settings of search window 15x15 and similar block size 7x7, to realize global self-similarity driven denoising processing. The final output image is uniformly resampled to 2240x1680 size, color space is converted to sRGB standard, forming a standardized input data set.
[0013] Next, the system performs a region of interest (ROI) cropping operation. To reduce the computational load of subsequent large model inference and focus on the key area, this embodiment uses a ROI extraction method based on superpixel segmentation. Specifically, the SLIC (Simple Linear Iterative Clustering) algorithm is used to over-segment the preprocessed image, with the number of superpixels K=300, the compactness parameter m=10, and the iteration convergence threshold 10 to the -5 power, generating a set of local regions with similar boundaries and colors. Each superpixel block contains an average of about 12560 pixels, with a clear spatial coordinate range and mean color feature. The system loads a pre-modeled three-dimensional geometric projection template library - a CAD model of key components for different vehicle types (such as A-type vehicles and B-type vehicles), and generates its two-dimensional contour mask under each camera perspective through perspective projection transformation. The mask adjusts the train's entry posture: using the wheel center coordinates and vehicle axis direction detected in the previous frame, the expected projection position of the current frame is dynamically updated, compensating for lateral shifts within a range of ±15 cm and posture rotations of ±3°. After intersection operation of the SLIC segmentation result and the dynamic mask, superpixel blocks that fall completely or partially within target regions such as bogie frames, brake disc outer edges, and coupling flanges are selected and merged to generate irregular polygonal ROIs. Only the region image blocks are subjected to subsequent high-cost visual large model inference, and the remaining background regions are directly discarded, reducing the effective processing area of a single image by about 62% and significantly reducing computational resource consumption.
[0014] The visual large model inference engine receives the cropped ROI image block and performs deep semantic analysis. In this embodiment, a ViT-Adapter structure is used as the core backbone network. The backbone is a large-scale pre-training model based on the Vision Transformer (ViT-L / 16) architecture, which has completed general visual representation learning on the ImageNet-21k and MVTec-AD industrial defect data sets. The model input size is adapted to 224x224, and the ROI image is divided into a 14x14 image block sequence. Each block is projected linearly and superimposed with position coding, and then fed into a 24-layer Transformer encoder. To enhance the model's sensitivity to the local texture characteristics of the subway components, two pluggable adapter layers are embedded between the 12th and 18th layers. Each adapter layer includes a lightweight convolution projection module and a channel attention gate unit: the convolution module is composed of 1x1 convolution kernels, and the channel number is first compressed to 1 / 8 of the original dimension (i.e., 384→48), then activated by GELU, and finally restored to the original dimension by 1x1 convolution, forming a bottleneck structure. The channel attention unit is designed based on the SE (Squeeze-and-Excitation) mechanism. By applying two fully connected layers (dimension reduction ratio r=16) to the feature vector after global average pooling, a channel weight factor is generated, which is multiplied with the original feature by channel to achieve key region response enhancement and irrelevant background activation suppression. The entire inference process freezes the parameters of the first 20 layers of the backbone network, and only the learnable parameters of the top 4 layers of the Transformer block, the classification head, and the two adapter layers are fine-tuned to prevent catastrophic forgetting and accelerate convergence.
[0015] During the forward propagation of the model, the feature vector corresponding to the [CLS] token output by the last Transformer layer is extracted, with a dimension of 1024, and spatial position information is fused through global average pooling to generate a fixed-length global feature representation . This feature vector contains the overall health status semantic information of the component in the image, including surface integrity, structural symmetry, assembly consistency, and other high-level attributes. For the same physical component (such as a bogie side beam weld), the system continuously collects 8 to 16 valid images during the train passing, corresponding to a time span of about 2.1 to 4.3 seconds (based on a train speed of 3 km / h), which is sufficient to cover at least one complete wheelset rotation period (typical rotation speed about 0.7 Hz), thereby avoiding instantaneous misjudgment caused by periodic occlusion or posture changes of the component.
[0016] The feature vectors of the above continuous frames are fed into the time series state modeling module for dynamic aggregation. This module uses a bidirectional long short-term memory network (Bi-LSTM) to encode the feature sequence within a sliding time window. Let the input sequence be , length n = 12, input dimension d = 512 at each time step. Bi-LSTM contains two independent branches of forward and reverse LSTM, each with independent forget gate, input gate, output gate and cell state update mechanism. The forward branch processes the sequence in chronological order, capturing historical context dependence; the reverse branch processes in reverse order, capturing future information feedback. The hidden layer dimension of each LSTM unit is set to 512, using weight matrices W_ii, W_hi, W_if, W_io (input-gate, hidden-gate, forget-gate, output-gate connection respectively) and bias terms b_i, b_f, b_o, b_c internally, and the activation function is selected in the form of tanh and sigmoid combination. To improve training stability, L2 norm clipping is applied to all gradient updates, with a threshold of 1.0; at the same time, the dropout mechanism is introduced, randomly shielding 30% of the neuron connections at the output end of the hidden layer to prevent overfitting. Finally, the hidden state h_t output at each time step is obtained by concatenating the forward and reverse hidden variables, with a dimension of 1024, representing the state representation of the integrated pre and post context awareness capability.
[0017] The fault discrimination decision unit receives the fusion representation h_t and executes the dual-path output mechanism. The first path is a multi-class classification task, which is mapped to a predefined 6-class fault type output space through a two-layer fully connected network (512-dimensional in the middle layer, ReLU activation): normal, crack, deformation, defect, loose, foreign object attachment. The Softmax function generates the probability distribution p(c|h_t) of each class, and the class corresponding to the maximum probability is selected as the preliminary diagnosis result. The second path performs abnormal score calculation, using a probability modeling method based on kernel density estimation (KDE) to evaluate the degree of deviation of the current sample from the normal mode. In the offline training stage, the system uses no less than 5000 images of subway components labeled as "normal" state to extract features and construct a normal class cluster feature database. On this basis, a Gaussian kernel function where, the normalized feature difference vector, the vector the square of the Euclidean norm of the vector the normalization constant of the Gaussian kernel, ensuring the integral equal to 1, the main body of the Gaussian function, the farther the distance, the smaller the weight, and the bandwidth parameter h is determined by the Silverman empirical rule to establish the probability density function:
[0018] where, the probability density estimate (i.e. the likelihood value) of the input feature on the "normal" class, reflecting whether the feature belongs to the normal distribution; the higher the value, the more "normal" it is, Feature vector of current sample to be detected (e.g. embedding vector extracted by CNN), dimension d, Feature vector of the first “normal” class sample, total samples (from offline training set), Number of normal class samples (≥ 5000 in this paper), Bandwidth parameter of kernel function (bandwidth), controls the degree of smoothing; too large will lead to underfitting, too small will lead to overfitting, Dimension of feature vector, Kernel function, here Gaussian kernel, Normalize the feature difference to unit scale, for comparison of different dimensions, in online inference, substitute the current feature into the above model to calculate its likelihood value under the normal distribution . If the value is lower than the 5th percentile of the historical normal sample distribution, it is determined to be a potential anomaly, and a continuous anomaly index A ∈ [0, 1] is generated, where , to quantify the risk level.
[0019] When any path output exceeds the preset safety threshold, the system triggers the alarm logic. The judgment condition is: if the classification path output is not “normal” class and the confidence is > 0.85, or the anomaly index A > 0.95, a structured alarm message is immediately generated. The message contains the following fields: fault location (accurate to car number + bogie number + component area code), suspected type (from classification head output), confidence value (the highest score of the two paths), occurrence time (UTC millisecond timestamp), associated image storage path (NFS network file system absolute address), feature vector hash value (SHA-256 encoding) and device operation log index number. The message is pushed to the central operation and maintenance platform through the RESTful API interface, supports direct docking with the asset management system (EAM), automatically creates a work order and allocates a maintenance team. At the same time, the system automatically saves the original image, preprocessing intermediate results, feature heat map and time series evolution curve to the local SSD array, with a retention period of not less than 180 days, for post-tracing analysis and model iteration optimization.
[0020] To continuously improve the system's recognition ability, the embodiment enables an online incremental update mechanism. After reviewing the alarm events, the operation and maintenance personnel can submit samples (including images and labeled tags) confirmed as new fault types to the negative sample library. The system performs incremental training procedures once a week: first, the new samples are data-augmented (including random rotation ±15°, brightness disturbance ±20%, and adding synthetic noise) to expand to 5 times the original number; then the visual large model is re-tuned using a curriculum learning strategy - in the initial stage, use public defect datasets (such as DAGM, NEU Surface Defect) to consolidate general features, and then gradually introduce subway-specific data, train progressively according to fault difficulty classification (Level 1: obvious cracks; Level 2: small peeling; Level 3: gradual wear), with each level training for at least 20 epochs and a batch size of 32. Gradient clipping (threshold 1.0) and weight decay (λ = 5e-4) constraints are applied during training, the optimizer is AdamW, the initial learning rate is 3e-5, and the cosine annealing scheduling strategy is used. After retraining, the new model is automatically put into operation to replace the old version after functional verification testing (accuracy improvement ≥0.5% and FPR ≤1.2%) to achieve closed-loop optimization.
[0021] According to the test, in the deployment environment of the embodiment, the system can complete all key component detection in 42.7 seconds per single vehicle scan, and the end-to-end response delay is controlled within 283 milliseconds, meeting the real-time requirements. Compared with the traditional small model scheme based on ResNet-50, the average detection accuracy of the system on the test set is improved by 19.7 percentage points, reaching 98.3%, the missed detection rate is reduced to 0.76%, and the false positive rate is reduced by 42%. Especially in identifying early fatigue cracks with a width of less than 0.1mm, it shows a significant advantage, successfully warning multiple potential fracture accidents, effectively ensuring train operation safety.
[0022] Embodiment Two In another application scenario, a certain subway company B line faces the complex operation and maintenance challenge of cross-line and mixed running of multiple vehicle types, its vehicle fleet includes 6 trains produced by different manufacturers (model covers A, C and linear motor driven vehicles), the key component dimensions, installation positions and material reflection characteristics of each vehicle type are significantly different, making it difficult for traditional customized detection systems to be reused, and each time the vehicle type is changed, the camera needs to be re-deployed, the parameters need to be re-calibrated, and the model needs to be rebuilt, the deployment period is as long as two weeks or more, which seriously affects the operation efficiency. Therefore, the embodiment proposes a variant structure of an AI visual large model fault detection system with strong compatibility and rapid adaptation ability, focusing on solving the cross-platform deployment problem.
[0023] The core innovation of the embodiment is to build a decoupled system architecture, which completely separates the data input interface from the core algorithm module. The data acquisition unit still uses a multi-view industrial camera array, but no longer requires uniform models or resolutions. The cameras deployed on site come from three different manufacturers, with resolutions ranging from 1920x1080 to 4096x3072, frame rates between 30fps and 120fps, and interface types including GigE Vision, USB3 Vision, and Camera Link. To achieve plug-and-play access of heterogeneous devices, the system designs a standardized data abstraction layer (DAL), through which all cameras are registered and accessed. The DAL has built-in device driver adapter components that automatically identify camera models and load corresponding SDKs, and uniformly convert them into a general image stream protocol (GISP) format output. The GISP specifies that image metadata must include: device ID, original resolution, color space, exposure time, gain value, gamma correction parameters, lens focal length, installation angle, geographic coordinate offset, and 32 other fields. Regardless of the underlying hardware, the upper module only interacts with the GISP interface, achieving complete decoupling of the physical layer and the logical layer.
[0024] The image preprocessing module adjusts the processing flow accordingly. Due to the varying resolutions of the input images, they cannot be directly cropped to a fixed size. The embodiment introduces an adaptive scaling mechanism based on image content awareness: first, according to the "installation angle" and "target component" fields in GISP, query the built-in view-component mapping table to determine the observation view category to which the current image belongs (such as "side-down 45° brake disc view"). Then call a lightweight CNN classifier (MobileNetV3-small structure) to perform coarse classification on the image content and determine whether it matches the expected imaging mode. If it matches, calculate the optimal scaling ratio s based on the standard reference size of the view (such as the brake disc view should present a circular area with a diameter of about 60% of the image width), so that the target component is close to the standard size 2240x1680 after scaling. Lanczos resampling algorithm is used for scaling to preserve high-frequency details. For polarized camera images, an additional depolarization fusion step is performed: if the same physical area is captured by two cameras with orthogonal polarization directions, extract their Stokes vector components , calculate the depolarization intensity image , effectively remove the specular reflection component and highlight the surface texture difference.
[0025] The ROI extraction mechanism is also reconstructed. Due to the large difference in geometric parameters of components of different vehicle models, it is not possible to rely on a single three-dimensional projection template. This embodiment adopts a method combining dynamic template matching and active learning. The system maintains a vehicle model-component template library, each record containing: vehicle model code, component name, CAD contour, key point coordinates, and allowed deformation range (±8% size float). When a new vehicle model first enters the station, the system starts the template initialization process: using the image sequence of the first 5 trains, perform unsupervised clustering (DBSCAN algorithm) to extract frequently appearing stable shape patterns, combined with a small number of seed samples labeled by operation and maintenance personnel (not less than 20 for each component), train a lightweight Mask R-CNN instance segmentation model for automatic labeling of the key component area of this vehicle model. The generated initial mask is stored in the template library and assigned a "to be verified" status. In subsequent train operation, the system continuously collects prediction results and manual correction feedback, when the cumulative correction times are less than the threshold (<3 times / 10 trips) and the IOU is stable above 0.92, automatically upgrade to "official" template, realize the ability of autonomous modeling without manual configuration.
[0026] The visual large model inference engine enables domain adaptive transfer mechanism in this embodiment. Although the backbone network still uses the ViT-Adapter structure, the design of the adaptation layer is more flexible. For different vehicle models, the system assigns each vehicle model a separate adapter parameter set, forming a multi-task learning architecture of "shared backbone, branch-specific". In the training stage, the federal fine-tuning strategy is adopted: first, use the aggregated normal samples on the central server to jointly optimize the backbone network; then, the updated weights are distributed to each site edge node, and each node uses the local vehicle model exclusive data to independently fine-tune its corresponding adapter module. The internal structure of the adapter remains consistent (1x1 convolution + SE attention), but the learnable parameters are stored independently. During inference, the system automatically loads the corresponding adapter weights according to the "vehicle model code" field in GISP, realizing seamless switching.
[0027] The timing state modeling module optimizes the sliding window strategy for mixed running scenarios. Due to the difference in speed (2.5 km / h~4.8 km / h) of different vehicle types, a fixed frame number window may result in insufficient coverage period. This embodiment uses a dynamic window mechanism based on physical displacement: the system calculates the minimum required time span T_min=2 / f according to the real-time speed v of the train (provided by the positioning system) and the component rotation frequency f (obtained by looking up the table), ensuring that the window length contains at least two complete rotation periods. For example, for a wheelset with a rotation speed of 0.6 Hz, T_min≈3.33 seconds; if the current vehicle speed is 3.6 km / h (i.e. 1 m / s), the corresponding spatial span is 3.33 meters, and according to the camera field of view width of 2.8 meters, at least 12 frames of images need to be collected (frame interval 83 ms). The system dynamically adjusts the Bi-LSTM input sequence length accordingly, with a range of 10~18 frames, ensuring the sufficiency of timing modeling.
[0028] The fault discrimination decision unit adds a cross-vehicle type normalized scoring mechanism. Due to the distribution of normal state characteristics of different vehicle types, direct comparison of abnormal indexes may lead to misjudgment. This embodiment introduces a Z-score standardization layer: for each vehicle-component combination, the mean μ and covariance matrix Σ of its normal sample features are maintained, and the Mahalanobis distance is calculated:
[0029] and the is mapped to a unified risk score S∈[0,1], the formula is where is a scale factor (empirical value =0.05). This score eliminates the statistical differences between vehicle types, allowing the alarm threshold to be set globally.
[0030] This embodiment also enhances the system's fault tolerance and abnormal handling capability. When a camera is disconnected or the image is blurred, the system starts a redundant compensation mechanism: using other perspective images to complete the missing view through three-dimensional reconstruction. Specifically, a stereo matching algorithm (SGM semi-global matching) is used to generate a depth map, which is combined with the known camera pose matrix P to project into the world coordinate system and reconstruct the component three-dimensional point cloud. Then project it into the perspective of the faulty camera to generate simulated images for subsequent reasoning, ensuring that the detection process does not be interrupted. Experiments show that in the case of single perspective failure, the system can still maintain a detection accuracy of 91.4%.
[0031] After six months of actual operation verification on Line B, the system successfully supports the automatic identification of six vehicle types, with an average deployment period shortened to 1.8 days (mainly physical installation time), which is 7.6 times faster than the traditional method. The model's cross-vehicle type generalization error is reduced to 5.3%, and the alarm consistency reaches 94.7%, significantly better than the control system without using the decoupling architecture.
[0032] Example Three In the third embodiment, the focus is on the robustness enhancement requirement under extremely harsh imaging conditions. A certain subway line passes through the coastal area, affected by humid salt fog, frequent rainfall and tunnel seepage, the surface of the car bottom parts is often covered with a composite pollution layer of oil stains, mud, water film, etc., which seriously obscures the original texture, leading to a sharp decline in the performance of conventional visual detection methods. In addition, the lighting conditions during the night maintenance period are unstable, and strong glare and deep shadows are easy to produce, further exacerbating the identification difficulty. This embodiment proposes an enhanced detection architecture that combines multi-spectral perception and self-supervised repair mechanisms, designed specifically to cope with low-visibility complex environments.
[0033] In this embodiment, the data acquisition unit is upgraded in hardware. On the basis of the original visible light camera array, a multi-spectral imaging subsystem is added. The specific configuration is as follows: 4 groups of multi-spectral cameras are added on both sides of the track, each group containing 6 channels with wavelengths of 450nm (blue), 550nm (green), 650nm (red), 750nm (near-infrared NIR), 850nm (short-wave infrared SWIR) and 1050nm (mid-wave infrared MWIR). Each channel uses a filter wheel switching design to ensure that full-spectrum imaging is completed at the same spatial location within a very short time (<20ms), with time synchronization accuracy guaranteed by the FPGA controller. At the same time, two polarized visible light cameras are retained, forming a "multi-spectral + polarization" composite perception array. All images are attached with accurate wavelength labels and polarization angle information, forming a six-dimensional data cube I(x,y,λ,p,t), where λ represents the wavelength dimension, and p represents the polarization angle (0°, 45°, 90°, 135°).
[0034] The image preprocessing module introduces a cross-modal alignment and fusion mechanism. Due to the differences in imaging resolution and field of view between different wavebands, spatial registration needs to be performed first. The system uses a non-rigid registration algorithm based on mutual information maximization (MI-maximization) to correct the affine transformation and B-spline grid deformation of other waveband images with the visible light image as the reference modality, with registration error controlled at the sub-pixel level (RMSE <0.3px). Subsequently, polarization-spectrum joint denoising is performed: for each pixel, a response vector is constructed under different wavelengths and polarization angles, and a tensor low-rank decomposition (Tucker decomposition) model is applied, assuming that the pure signal has a low-rank structure and the noise is a sparse disturbance, solving the optimization problem:
[0035] where is the low-rank tensor (representing the true surface reflection characteristics), is the sparse error tensor (representing noise and interference), is the tensor core norm, is the L1 norm, Balancing parameters (empirical values =0.15). The abnormal response affected by pollution is effectively separated by iterative solution by the alternating direction multiplier method (ADMM).
[0036] To further restore the hidden surface details, the embodiment develops an image restoration module based on a self-supervised generative adversarial network. The module does not need to train paired clean-pollution images, but uses a large number of single-mode pollution images to construct a self-supervised task. The network structure adopts a U-Net generator G and a PatchGAN discriminator D. The training strategy is: the "stable area" (such as the frame fixed area without moving parts) is extracted from the historical images of the same vehicle passing through multiple times, and it is assumed that its long-term appearance is unchanged; when the area is covered by mud in some imaging, it is taken as input x, and G(x) is expected to restore the un-polluted state . The loss function consists of three parts: pixel-level L1 loss, perception loss (based on VGG16 high-level features), and adversarial loss. In particular, a context attention mechanism is introduced to enable the generator to "borrow" texture information from the surrounding clear area to fill the occluded area. After training, the module can convert the pollution image into a "virtual clean" image in real time for subsequent large model inference.
[0037] The visual large model inference engine is correspondingly expanded to a multi-modal input architecture. The original ViT-Adapter structure is modified to support multi-channel input: the repaired visible light image, NIR image, and polarization difference image ) are spliced along the channel dimension to form a 12-channel input tensor. To adapt to the new modalities, the first layer convolution kernel is expanded from the standard 3x3x3 to 3x3x12, and the weight initialization uses the Xavier method. The channel attention mechanism in the adaptation layer is upgraded to a cross-modal attention gate, which calculates the correlation weight matrix between modalities to dynamically adjust the importance of features in different wavebands. For example, in an oil pollution environment, the SWIR waveband has stronger penetration ability, and the system automatically increases its feature weight ratio to more than 68%.
[0038] The time series state modeling module introduces a multi-modal time series alignment mechanism. Since different sensors may have micro-latency, the system uses the dynamic time warping (DTW) algorithm to nonlinearly align the feature sequences of each modality, ensuring that the same physical event is at the same time position in all channels. The Bi-LSTM input is changed to a multi-branch structure, and each modality is encoded separately before cross-modal attention fusion to generate a unified state representation.
[0039] The fault discrimination decision unit adds a multi-evidence fusion decision mechanism. In addition to traditional classification and abnormal score, an auxiliary criterion based on spectral fingerprint analysis is added: for suspicious areas, extract their reflectivity curves in 6 bands, match with the standard material library (steel, aluminum alloy, rubber, composite material), if abnormal absorption peak (such as rust absorption enhancement at 750 nm), increase the alarm priority. Finally, three indicators are integrated: classification confidence, abnormal index, spectral deviation, D-S evidence theory is used for fusion decision, and the final determination result is output.
[0040] In the continuous three-month rainy season test, the system still maintains an average detection accuracy of 93.2% in the face of poor environment with an average visibility of less than 15 meters, which is 21.8 percentage points higher than the control system using only visible light. Especially in identifying early corrosion points covered with oil sludge, the early warning capability is extended to 7-10 days, fully verifying the effectiveness and engineering practicability of the multispectral fusion and self-supervised repair mechanism.
[0041] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content without departing from the scope of the present application, and any equivalent embodiments with equivalent changes are still within the scope of the present application.
Claims
1. A method for fault detection of key components in subway systems using AI visual large-scale models, characterized in that, Includes the following steps: By deploying a multi-view industrial camera array next to the track, the original visible light image sequence of key components is simultaneously acquired during the low-speed entry into the station or return to the depot of the train, and a timestamp and carriage number information are added to each frame of the image. The acquired raw images are input into the image preprocessing module for distortion correction, illumination normalization, noise suppression, and region of interest cropping to generate a standard image set with uniform size, consistent color space, and reduced background interference. The standard image set is input into the visual large model inference engine pre-trained with massive industrial defect data. Multi-level spatial semantic features are extracted using a deep neural network based on the Transformer architecture, and a fixed-dimensional global feature vector is obtained through global pooling. The global feature vectors extracted from the same physical component in a continuous time period are arranged in chronological order to form an input sequence, which is then fed into a bidirectional long short-term memory network for context encoding. The appearance evolution patterns between consecutive frames are dynamically aggregated to generate a fused spatial temporal state representation. Based on the state representation, classification and anomaly scoring are performed in parallel. The classification path outputs a predefined fault type label, and the anomaly scoring path quantifies the degree to which the current sample deviates from the normal cluster. When the judgment result meets the alarm triggering conditions, a structured alarm message containing the fault location, suspected type, confidence value and associated image path is generated and pushed to the operation and maintenance management platform. At the same time, intermediate data is saved for traceability analysis and model optimization.
2. The method according to claim 1, characterized in that, The large-scale visual model inference engine adopts the ViT-Adapter structure, whose backbone network is a large-scale Vision Transformer model with frozen parameters. Pluggable adaptation layers are embedded in specific network layers to achieve domain transfer adaptation. The adaptation layer includes a lightweight convolutional projection module and a channel attention gating unit. The convolutional projection module reduces computational overhead through a bottleneck structure of dimensionality reduction-activation-upgrading, and the channel attention gating unit generates channel weight factors based on global pooling results and performs weighted modulation on the original features.
3. The method according to claim 1, characterized in that, The illumination normalization operation in the image preprocessing module adopts an adaptive illumination compensation algorithm based on Retinex theory. Specifically, it includes: applying multi-scale Gaussian filtering to the input image to generate several smooth images, taking the maximum value of the corresponding pixel at each scale as the illumination component estimate; separating the reflection component from the original image through pixel-by-pixel division operation, and reconstructing the output image by combining the dynamic range compression factor.
4. The method according to claim 1, characterized in that, The region of interest (ROI) cropping operation in the image preprocessing module employs a selective processing strategy based on superpixel segmentation, including: using the SLIC algorithm to over-segment the preprocessed image to generate a uniform and compact set of local regions; loading a 3D geometric projection template matching the current vehicle model as a priori mask, and compensating for vehicle attitude offset based on the detection results of the preceding frame through a dynamic update mechanism; performing intersection operations between the superpixel blocks and the updated mask, filtering out regions located within the contour range of the bogie, brake disc, or coupling, and merging them to generate irregular polygonal ROIs, and performing subsequent high-cost model inference only on the ROI.
5. The method according to claim 1, characterized in that, The input sequence length of the bidirectional long short-term memory network is set to a time window covering at least one complete wheel pair rotation cycle. The input at each time step is a high-level semantic feature vector extracted from the same component region at different time points. The network processes the sequence along the forward and reverse directions respectively, capturing historical dependencies and future context information, and finally splices the hidden states in the two directions to form a state output containing context perception capabilities.
6. The method according to claim 1, characterized in that, The anomaly score calculation adopts a probabilistic modeling method based on kernel density estimation, which specifically includes: in the offline stage, using images of subway components labeled as normal to extract features and construct a normal cluster database; in the online stage, using a Gaussian kernel function and a bandwidth parameter determined by the Silverman rule to establish a probability density function, and substituting the current features into the model to calculate its likelihood value.
7. The method according to claim 1, characterized in that, The multi-view industrial camera array is arranged in a three-dimensional distribution, covering top-down, side-down, and low-angle observation views. At least some of the cameras are equipped with polarizing filters to reduce the effect of specular reflection on the surface of metal parts. All cameras adopt a hard-triggered synchronization mechanism, and the trigger signal comes from the position pulse output by the train positioning system.
8. A detection system for a subway key component fault detection method applied to the AI visual large model described in any one of claims 1-7, characterized in that, include: The system comprises the following components: a data acquisition unit, which synchronously acquires raw visual image sequences of key components via a multi-view industrial camera array alongside the track when a train passes, and records the corresponding timestamps and carriage numbers; an image preprocessing module, connected to the data acquisition unit, sequentially performs distortion correction, illumination equalization, noise filtering, and ROI extraction operations on the raw images to generate standardized input data; a large-scale visual model inference engine, which receives the output from the image preprocessing module and uses a large-scale convolutional neural network based on the Transformer architecture to perform deep semantic analysis on the images, extracting multi-level spatial features and pooling them into global feature vectors; a temporal state modeling module, coupled with the large-scale visual model inference engine, uses a bidirectional LSTM network to dynamically aggregate feature vector sequences of the same component region across consecutive frames to capture trend patterns of appearance changes; a fault discrimination decision unit, which integrates a classification head and anomaly scoring mechanism, outputs component health status labels and confidence scores based on the fused spatial-temporal features; and an alarm and feedback interface, which pushes the judgment results to the operation and maintenance management platform in real time and supports automatic optimization of the discrimination threshold parameters based on historical cases.
9. The system according to claim 8, characterized in that, A standardized data abstraction layer is set between the image preprocessing module and the core algorithm module. The industrial cameras in the data acquisition unit come from different manufacturers and have heterogeneous resolutions, frame rates and interface types. The standardized data abstraction layer has a built-in device driver adapter component, which can automatically identify the camera model and load the corresponding SDK, and convert various image streams into a common image stream protocol format for output.
10. The system according to claim 8, characterized in that, The data acquisition unit further includes a multispectral imaging subsystem, which is equipped with cameras with multiple wavelength channels to complete full-spectrum imaging at the same spatial location in a very short time, and adds wavelength labels and polarization angle information; the image preprocessing module also includes a cross-modal alignment and fusion unit, which uses a non-rigid registration algorithm based on maximizing mutual information to perform spatial correction on images of different bands, and applies a tensor low-rank decomposition model to separate the real surface reflection characteristics and noise interference; the visual large model inference engine is extended to support a multi-channel input architecture, which stitches the repaired visible light, near-infrared and polarization difference images along the channel dimension and inputs them into the network, and dynamically adjusts the feature weights of each band through a cross-modal attention mechanism.
Citation Information
Cited By
PVC sizing material defect detection method based on AI visual identification and intelligent sorting system
CN121661057A
Abnormal signal identification method and system of electric control cabinet
CN121682637A
Elevator door system and well key component inspection system based on depth vision
CN121757699A
Elevator door system based on deep vision and hoistway key component inspection system
CN121757699B