Intelligent detection method for catenary concealment defect based on multi-modal large model
By fusion of images, point clouds and sound wave data by multimodal large models, the accuracy and intelligence of contact network concealment defect detection are solved, and efficient and intelligent defect recognition and positioning are achieved, and it is suitable for various rail transit scenarios such as high-speed railways.
Patent Information
- Application Number
- CN202510460184.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-29
AI Technical Summary
The existing contact network defect detection methods are difficult to identify early, small, and concealed defects, and the level of intelligence is insufficient, resulting in potential safety hazards.
Multimodal large-scale model is used to integrate images, laser point clouds, infrared heat maps and structural acoustic data, and feature fusion and deep semantic modeling are performed through the Transformer architecture, combined with expert knowledge feedback and continuous learning mechanisms to achieve high-precision identification and positioning of hidden defects in the contact network.
It significantly improves the recognition rate of hidden defects, enhances the robustness and adaptability of the model, supports continuous learning and expert feedback, and is suitable for a variety of rail transit scenarios.
Smart Images

Figure CN120387067A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of high-speed railway catenary, and particularly relates to an intelligent detection method for concealed defects of catenary based on a multi-modal large model. Background Art
[0002] As a key component of the rail transit power supply system, the structural health status of the catenary is directly related to the safety of train operation and the reliability of the power supply system. During long-term operation, the catenary is susceptible to various factors such as environmental corrosion, mechanical fatigue, and climate change, resulting in various types of structural or functional defects, such as wear of the contact wire, loosening of the suspension structure, cracking of insulators, and corrosion of fittings. In particular, some early, tiny, non-obvious concealed defects, such as microcracks, corrosion initiation points, fatigue micro-deformations, and insulation aging, if not detected in time, may evolve into sudden wire breaks, discharges, or equipment failures, seriously threatening the safety of train operation.
[0003] Currently, the defect detection of the catenary mainly relies on the following methods: (1) Manual inspection: By visually inspecting manually and observing with a rangefinder or telescope. Although it has strong flexibility, it has low efficiency, strong subjectivity, and is difficult to detect tiny or concealed defects. (2) Single-sensor detection: For example, using a single sensor such as a laser rangefinder, an infrared thermal imager, or ultrasonic waves to monitor specific parameters of the catenary. However, these methods have problems such as limited information dimension, local detection ability, and poor environmental robustness, and it is difficult to comprehensively identify multiple types of defects under complex working conditions. (3) Traditional machine vision detection: Some railway lines have deployed image recognition or machine learning methods for anomaly detection, but they are mostly limited to visible light images and have limited adaptability to background interference, light changes, and few-sample defects. The above traditional methods have deficiencies in concealed defect recognition, such as limited detection dimension, poor robustness, low intelligence level, and lack of continuous learning ability.
[0004] With the development of artificial intelligence and large model technology, multi-modal large models (MLLMs) have demonstrated powerful feature fusion and semantic understanding capabilities in the fields of computer vision, natural language processing, and industrial inspection. By introducing multi-modal large models, different modalities of information such as images, point clouds, infrared, and sounds are deeply fused, which can not only improve the recognition rate of concealed defects but also enhance the generalization ability and continuous learning ability of the model. Therefore, there is an urgent need for a new method for detecting concealed defects of catenary based on multi-modal perception and large model intelligent analysis to improve the automation, intelligence, and refinement level of catenary inspection and ensure the long-term safe and stable operation of the rail transit system. Summary of the Invention
[0005] The present invention relates to an intelligent detection method for hidden defects of catenary based on a multimodal large model, aiming to solve the problems existing in the existing catenary defect detection methods, such as low recognition accuracy, difficulty in detecting hidden defects, and low intelligent level. This method realizes high-precision recognition and automatic annotation of early, subtle, and hidden structural defects in the catenary system by integrating multimodal information such as images, lidar point clouds, infrared thermal maps, and acoustic signals, and by leveraging the semantic understanding and cross-modal fusion capabilities of deep learning large models.
[0006] To achieve the above object, the technical solution adopted by the present invention is: an intelligent detection method for hidden defects of catenary based on a multimodal large model, including the steps:
[0007] Through a variety of sensors such as high-resolution industrial cameras, lidars, infrared thermal imagers, and sound sensors installed on work vehicles or inspection equipment, synchronously collect visible light images, three-dimensional point clouds, infrared spectra, and structural sound signals of the catenary to form a multimodal original data set; perform spatio-temporal registration, denoising, coordinate unification, and feature normalization processing on the collected multimodal data to construct a standardized multimodal input vector; input the preprocessed multimodal data into a pre-trained multimodal large model, and use the attention mechanism (Transformer architecture) to perform inter-modal semantic feature fusion to extract deep semantic features related to hidden defects; classify and locate the fused features through the downstream detection head of the multimodal large model to identify hidden defects including crack initiation points, micro-corrosion, micro-damage of insulators, contact surface corrosion, etc.; compare the detection results with manually reviewed data, construct an expert knowledge feedback module, and fine-tune and enhance the accuracy of the large model through a continuous learning mechanism; visualize the detection results in the form of pictures and texts and upload them to the cloud operation and maintenance platform to support remote diagnosis by experts and work order dispatching.
[0008] Furthermore, the multimodal large model is an improved multimodal Transformer model jointly trained based on images, texts, and point clouds, and integrates the BERT encoder and the Vision Encoder module to enhance the context understanding ability.
[0009] Furthermore, the multimodal input includes at least the following four modalities: visible light images, three-dimensional lidar point clouds, infrared spectra, and structural sound spectrograms.
[0010] Furthermore, the defect types include but are not limited to: fine cracks in insulators, local corrosion or oxidation of contact wires, fatigue aging of fittings, micro-displacement of poles, relaxation of suspension structures, etc.
[0011] Furthermore, the data preprocessing module includes the following sub-modules: point cloud denoising and voxel filtering, image color normalization and enhancement, infrared spectrum temperature difference contrast stretching, acoustic spectrum Fourier transform and envelope analysis.
[0012] Furthermore, the defect detection module adopts a multi-task learning framework and outputs the defect type, location coordinates, confidence level, and recommended handling level simultaneously.
[0013] Furthermore, the expert knowledge feedback module uses a supervised contrastive learning mechanism to incrementally optimize the model and supports data closed-loop retraining.
[0014] Furthermore, the method is deployed between the edge computing unit and the cloud platform. The edge side performs preprocessing and rapid discrimination, and the cloud is responsible for model inference optimization and expert scheduling.
[0015] The beneficial effects of adopting this technical solution are mainly reflected in the following aspects:
[0016] 1. Strong detection ability: By fusing multi-source information and introducing large model intelligent analysis, the recognition rate of concealed defects (such as micro-cracks, early corrosion, etc.) is significantly improved;
[0017] 2. High robustness: It has good anti-interference ability and can adapt to complex lighting, background, and operating environments;
[0018] 3. High level of intelligence: It supports continuous learning and expert feedback mechanisms, the model can be continuously optimized, and the detection ability is enhanced over time;
[0019] 4. Wide range of applications: It is applicable to various track power supply system scenarios such as high-speed railways, regular-speed railways, and urban rail transit.
[0020] In summary, the present invention combines cutting-edge artificial intelligence technology with the engineering requirements of the track power supply system, providing an efficient, intelligent, and accurate solution for the detection of concealed defects in catenaries, and has broad engineering application value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic flow diagram of an intelligent detection method for concealed defects in catenaries based on a multi-modal large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS [[ID=3,6]]
[0022] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.
[0023] In this embodiment, as shown in Figure 1 The present invention proposes an intelligent detection method for concealed defects in catenaries based on a multi-modal large model, including the steps:
[0024] An intelligent detection method for concealed defects in catenaries based on a multi-modal large model, characterized by including the following steps:
[0025] S100: Data Acquisition: Synchronously acquire visible light images, 3D point clouds, infrared spectra, and structural sound signals of the catenary through various sensors such as high-resolution industrial cameras, lidar, infrared thermal imagers, and sound sensors deployed on operation vehicles or inspection equipment to form a multi-modal raw dataset;
[0026] S200: Data Preprocessing: Perform spatio-temporal registration, denoising, coordinate unification, and feature normalization on the acquired multi-modal data to construct a standardized multi-modal input vector;
[0027] S300: Feature Fusion and Semantic Enhancement: Input the preprocessed multi-modal data into a pre-trained multi-modal large model, and use the attention mechanism (Transformer architecture) to perform inter-modal semantic feature fusion to extract deep semantic features related to hidden defects;
[0028] S400: Defect Discrimination and Location: Classify and locate the fused features through the downstream detection head of the multi-modal large model to identify hidden defects including crack initiation points, micro-corrosion, micro-damage of insulators, contact surface corrosion, etc.;
[0029] S500: Defect Annotation and Iterative Learning: Compare the detection results with manually reviewed data, construct an expert knowledge feedback module, and fine-tune and enhance the accuracy of the large model through a continuous learning mechanism;
[0030] S600: Result Display and Remote Push: Visualize the detection results in the form of pictures and texts, upload them to the cloud operation and maintenance platform, and support expert remote diagnosis and work order dispatching.
[0031] Furthermore, the multi-modal large model is an improved multi-modal Transformer model based on joint training of image-text-point cloud, and integrates a BERT encoder and a Vision Encoder module to enhance the context understanding ability.
[0032] Furthermore, the multi-modal input includes at least the following four modalities: visible light images, 3D laser point clouds, infrared spectra, and structural sound spectrograms.
[0033] Furthermore, the defect types include but are not limited to: fine cracks in insulators, local corrosion or oxidation of catenary wires, fatigue aging of fittings, micro-displacement of poles, slack of suspension structures, etc.
[0034] Furthermore, the data preprocessing module includes the following sub-modules: point cloud denoising and voxel filtering, image color normalization and enhancement, infrared spectrum temperature difference contrast stretching, sound spectrum Fourier transform and envelope analysis. Among them:
[0035] (1) The point cloud voxel filtering is based on the following formula:
[0036]
[0037] In the formula, V k represents the point set within the k-th voxel grid, and p i is the spatial coordinate of each point within the grid, and P′ is the representative point retained after filtering.
[0038] (2) The contrast stretching of the infrared spectrum adopts a linear mapping formula:
[0039]
[0040] In the formula, I′(x,y) represents the temperature value of the pixel point in the original image, T max and T min are respectively the minimum and maximum temperature values of the current image.
[0041] (3) The acoustic spectrum envelope extraction uses short-time Fourier transform (STFT) and frequency-domain envelope calculation:
[0042]
[0043] Among them, x(n) is the sampling signal, w(n) is the window function, S(t,f) is the time-frequency energy diagram, and E(f) is the frequency-domain envelope line.
[0044] Furthermore, the defect detection module adopts a multi-task learning framework and outputs the defect type, location coordinates, confidence level, and recommended disposal level simultaneously. Its loss function is in the following combined loss form:
[0045] L = λ1·L cls + λ2·L loc + λ3·L conf + λ4·L priority
[0046] In the formula, L cls is the cross-entropy loss of the defect category; L loc is the location regression loss; L conf is the confidence loss; L priority is the disposal priority loss combined with expert rules; λ1~λ4 are the weight hyperparameters of each sub-task.
[0047] Furthermore, the expert knowledge feedback module adopts a supervised contrast learning mechanism to incrementally optimize the model, and the contrast loss function is as follows:
[0048]
[0049] In the formula, z i and z j are the embedding vectors of different modal enhanced views of the same defect sample; sim(zi , z j ) represents the cosine similarity; τ is the temperature coefficient that controls the distribution smoothness; K is the total number of contrast samples.
[0050] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent detection method for concealed defects of catenary based on multi-modal large models, characterized in that, Including the following steps: S100: Data acquisition: Synchronously collect visible light images, 3D point clouds, infrared spectra, and structural sound signals of the catenary through a variety of sensors such as high-resolution industrial cameras, lidar, infrared thermal imagers, and sound sensors deployed on operation vehicles or inspection equipment to form a multi-modal raw data set; S200: Data preprocessing: Perform spatio-temporal registration, denoising, coordinate unification, and feature normalization on the collected multi-modal data to construct a standardized multi-modal input vector; S300: Feature fusion and semantic enhancement: Input the preprocessed multi-modal data into a pre-trained multi-modal large model, and use the attention mechanism (Transformer architecture) to perform inter-modal semantic feature fusion to extract deep semantic features related to hidden defects; S400: Defect discrimination and localization: Classify and locate the fused features through the downstream detection head of the multi-modal large model to identify hidden defects including crack initiation points, micro-corrosion, micro-damage of insulators, contact surface corrosion, etc.; S500: Defect annotation and iterative learning: Compare the detection results with the manually reviewed data, construct an expert knowledge feedback module, and fine-tune and enhance the accuracy of the large model through a continuous learning mechanism; S600: Result display and remote push: Visualize the detection results in the form of pictures and texts, upload them to the cloud operation and maintenance platform, and support expert remote diagnosis and work order dispatch.
2. The intelligent detection method for concealed defects of catenary based on multi-modal large model according to claim 1, characterized in that, The multi-modal large model is an improved multi-modal Transformer model based on joint training of images, texts, and point clouds, and integrates the BERT encoder and Vision Encoder module to enhance the context understanding ability.
3. The intelligent detection method for concealed defects of catenary based on multi-modal large model according to claim 1, wherein, The multi-modal input includes at least the following four modalities: visible light images, 3D laser point clouds, infrared spectra, and structural sound spectrograms.
4. The method for intelligent detection of hidden defects of contact network based on multimodal large model according to claim 1 is characterized in that: The defect types include but are not limited to: fine cracks in insulators, local corrosion or oxidation of contact wires, fatigue aging of fittings, micro-displacement of poles, and relaxation of suspension structures.
5. The intelligent detection method for hidden defects of catenary based on multi-modal large model according to claim 1, characterized in that, The data preprocessing module includes the following sub-modules: point cloud denoising and voxel filtering, image color normalization and enhancement, infrared spectrum temperature difference contrast stretching, and acoustic spectrum Fourier transform and envelope analysis.
6. The intelligent detection method for concealed defects of catenary based on multimodal large model according to claim 1, wherein The defect detection module adopts a multi-task learning framework and outputs the defect type, position coordinates, confidence level, and recommended disposal level at the same time.
7. The intelligent detection method for hidden defects of catenary based on multi-modal large model according to claim 1, characterized in that, The expert knowledge feedback module uses a supervised contrast learning mechanism to perform incremental optimization on the model and supports data closed-loop re-training.
8. The intelligent detection method for concealed defects of catenary based on multi-modal large model according to claim 1, characterized in that, The method is deployed between the edge computing unit and the cloud platform. The edge side performs preprocessing and rapid discrimination, and the cloud is responsible for model inference optimization and expert scheduling.
9. The intelligent detection method for concealed defects of catenary based on multi-modal large model according to claim 5, characterized in that, The point cloud voxel filtering is based on the following formula: where V k represents the point set within the k-th voxel grid, p i is the spatial coordinate of each point within the grid, and P ′ is the representative point retained after filtering; The infrared spectrum contrast stretching adopts a linear mapping formula: Where, I ′ (x, y) represents the temperature value of the pixel point in the original image, T max and T min are respectively the minimum and maximum temperature values of the current image; The acoustic spectrum envelope extraction uses the short-time Fourier transform (STFT) and frequency-domain envelope calculation: Where x(n) is the sampled signal, w(n) is the window function, S(t,f) is the time-frequency energy map, and E(f) is the frequency-domain envelope line.
Citation Information
Cited By
Body armor composite interlayer defect positioning method and system based on multi-mode nondestructive testing, electronic equipment and storage medium
CN120801638A
Sliding rail quality inspection method and system based on artificial intelligence
CN121121439A