Pantograph-catenary defect detection and feature fusion method based on deep transfer learning
Through deep transfer learning and multimodal data fusion, the robustness and accuracy of bow net defect detection in traditional methods are solved, and efficient and accurate identification and evaluation of bow net defects are achieved.
Patent Information
- Application Number
- CN202510382054.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional single-modal detection methods are difficult to fully capture the complex characteristics of bow net defects, and are not robust under large environmental interference or extreme conditions. It is difficult for the existing technology to achieve effective fusion and accurate detection of multimodal data.
Using a deep transfer learning method, multimodal data (visible light images, infrared thermal imaging, and laser 3D point cloud) are collected and aligned, combined with domain adaptive feature extraction and environmental data encoding, dynamic weights are generated for feature fusion, and finally output defect categories, locations and severity by detecting the classification network.
The comprehensive capture of multi-dimensional defect characteristics is achieved, the accuracy and environmental adaptability of bow network defect detection are improved, the dependence on labeled data is reduced, and the intelligence level and operation and maintenance efficiency of bow network equipment status evaluation are improved.
Smart Images

Figure CN120372529A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of defect detection, and more specifically, to a pantograph-catenary defect detection and feature fusion method based on deep transfer learning. Background Art
[0002] Traditional single-modal detection methods are difficult to comprehensively capture the complex features of pantograph-catenary defects, while multi-modal data fusion can provide a more complete defect representation. However, facing the domain difference problem between laboratory simulation and actual line data, directly applying deep learning models is prone to overfitting.
[0003] The Chinese patent application with the authorization publication number CN105652154B discloses a catenary operation status safety monitoring and analysis system: The system at least includes a first camera, a second camera, an image correction unit, a pantograph recognition unit, a contact wire recognition unit, a model database, a geometric parameter calculation unit, and a defect recognition unit; The first camera and the second camera respectively collect video images of the pantograph from two different angles and output a first perspective image and a second perspective image; The image correction unit respectively performs perspective correction on the first perspective image and the second perspective image according to the calibration of the pantograph, so that the pantograph in the first perspective image and the second perspective image is in a left-right symmetric form; The model database is used to store pantograph models, and the pantograph models include a first perspective pantograph model corresponding to the first perspective image and a second perspective pantograph model corresponding to the second perspective image; The pantograph recognition unit respectively recognizes the pantograph in the first perspective image and the second perspective image according to the first perspective pantograph model and the second perspective pantograph model, and locates the pantograph area; The contact wire recognition unit is used to respectively recognize the straight lines suspected of being contact wires in the first perspective image and the second perspective image, and compare the pantograph areas of the first perspective image and the second perspective image at the same scale, find out the straight lines suspected of being contact wires that intersect the top plane of the pantograph, and determine this straight line as the contact wire; The geometric parameter calculation unit respectively calculates the geometric parameters of the catenary in the first perspective image and the second perspective image according to the pantograph recognized by the pantograph recognition unit and the contact wire recognized by the contact wire recognition unit, and outputs the optimal geometric parameters according to the smoothness and / or similarity characteristics; The defect recognition unit respectively and in real time detects and recognizes the defects existing in the catenary according to the pantograph recognized by the pantograph recognition unit, the contact wire recognized by the contact wire recognition unit, and / or the geometric parameters output by the geometric parameter calculation unit, including catenary defects, pantograph defects, and pantograph-catenary relationship defects; The defect recognition unit at least includes any one or a combination of more than one of a pantograph deformation defect recognition unit, a component shedding defect recognition unit, a pantograph pulling-out overlimit defect recognition unit, an arcing defect recognition unit, a high-temperature interference recognition unit, and an accidental pantograph lowering defect recognition unit. This invention can timely discover the abnormal defects existing in the pantograph-catenary, and can in real time discover pantograph-catenary relationship defects, catenary defects, pantograph defects, and other operation defects, etc.
[0004] Although the above method can meet most scenarios, through research and practical application of the above method and the existing technology, it is found that the above method and the existing technology at least have the following partial defects:
[0005] It mainly relies on binocular visible light cameras for geometric parameter measurement and defect recognition, lacking the ability to fuse multi-modal data, making it difficult to comprehensively capture the multi-dimensional features of defects, resulting in limited detection accuracy for complex defects; the system performs matching recognition based on a preset pantograph and catenary model library, and is prone to model overfitting or misjudgment in scenarios with large environmental interference; under extreme weather or high-speed operation conditions, the robustness of geometric parameter measurement and defect recognition is insufficient.
[0006] In view of this, the present invention proposes a pantograph-catenary defect detection and feature fusion method based on deep transfer learning to solve the above problems. Summary of the Invention
[0007] To overcome the above defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A pantograph-catenary defect detection and feature fusion method based on deep transfer learning, comprising the following steps:
[0008] Collect sensor data, where the sensor data includes detection data and environmental data; the detection data includes visible light images, infrared thermal images, and laser 3D point clouds of the object to be detected; the environmental data includes temperature and humidity;
[0009] Align the detection data to obtain an aligned multi-modal data set;
[0010] Perform domain adaptive feature extraction on the aligned multi-modal data set to obtain multi-modal domain-invariant feature vectors;
[0011] Encode the environmental data, generate dynamic weights through an attention module, and perform weighted fusion on the multi-modal domain-invariant feature vectors according to the dynamic weights to obtain a unified feature representation after fusion;
[0012] Use the unified feature representation after fusion as the input of the detection and classification network to obtain the probability distribution of pantograph-catenary defect categories, the regression values of position and size, and the severity score.
[0013] Further, the method for obtaining multi-modal domain-invariant feature vectors includes:
[0014] Use the aligned multi-modal data sets in the source domain and the target domain as the input of a hybrid backbone network to obtain corresponding multi-modal features, and calculate the MMD loss between the multi-modal features of the source domain and the target domain according to a preselected kernel function.
[0015] Input the source domain multi-modal features and the target domain multi-modal features into a domain classifier to obtain the probability that the input multi-modal features come from the source domain; combine the cross-entropy loss to obtain the domain loss of the domain classifier.
[0016] Add the MMD loss and the domain loss to obtain the domain adaptation loss, optimize the hybrid backbone network and the domain classifier with the minimum domain adaptation loss as the optimization goal, and use the enhanced multi-modal features output by the domain classifier when the domain adaptation loss is minimized as the multi-modal domain-invariant feature vector.
[0017] Further, the training method of the hybrid backbone network includes:
[0018] Pre-collect B groups of hybrid training data, where the hybrid training data includes the aligned multi-modal data set and multi-modal features; the aligned multi-modal data set includes the aligned visible light images, the aligned infrared thermal images, and the aligned laser 3D point clouds; the multi-modal features include visible light features, infrared features, and point cloud features.
[0019] The hybrid backbone network includes a visible light branch, an infrared branch, and a point cloud processing branch. Take the aligned visible light images as the input of the visible light branch and the corresponding visible light features as the output of the visible light branch; take the aligned infrared thermal images as the input of the infrared branch and the corresponding infrared features as the output of the infrared branch; take the aligned laser 3D point clouds as the input of the point cloud processing branch and the corresponding point cloud features as the output of the point cloud processing branch. With the goal of minimizing the error between the output multi-modal features and the actual multi-modal features, optimize the network parameters of the hybrid backbone network through a nature-inspired optimization algorithm, obtain the network parameters corresponding to the minimum error between the multi-modal features output by the hybrid backbone network and the actual multi-modal features, and use the hybrid backbone network constructed with the corresponding network parameters as the trained hybrid backbone network.
[0020] Further, the aligned multi-modal data set includes the aligned visible light images, the aligned thermal imaging images, and the aligned laser 3D point clouds.
[0021] Further, the method for obtaining the aligned multi-modal data set includes:
[0022] Obtain the GPS timestamp of the sensor at the data acquisition time step, generate an intermediate frame by weighted averaging the pixel values of two adjacent frames of visible light images before and after the GPS timestamp, and use the intermediate frame as the aligned visible light image;
[0023] Take the thermal imaging image as the input of the image alignment network to obtain the aligned thermal imaging image;
[0024] Convert the laser 3D point cloud from the lidar coordinate system to the camera coordinate system through a rotation matrix and a translation vector to obtain the 3D camera coordinates, and then project the 3D camera coordinates onto the 2D image plane through the camera intrinsic matrix to obtain the projected point cloud. Filter the point cloud in the projected point cloud that exceeds the boundary of the visible light image to obtain the aligned laser 3D point cloud.
[0025] Further, the training method of the image alignment network includes:
[0026] Pre-collect a set of image training data, where the image training data includes thermal images and corresponding aligned thermal images;
[0027] Use the thermal image as the input of the image alignment network, and the aligned thermal image as the output of the image alignment network. Calculate the temperature difference between the high-resolution thermal image and the thermal image, calculate the constraint loss function through the mean square error, aiming to minimize the constraint loss function, and optimize the network parameters of the image alignment network through a nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the aligned thermal image output by the image alignment network and the actual aligned thermal image. The image alignment network constructed with the corresponding network parameters is used as the trained image alignment network.
[0028] Further, the method for obtaining the fused unified feature representation includes:
[0029] Encode the environmental data to obtain an environmental embedding vector. Use the multi-modal domain-invariant feature vector and the environmental embedding vector as the input of the attention module to obtain the attention weights corresponding to the multi-modal domain-invariant feature vector and the environmental embedding vector, and perform weighted fusion on the multi-modal domain-invariant feature vector and the environmental embedding vector according to the attention weights to obtain the fused unified feature representation.
[0030] Further, the training method of the attention module includes:
[0031] Pre-collect a set of weight training data, where the weight training data includes multi-modal domain-invariant feature vectors, environmental embedding vectors, and corresponding attention weights;
[0032] Use the multi-modal domain-invariant feature vector and the environmental embedding vector as the input of the attention module, and the corresponding attention weights as the output of the attention module. Aiming to minimize the error between the output corresponding attention weights and the actual corresponding attention weights, optimize the network parameters of the attention module through a nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the output corresponding attention weights of the attention module and the actual corresponding attention weights. The attention module constructed with the corresponding network parameters is used as the trained attention module.
[0033] Further, the pantograph-catenary defect types include catenary wear, pantograph slider crack, arcing ablation, foreign object attachment, mechanical deformation, and oxidation corrosion; the severity is classified using a preset number of levels.
[0034] Further, the training method of the detection and classification network includes:
[0035] Pre-collect K groups of training data. The training data includes the fused unified feature representation and the output results. The output results include the defect category probability distribution, the position and size regression value, and the severity score.
[0036] Use the fused unified feature representation as the input of the detection and classification network, and use the output results as the output of the detection and classification network. With the goal of minimizing the error between the output results and the actual results, optimize the network parameters of the detection and classification network through a nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the output results output by the detection and classification network and the actual output results. Construct the detection and classification network with the corresponding network parameters as the trained detection and classification network. Among them, the detection and classification network includes a multi-task detection head, and the multi-task detection head includes a classification head, a regression head, and a severity grading head.
[0037] The technical effects and advantages of the catenary defect detection and feature fusion method based on deep transfer learning of the present invention:
[0038] Through multi-modal data acquisition and alignment processing, the present invention realizes the spatio-temporal synchronization of visible light, infrared thermal imaging, and laser point cloud, ensuring the multi-dimensional complementarity of defect features; combines domain adaptation technology to eliminate the data distribution differences between the laboratory and the actual scene, and improves the generalization ability of the model under complex working conditions; enhances the robustness of features to environmental interferences such as temperature and humidity through dynamic weighted fusion of environmental parameters; finally, uses a multi-task network to output the defect category, position and size, and severity, providing a complete decision-making basis for catenary operation and maintenance; breaks through the limitations of single-modal detection, improves the defect recognition accuracy and environmental adaptability, reduces the dependence on labeled data, realizes the collaborative output of qualitative and quantitative analysis of defects, and significantly improves the intelligent level and operation and maintenance efficiency of catenary equipment status assessment. Description of the Drawings
[0039] Figure 1 It is a schematic flow chart of the catenary defect detection and feature fusion method based on deep transfer learning of the present invention;
[0040] Figure 2 It is a schematic logical architecture diagram of the present invention;
[0041] Figure 3 It is a schematic diagram of the hybrid backbone network structure framework of the present invention;
[0042] Figure 4 It is a schematic diagram of the detection and classification network architecture of the present invention. Detailed Embodiments
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] Embodiment 1
[0045] Please refer to Figure 1 - Figure 2 As shown, the method for pantograph-catenary defect detection and feature fusion based on deep transfer learning in this embodiment includes the following steps:
[0046] Collect sensor data, where the sensor data includes detection data and environmental data; the detection data includes visible light images, infrared thermal images, and laser 3D point clouds of the object to be detected; the visible light images are collected by a high-precision camera, which can provide intuitive visual information such as the surface texture and color of the pantograph-catenary, and help extract the features of surface defects such as contact wire wear and foreign object attachment; the infrared thermal image is collected by an infrared thermal imaging device, such as a FLIR A700 device, to capture the temperature distribution of the pantograph-catenary components, and detect hidden defects through temperature anomalies (such as high temperatures at arcing points), which can provide key evidence for the analysis of heat-related defects; the laser 3D point cloud is collected by a depth camera, which can obtain the three-dimensional spatial structure data of the pantograph-catenary, accurately locate spatial geometric defects such as mechanical deformation and dimensional deviation, and improve the spatial dimension information of defect features; the environmental data includes temperature and humidity; the temperature can be collected by a temperature sensor, and the humidity can be collected by a humidity sensor. As environmental parameters, temperature and humidity can affect the state of the pantograph-catenary equipment and data features, and are used to adjust the weights of different modality data during feature fusion to improve the adaptability of the model to complex environments;
[0047] Align the detection data to obtain an aligned multi-modal data set;
[0048] The method for obtaining the aligned multi-modal data set includes:
[0049] Obtain the GPS timestamp of the sensor at the data collection time step, generate an intermediate frame by weighted averaging the pixel values of two adjacent frames of visible light images before and after the GPS timestamp, and use the intermediate frame as the aligned visible light image;
[0050] Use the thermal imaging image as the input of the image alignment network to obtain the aligned thermal imaging image;
[0051] The training method of the image alignment network includes:
[0052] Pre-collect a set of image training data, where the image training data includes thermal imaging images and the corresponding aligned thermal imaging images;
[0053] Use the thermal imaging image as the input of the image alignment network, and use the aligned thermal imaging image as the output of the image alignment network. Calculate the temperature difference between the high-resolution thermal imaging image and the thermal imaging image, and calculate the constraint loss function through the mean square error. With the goal of minimizing the constraint loss function, optimize the network parameters of the image alignment network through a nature-inspired optimization algorithm, and obtain the network parameters corresponding to the minimum error between the aligned thermal imaging image output by the image alignment network and the actually aligned thermal imaging image. Use the image alignment network constructed with the corresponding network parameters as the trained image alignment network.
[0054] Convert the laser 3D point cloud from the lidar coordinate system to the camera coordinate system through the rotation matrix and translation vector to obtain the 3D camera coordinates, and then project the 3D camera coordinates onto the 2D image plane through the camera internal parameter matrix to obtain the projected point cloud. Filter the point cloud that exceeds the visible light image boundary in the projected point cloud to obtain the aligned laser 3D point cloud.
[0055] For the multi-modal dataset obtained after aligning the detection data, through spatio-temporal synchronization (such as aligning time based on GPS timestamps and unifying space through coordinate system conversion), make data such as visible light images, infrared thermal imaging, and laser 3D point clouds accurately correspond to the same pantograph-catenary physical scene, laying a core foundation for feature fusion. This operation eliminates the misalignment caused by sampling differences in the original data, ensuring that each modal feature (such as visual texture, temperature distribution, spatial structure) can be truly complementary during fusion, enabling the model to comprehensively capture multi-dimensional features of defects such as contact wire wear and arc ablation during defect detection, improving the integrity and consistency of feature expression, and ultimately enhancing the accuracy and reliability of defect recognition, contributing to a more accurate pantograph-catenary state assessment.
[0056] Perform domain adaptive feature extraction on the aligned multi-modal dataset to obtain multi-modal domain-invariant feature vectors; it can eliminate the distribution differences of data in different scenarios such as laboratory simulations (source domain) and actual lines (target domain), enabling features to focus on the essential attributes of pantograph-catenary defects (such as common features like wear morphology and temperature anomalies). This provides a unified and stable foundation for feature fusion, ensuring that the model is not interfered by domain differences when fusing multi-modal information such as vision, thermal imaging, and point clouds, accurately capturing multi-dimensional features of defects, improving the model's generalization detection ability in complex actual scenarios, avoiding misjudgments caused by data distribution deviations, and ultimately achieving more accurate and reliable pantograph-catenary defect recognition and classification.
[0057] Methods for obtaining multi-modal domain-invariant feature vectors include:
[0058] Use the aligned multi-modal datasets in the source domain and the target domain as the input of the hybrid backbone network to obtain the corresponding multi-modal features, and calculate the MMD loss between the multi-modal features in the source domain and the target domain according to the preselected kernel function, such as:
[0059]
[0060] Among them, mmd loss is the MMD loss; feat src is the source domain multimodal feature; feat tgt is the target domain multimodal feature; n s is the number of source domain multimodal features; is the i-th multimodal feature in the source domain; n t is the number of target domain multimodal features; is the j-th multimodal feature in the target domain; φ is the mapping function that maps features to the reproducing kernel Hilbert space; H is the reproducing kernel Hilbert space; ||·||2 is the norm in the reproducing kernel Hilbert space H;
[0061] Input the source domain multimodal features and the target domain multimodal features into the domain classifier to obtain the probability that the input multimodal features come from the source domain; combine the cross-entropy loss to obtain the domain loss of the domain classifier; for example: Among them, is the domain loss; y src is the label of the i-th multimodal feature in the source domain; y tgt is the label of the j-th multimodal feature in the target domain; p src is the output of the domain classifier for the i-th multimodal feature in the source domain; p tgt is the output of the domain classifier for the j-th multimodal feature in the target domain;
[0062] Add the MMD loss and the domain loss to obtain the domain adaptation loss, and optimize the hybrid backbone network and the domain classifier with the minimum domain adaptation loss as the optimization goal. Use the enhanced multimodal features output by the domain classifier when the domain adaptation loss is the smallest as the multimodal domain-invariant feature vector.
[0063] Refer to Figure 3 , the training method of the hybrid backbone network includes:
[0064] Pre-collect B groups of mixed training data. The mixed training data includes the aligned multimodal dataset and multimodal features; the aligned multimodal dataset includes the aligned visible light images, the aligned infrared thermal images, and the aligned laser 3D point clouds; the multimodal features include visible light features, infrared features, and point cloud features;
[0065] The hybrid backbone network includes a visible light branch, an infrared branch, and a point cloud processing branch. The aligned visible light image is used as the input of the visible light branch, and the corresponding visible light features are used as the output of the visible light branch. The aligned infrared thermal image is used as the input of the infrared branch, and the corresponding infrared features are used as the output of the infrared branch. The aligned laser 3D point cloud is used as the input of the point cloud processing branch, and the corresponding point cloud features are used as the output of the point cloud processing branch. Taking the minimum error between the output multimodal features and the actual multimodal features as the goal, the network parameters of the hybrid backbone network are optimized through a nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the multimodal features output by the hybrid backbone network and the actual multimodal features. The hybrid backbone network constructed with the corresponding network parameters is used as the trained hybrid backbone network.
[0066] Encode the environmental data, generate dynamic weights through the attention module, and weighted fuse the multimodal domain-invariant feature vectors according to the dynamic weights to obtain a unified feature representation after fusion.
[0067] The method for obtaining the unified feature representation after fusion includes:
[0068] Encode the environmental data to obtain an environmental embedding vector. Use the multimodal domain-invariant feature vector and the environmental embedding vector as the input of the attention module to obtain the attention weights corresponding to the multimodal domain-invariant feature vector and the environmental embedding vector. Weighted fuse the multimodal domain-invariant feature vector and the environmental embedding vector according to the attention weights to obtain a unified feature representation after fusion.
[0069] The training method of the attention module includes:
[0070] Pre-collect C groups of weight training data, where the weight training data includes multimodal domain-invariant feature vectors, environmental embedding vectors, and corresponding attention weights.
[0071] Use the multimodal domain-invariant feature vector and the environmental embedding vector as the input of the attention module, and the corresponding attention weights as the output of the attention module. Taking the minimum error between the output corresponding attention weights and the actual corresponding attention weights as the goal, optimize the network parameters of the attention module through a nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the corresponding attention weights output by the attention module and the actual corresponding attention weights. The attention module constructed with the corresponding network parameters is used as the trained attention module.
[0072] After encoding the environmental data, the environmental information is transformed into a feature dimension understandable by the model. By means of an attention module, the association between the environment and multi-modal domain-invariant features (such as visual, thermal imaging, geometric features) is analyzed to generate dynamic weights. Based on this weight, feature weighted fusion is carried out, enabling the model to dynamically adjust the importance of each modal feature according to the real-time environment (such as temperature and humidity, running speed); for example, in a humid environment, the weight of the temperature anomaly feature of infrared thermal imaging is strengthened, and when running at high speed, the weight of the geometric structure feature of laser point cloud is emphasized. The finally obtained unified feature representation is closely combined with the environmental working conditions, enabling the model to be more in line with the actual scenario when fusing multi-modal information, enhancing the adaptability to complex environments, and significantly improving the accuracy and reliability of pantograph-catenary defect (such as catenary wear, arcing ablation, etc.) detection.
[0073] The fused unified feature representation is used as the input of the detection and classification network to obtain the probability distribution of pantograph-catenary defect categories, the regression values of position and size, and the severity score; among them, the types of pantograph-catenary defects include catenary wear, crack of pantograph slider, arcing ablation, foreign object attachment, mechanical deformation, and oxidation corrosion; the multi-dimensional annotation standards include defect type, position, size, and severity, where the severity is classified using a preset number of levels (such as 4 levels).
[0074] Refer to Figure 4 , the training method of the detection and classification network includes:
[0075] K groups of training data are collected in advance, and the training data includes the fused unified feature representation and the output results, and the output results include the probability distribution of defect categories, the regression values of position and size, and the severity score;
[0076] The fused unified feature representation is used as the input of the detection and classification network, and the output results are used as the output of the detection and classification network. With the goal of minimizing the error between the output results and the actual results, the network parameters of the detection and classification network are optimized through a nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the output results output by the detection and classification network and the actual output results, and the detection and classification network constructed with the corresponding network parameters is used as the trained detection and classification network; among them, the detection and classification network includes a multi-task detection head, and the multi-task detection head includes a classification head, a regression head, and a severity grading head.
[0077] The fused unified feature representation is input into the detection and classification network, and the probability distribution of defect categories, the regression values of position and size, and the severity score are output in parallel through multiple task heads, achieving accurate classification and quantitative evaluation of pantograph-catenary defects. Among them, the category probability distribution comprehensively judges the defect type based on multi-modal features (such as contact wire wear, arcing ablation, etc.), the position and size regression values locate the defect spatial position and measure its size through geometric features, and the severity score evaluates the defect development stage by combining features such as temperature gradient and texture complexity. The three work together to provide both qualitative diagnosis (category) of defects and output quantitative analysis (position / size / severity), providing a complete decision-making basis for pantograph-catenary operation and maintenance. At the same time, multi-task learning improves the model efficiency by sharing features and enhances the robustness of feature fusion, ultimately achieving high precision and high reliability in defect detection.
[0078] Embodiment 2
[0079] This embodiment provides a method for dynamic multi-modal data alignment and enhancement, including the following steps:
[0080] Take the detection data as the input of the two-channel self-supervised network to obtain calibrated detection data.
[0081] Convert the calibrated visible light image back to the original modality, compare it with the input visible light image to obtain the cycle consistency loss, project the three-dimensional coordinates of the laser 3D point cloud onto the visible light image plane, calculate the reprojection error, and obtain the geometric constraint loss; take the sum of the cycle consistency loss and the geometric constraint loss as the minimum target to adjust the transformation matrix between sensors to adapt to environmental changes.
[0082] Take the calibrated detection data as the input of CycleGAN to obtain enhanced detection data; among them, the discriminator of CycleGAN simultaneously distinguishes the real calibrated detection data from the synthetic calibrated detection data, forcing the generator to learn domain-invariant features (such as defect geometry); CycleGAN also introduces an identity loss to ensure that the generator keeps the calibrated detection data unchanged when the input is in the same modality, avoiding feature distortion.
[0083] Through the above closed-loop optimization of dynamic calibration and data enhancement, it is possible to achieve the enhancement of the spatio-temporal consistency of detection data, alleviate the problem of small-sample defect detection, and significantly improve the defect recognition accuracy of the model in unseen scenarios.
[0084] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0085] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for pantograph-catenary defect detection and feature fusion based on deep transfer learning, characterized in that The method includes the following steps: Collect sensor data, where the sensor data includes detection data and environmental data; the detection data includes visible light images, infrared thermal images, and laser 3D point clouds of the object to be detected; the environmental data includes temperature and humidity; Perform alignment processing on the detection data to obtain an aligned multi-modal data set; Perform domain adaptation feature extraction on the aligned multi-modal data set to obtain multi-modal domain-invariant feature vectors; Encode the environmental data, generate dynamic weights through an attention module, and perform weighted fusion on the multi-modal domain-invariant feature vectors according to the dynamic weights to obtain a fused unified feature representation; Use the fused unified feature representation as the input of the detection and classification network to obtain the probability distribution of catenary defect categories, the regression values of position and size, and the severity score.
2. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 1, wherein The method for obtaining the multi-modal domain-invariant feature vectors includes: Use the aligned multi-modal data sets in the source domain and the target domain as the input of the hybrid backbone network to obtain corresponding multi-modal features, and calculate the MMD loss between the multi-modal features in the source domain and the target domain according to a preselected kernel function; Input the source domain multi-modal features and the target domain multi-modal features into the domain classifier to obtain the probability that the input multi-modal features come from the source domain; combine the cross-entropy loss to obtain the domain loss of the domain classifier; Add the MMD loss and the domain loss to obtain the domain adaptation loss, optimize the hybrid backbone network and the domain classifier with the minimum domain adaptation loss as the optimization goal, and use the enhanced multi-modal features output by the domain classifier when the domain adaptation loss is the smallest as the multi-modal domain-invariant feature vectors.
3. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 2, wherein The training method of the hybrid backbone network includes: Pre-collect B groups of hybrid training data, where the hybrid training data includes the aligned multi-modal data set and multi-modal features; the aligned multi-modal data set includes aligned visible light images, aligned infrared thermal images, and aligned laser 3D point clouds; the multi-modal features include visible light features, infrared features, and point cloud features; The hybrid backbone network includes a visible light branch, an infrared branch, and a point cloud processing branch. Use the aligned visible light image as the input of the visible light branch and the corresponding visible light feature as the output of the visible light branch; use the aligned infrared thermal image as the input of the infrared branch and the corresponding infrared feature as the output of the infrared branch; use the aligned laser 3D point cloud as the input of the point cloud processing branch and the corresponding point cloud feature as the output of the point cloud processing branch. With the goal of minimizing the error between the output multi-modal features and the actual multi-modal features, optimize the network parameters of the hybrid backbone network through a nature-inspired optimization algorithm, obtain the network parameters corresponding to the minimum error between the multi-modal features output by the hybrid backbone network and the actual multi-modal features, and use the hybrid backbone network constructed with the corresponding network parameters as the trained hybrid backbone network.
4. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 1, wherein The aligned multi-modal data set includes aligned visible light images, aligned thermal images, and aligned laser 3D point clouds.
5. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 4, wherein The method for obtaining the aligned multi-modal data set includes: Obtain the GPS timestamp of the sensor at the data acquisition time step, generate an intermediate frame by weighted averaging the pixel values of two adjacent frames of visible light images before and after the GPS timestamp, and use the intermediate frame as the aligned visible light image; Use the thermal imaging image as the input of the image alignment network to obtain the aligned thermal imaging image; Convert the lidar 3D point cloud from the lidar coordinate system to the camera coordinate system through the rotation matrix and translation vector to obtain the 3D camera coordinates, and then project the 3D camera coordinates onto the 2D image plane through the camera internal parameter matrix to obtain the projected point cloud. Filter the point cloud that exceeds the boundary of the visible light image in the projected point cloud to obtain the aligned lidar 3D point cloud.
6. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 5, characterized in that The training method of the image alignment network includes: Pre-collect A groups of image training data, where the image training data includes thermal imaging images and the corresponding aligned thermal imaging images; Use the thermal imaging image as the input of the image alignment network, use the aligned thermal imaging image as the output of the image alignment network, calculate the temperature difference between the high-resolution thermal imaging image and the thermal imaging image, calculate the constraint loss function through the mean square error, aim at minimizing the constraint loss function, and optimize the network parameters of the image alignment network through the nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the aligned thermal imaging image output by the image alignment network and the actual aligned thermal imaging image, and use the image alignment network constructed by the corresponding network parameters as the trained image alignment network.
7. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 1, wherein The method for obtaining the fused unified feature representation includes: Encode the environmental data to obtain an environmental embedding vector, use the multi-modal domain-invariant feature vector and the environmental embedding vector as the input of the attention module to obtain the attention weights corresponding to the multi-modal domain-invariant feature vector and the environmental embedding vector, and perform weighted fusion on the multi-modal domain-invariant feature vector and the environmental embedding vector according to the attention weights to obtain the fused unified feature representation.
8. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 7, wherein The training method of the attention module includes: Pre-collect C groups of weight training data, where the weight training data includes multi-modal domain-invariant feature vectors, environmental embedding vectors, and the corresponding attention weights; Use the multi-modal domain-invariant feature vector and the environmental embedding vector as the input of the attention module, use the corresponding attention weights as the output of the attention module, aim at minimizing the error between the output corresponding attention weights and the actual corresponding attention weights, optimize the network parameters of the attention module through the nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the corresponding attention weights output by the attention module and the actual corresponding attention weights, and use the attention module constructed by the corresponding network parameters as the trained attention module.
9. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 1, wherein The types of pantograph-catenary defects include catenary wear, pantograph slider crack, arcing ablation, foreign object attachment, mechanical deformation, and oxidation corrosion; the severity is classified using a preset number of levels.
10. The method for pantograph-catenary defect detection and feature fusion based on deep transfer learning according to claim 1, wherein, The training method of the detection and classification network includes: Pre-collect K groups of training data, where the training data includes the fused unified feature representation and the output results, and the output results include defect category probability distribution, position and size regression values, and severity scores; Use the fused unified feature representation as the input of the detection and classification network, and use the output result as the output of the detection and classification network. With the goal of minimizing the error between the output result and the actual result, optimize the network parameters of the detection and classification network through a nature-inspired optimization algorithm to obtain the network parameters corresponding to the minimum error between the output result output by the detection and classification network and the actual output result. Construct the detection and classification network with the corresponding network parameters as the trained detection and classification network; among them, the detection and classification network includes a multi-task detection head, and the multi-task detection head includes a classification head, a regression head, and a severity grading head.
Citation Information
Patent Citations
Overhead Contact Line Operation Status Safety Monitoring and Analysis System
CN105652154B
Cited By
Catenary thermal defect geometric parameter false alarm filtering method based on visual image
CN121544628A