Catenary multi-source multi-dimensional data feature extraction and processing method
By constructing the contact network data set and using multi-source feature extraction and fusion analysis methods, the problem of multi-source multi-dimensional data fusion is solved, and comprehensive, real-time and accurate monitoring of the contact network status is achieved to adapt to the application needs of different scenarios.
Patent Information
- Application Number
- CN202510433787.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-26
AI Technical Summary
The data formats, sampling frequency and feature expression methods of different data sources vary greatly, which makes it difficult to effectively integrate multi-source and multi-dimensional data on the contact network, affecting the comprehensiveness and accuracy of detection.
A contact network data set is constructed, including image data, dynamic geometric data and defect text data, and an image feature extraction model, a dynamic geometric feature extraction model and a defect text feature extraction model are used, and feature fusion analysis is performed in combination with a ternary comparison loss function to achieve deep fusion of multi-source and multi-dimensional data.
It realizes comprehensive, real-time and accurate monitoring of the status of the contact network, provides richer status information, is versatile and scalable, and can adapt to application needs in different scenarios.
Smart Images

Figure CN120541745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning, image processing, and signal processing, and more specifically to a method for extracting and processing features of multi-source and multi-dimensional data of a contact network. Background Art
[0002] In modern railway and urban rail transit systems, the overhead contact network is an important component of electric traction power supply, and its operating status is directly related to the safe operation of trains and the reliability of power supply.
[0003] With the rapid development of sensor technology, image processing technology, and artificial intelligence algorithms, catenary inspection is gradually moving towards multi-source data fusion and intelligentization. By fusing multi-source and multi-dimensional data such as image data, defect text descriptions, and sensor signals, it is possible to more comprehensively and accurately extract the state characteristics of the catenary, thereby achieving efficient data processing and defect identification.
[0004] However, the data formats, sampling frequencies and feature expressions of different data sources vary greatly, and how to effectively integrate these heterogeneous data is a technical difficulty. Summary of the Invention
[0005] In order to overcome the defects existing in the above-mentioned prior art, the present invention discloses a method for extracting and processing multi-source and multi-dimensional data features of the contact network, aiming to realize feature extraction and deep fusion of defect text data, image data, and dynamic data, so as to achieve comprehensive, real-time and accurate monitoring of the contact network status.
[0006] In order to achieve the above objectives, the present invention adopts the following technical solutions:
[0007] A method for extracting and processing features of multi-source and multi-dimensional data of a contact network comprises the following steps:
[0008] 1. Building a Dataset
[0009] S1. Constructing a contact network dataset, wherein the contact network dataset includes image data, dynamic geometry data, and defect text data of the contact network;
[0010] Preferably, constructing the contact network data set includes: using cameras and sensors to collect contact network image data and contact network dynamic geometry data respectively, using defect reports to record contact network defect text data, the contact network image data, contact network dynamic geometry data and contact network defect text data form image-dynamic geometry-defect text data triples, and in the image-dynamic geometry-defect text data triples, each contact network image data, contact network dynamic geometry data and contact network defect text data are paired by timestamps and spatial coordinates.
[0011] In the present invention, the contact network image data refers to various contact network defect images and normal images collected by the camera. Its function is to enable the feature fusion analysis model to learn relevant information in the contact network image, including defect features, normal status and other useful information in the image.
[0012] The dynamic geometric data of the contact network refers to various vehicle operation data collected by sensors during vehicle operation. Its function is to help the feature fusion analysis model to more accurately understand the actual operating status of the contact network, thereby realizing the accurate detection and positioning of contact network defects.
[0013] Contact network defect text data refers to text description information related to the contact network, such as equipment name, defect type, defect location, defect degree, etc. Its function is to provide the feature fusion analysis model with descriptions related to contact network knowledge and text descriptions of defects.
[0014] 2. Feature Extraction
[0015] S2. Using an image feature extraction model, a dynamic geometry feature extraction model, and a defect text feature extraction model, respectively, extract image features from the image data, dynamic geometry features from the dynamic geometry data, and defect text features from the defect text data to obtain high-dimensional feature vectors of the image data, the dynamic geometry data, and the defect text data;
[0016] 2.1 Image Feature Extraction
[0017] Preferably, using an image feature extraction model to extract image features from image data to obtain a high-dimensional feature vector of the image data includes the following steps:
[0018] S211, pre-processing the contact network image data collected by the camera;
[0019] Preferably, the preprocessing of the contact network image data collected by the camera includes: denoising, enhancement, cropping and normalization operations.
[0020] Preferably, the denoising comprises: applying Gaussian filtering to remove image noise;
[0021] The cropping and normalization include: cropping the image to 224×224 pixels and normalizing it to the range of [0, 1]. The size of each image is 224×224×3, representing RGB color channels.
[0022] S212, using the image encoder of the CLIP model to construct an image feature extraction model;
[0023] Preferably, the image feature extraction model adopts the image encoder part of the CLIP model, is based on a deep learning architecture, and is pre-trained through an image dataset to learn the underlying visual features and high-level semantic features of the image.
[0024] Preferably, in the image feature extraction model, ResNet-50 is used as the image encoder, and the ResNet-50 architecture uses residual connections. In the image feature extraction model:
[0025] Input: preprocessed image, size 224×224×3;
[0026] Convolutional layer: The initial convolutional layer is 7×7 convolution with a stride of 2 and an output size of 112×112×64;
[0027] Residual module: includes multiple residual units, each of which includes two to three layers of convolution;
[0028] Fully connected layer: Output a 2048-dimensional feature vector fimg∈R through the final fully connected layer 2048 , which is used for subsequent fusion with other modality data.
[0029] S213 , inputting the pre-processed contact network image data into an image feature extraction model to obtain a high-dimensional feature vector of the image data.
[0030] 2.2 Dynamic Geometric Feature Extraction
[0031] Preferably, the dynamic geometric feature extraction model is used to extract dynamic geometric features from the dynamic geometric data to obtain a high-dimensional feature vector of the dynamic geometric data, including the following steps:
[0032] S221, preprocessing the contact network dynamic geometry data collected by the sensor;
[0033] Preferably, the preprocessing of the contact network dynamic geometric data collected by the sensor includes: data calibration, filtering, denoising and normalization operations.
[0034] Preferably, the calibration includes: correcting the sensor acquisition time and accuracy;
[0035] The filtering and denoising include: removing high-frequency noise by low-pass filtering;
[0036] The normalization includes: standardizing the contact force, current, and vibration data so that their mean is 0 and their variance is 1.
[0037] S222. Constructing a dynamic geometric feature extraction model using long short-term memory networks and convolutional neural networks;
[0038] Preferably, in the dynamic geometric feature extraction model, for time series data, a long short-term memory network is used to capture the temporal dependencies in the data; for spatial distribution data, a convolutional neural network is used to extract the spatial features of the data.
[0039] Preferably, in the dynamic geometric feature extraction model, the dynamic geometric data X geom ∈R 100×3 , represents the three sensor features within 100 time steps: contact force, current, and vibration;
[0040] The long short-term memory network includes two layers, each layer includes an input gate, a forget gate and an output gate, and the number of hidden units in each layer is 256. The dynamic geometric data is processed by the long short-term memory network, and the output feature vector is f geom ∈R 256 .
[0041] S223. Input the pre-processed contact network dynamic geometry data into a dynamic geometry feature extraction model to obtain a high-dimensional feature vector of the dynamic geometry data.
[0042] 2.3. Defect Text Feature Extraction
[0043] Preferably, using a defect text feature extraction model to extract defect text features from defect text data to obtain a high-dimensional feature vector of the defect text data includes the following steps:
[0044] S231, pre-processing the defect text data recorded in the defect report;
[0045] Preferably, the preprocessing of the defect text data recorded in the defect report includes: text cleaning, word segmentation, stop word removal and stemming operations.
[0046] Preferably, the text cleaning includes: removing punctuation marks and unnecessary symbols in the ledger text, retaining key information, and processing spelling errors;
[0047] The word segmentation includes: using the jieba word segmentation tool to segment the text into words.
[0048] S232. Using the text encoder of the CLIP model, a defect text feature extraction model is constructed;
[0049] Preferably, the defect text feature extraction model learns the semantic information and contextual association of the text by pre-training on a large-scale text dataset.
[0050] S233. Input the preprocessed contact network defect text data into a defect text feature extraction model to obtain a high-dimensional feature vector of the defect text data.
[0051] In the present invention, the image features in the image data, the dynamic geometry features in the dynamic geometry data, and the defect text features in the defect text data are extracted to obtain high-dimensional feature vectors of the image data, the dynamic geometry data, and the defect text data because:
[0052] The high-dimensional feature vectors of image data can accurately represent the content of the image and provide a basis for subsequent feature fusion and analysis.
[0053] The high-dimensional feature vectors of dynamic geometric data can accurately represent the spatiotemporal characteristics of dynamic geometric data and provide a basis for subsequent feature fusion and analysis.
[0054] The high-dimensional feature vectors of defect text data can accurately represent the semantic content and contextual information of the defect, providing a basis for subsequent feature fusion and analysis.
[0055] 3. Model Training
[0056] S3. Construct image sensor text positive and negative sample pairs, ternary contrast loss function, and feature fusion analysis model, and train the feature fusion analysis model using high-dimensional feature vectors, image sensor text positive and negative sample pairs, and ternary contrast loss function.
[0057] Preferably, the ternary contrast loss function includes:
[0058]
[0059] Among them, L(θ) is the ternary contrast loss function; θ is the model parameter; x, y, and z represent images, text, and dynamic geometric data, respectively; f, g, and h represent the embedding functions of images, text, and dynamic geometric data, respectively, and the embedding function maps the original input into a shared feature space; i is the corresponding modal data; j is the positive sample data of the other modality corresponding to i, which matches the modal data i; k is the negative sample data of the other modality corresponding to i, which does not match the modal data i; sim represents the similarity measurement function, which uses cosine similarity to calculate the similarity between two embedding vectors; m is a positive threshold used to control the interval between positive and negative sample pairs.
[0060] In the present invention, by setting the above-mentioned ternary contrast loss function, the model is allowed to learn effective embedding representations between data of different modalities (such as images, text, and dynamic geometric data), so that samples from different modalities but semantically similar are closer in the embedding space to achieve the purpose of cross-modal learning; at the same time, by introducing negative samples, the model is forced to learn to distinguish data of different categories, which helps to enhance the robustness of the model.
[0061] Preferably, the constructing of the image sensor text positive and negative sample pairs includes constructing positive samples and constructing negative samples;
[0062] The positive sample construction includes constructing a positive sample pair, wherein the positive sample pair refers to images, texts and dynamic geometry data of the same defect category;
[0063] The negative sample construction includes constructing negative sample pairs and constructing hard negative samples. The negative sample pairs refer to images, text, and dynamic geometry data all belonging to different defect categories; the hard negative samples refer to two of the three types of data, namely, images, text, and dynamic geometry data, belonging to the same defect category, and the other one belonging to a different defect category.
[0064] Preferably, the training of the feature fusion analysis model includes:
[0065] For images, texts, and dynamic geometry data, image feature extraction models, defect text feature extraction models, and dynamic geometry feature extraction models are used for cleaning and feature extraction, and a fully connected layer is used to map images, texts, and dynamic geometry data to the same dimension R. 512 ;
[0066] Construct image, text, and dynamic geometry data pairs through positive and negative sample pair construction;
[0067] The ternary contrast loss function is calculated using image, text, and dynamic geometry data pairs. The gradient calculated by the ternary contrast loss function is back-propagated to each feature extraction model to update the parameters of each modal feature extraction model. The Adam optimizer is used to update the gradient, adjust the weights and biases of the feature fusion analysis model, and gradually optimize the accuracy of feature extraction and similarity calculation.
[0068] The training process is optimized, including learning rate adjustment and batch training. In the learning rate adjustment, the initial learning rate is set to 1e-4 and gradually reduced through a learning rate decay strategy. In the batch training, in each training cycle, the model processes a batch of data, each batch includes 64 samples, and the samples in each batch are calculated using a ternary contrastive loss function to calculate the loss of positive and negative sample pairs, and then update the model.
[0069] Hyperparameter adjustment includes adjusting the interval m in the ternary contrastive loss and adjusting the positive-negative sample ratio. In the interval m adjustment in the ternary contrastive loss, m is set to 0.2; in the positive-negative sample ratio adjustment, balanced sampling is used to adjust the distribution of positive and negative samples.
[0070] 4. Downstream Use
[0071] S4. Provide the image feature extraction model, dynamic geometric feature extraction model, defect text feature extraction model, and feature fusion analysis model for downstream use.
[0072] Preferably, when the model is used downstream, the feature fusion analysis model learns the matching relationship between image, text and sensor data through the optimization of the ternary contrast loss function; when querying, the user provides a modal data, and the feature fusion analysis model uses the embedding space learned by the ternary contrast loss to return other modal data that is most relevant to the query modality. This query method is used for defect detection, cross-modal information retrieval and multimodal anomaly detection tasks.
[0073] Preferably, the model for downstream use includes:
[0074] Given a query modality data, the query modality data includes image data, dynamic geometry data or defect text data of the contact network;
[0075] The query modality data is converted into a feature vector using a feature extraction model;
[0076] Calculate the similarity between the query modality feature vector and other modalities in the database;
[0077] Based on similarity ranking, the most relevant images, texts, or sensor data are returned as query results.
[0078] Beneficial effects of the present invention:
[0079] This invention employs multimodal data fusion: for the first time, the CLIP concept is applied to the processing of multi-source, multi-dimensional data from the contact network, achieving deep fusion of image, dynamic geometry, and text data. It also employs multi-source data feature extraction: a feature extraction technique that integrates image data collected by cameras, dynamic geometry data collected by sensors, and related descriptive text data. It also employs a multimodal feature training strategy: a ternary contrast loss function training strategy.
[0080] This invention is comprehensive: by collecting and fusing multimodal data, it provides more comprehensive and richer information on the state of the overhead line. It is also versatile: unsupervised learning can identify different scenarios or features that have never appeared before, and has strong transferability. It is scalable: the model can be expanded and upgraded based on actual needs to adapt to the application requirements of different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 Schematic diagram of the method for extracting and processing multi-source and multi-dimensional data features of the contact network according to the present invention. DETAILED DESCRIPTION
[0082] The following will provide a clear and complete description of the concept, specific structure and technical effects of the present invention in conjunction with the embodiments and drawings, so as to fully understand the purpose, features and effects of the present invention.
[0083] Example 1
[0084] A method for extracting and processing features of multi-source and multi-dimensional data of contact network, such as Figure 1 As shown, the following steps are included:
[0085] S1. Constructing a contact network dataset, wherein the contact network dataset includes image data, dynamic geometry data, and defect text data of the contact network;
[0086] S2. Using an image feature extraction model, a dynamic geometry feature extraction model, and a defect text feature extraction model, respectively, extract image features from the image data, dynamic geometry features from the dynamic geometry data, and defect text features from the defect text data to obtain high-dimensional feature vectors of the image data, the dynamic geometry data, and the defect text data;
[0087] S3. Construct image sensor text positive and negative sample pairs, ternary contrast loss function, and feature fusion analysis model, and train the feature fusion analysis model using high-dimensional feature vectors, image sensor text positive and negative sample pairs, and ternary contrast loss function.
[0088] S4. Provide the image feature extraction model, dynamic geometric feature extraction model, defect text feature extraction model, and feature fusion analysis model for downstream use.
[0089] Example 2
[0090] Based on Example 1, this example further elaborates on step S1. Step S1 is to construct a data set. The camera and sensor collect images and dynamic geometric data of the contact network, and combine them with text labels to form a matching image-text-dynamic geometric data triple.
[0091] In the catenary fault detection task, images, dynamic geometric data, and text defect reports constitute a multimodal dataset. Each data sample includes:
[0092] Image data: The contact network image is collected by a high-precision camera. The image data is collected by a high-definition camera and captures the image of the contact network. It may contain various situations, such as the location of the contact network, breakage or looseness and other faults.
[0093] Dynamic geometric data: Dynamic geometric data is collected in real time by sensors (contact force, current, vibration sensors) to monitor the mechanical properties of the contact network.
[0094] Defect text data: Fault reports recorded by personnel and automated systems, including detailed descriptions of catenary faults.
[0095] Each image, dynamic geometry data, and text description are paired using information such as timestamps and spatial coordinates to ensure alignment and consistency among the three.
[0096] Example 3
[0097] This embodiment further elaborates on step S2 based on embodiment 2. Step S2 is feature extraction, including image feature extraction, dynamic geometric feature extraction, and defect text feature extraction, as follows:
[0098] 1. Image feature extraction:
[0099] Preprocessing: First, the contact network images captured by the camera are preprocessed, including denoising, enhancement, cropping and normalization operations, to improve image quality and reduce the computational complexity of subsequent processing.
[0100] Denoising: Apply Gaussian filtering to remove image noise.
[0101] Cropping and normalization: The images are cropped to 224×224 pixels and normalized to the range of [0,1]. The size of each image is 224×224×3, representing the RGB color channels.
[0102] Feature extraction model: This uses the image encoder portion of the CLIP model, which is typically based on deep learning architectures such as ResNet and ViT (Vision Transformer). These models are pre-trained on large-scale image datasets to learn both low-level visual features (such as edges and textures) and high-level semantic features (such as objects and scenes).
[0103] ResNet-50 is used as the image encoder. The ResNet-50 architecture uses skip connections to avoid the gradient vanishing problem and effectively extract multi-level image features.
[0104] Input: Preprocessed image, size 224×224×3.
[0105] Convolutional layer: The initial convolutional layer is a 7×7 convolution with a stride of 2 and an output size of 112×112×64.
[0106] Residual module: It consists of multiple residual units (each residual unit contains two to three layers of convolution), which makes the network deeper and more efficient to train.
[0107] Fully connected layer: Output a 2048-dimensional feature vector fimg∈R through the final fully connected layer 2048 , which is used for subsequent fusion with other modality data.
[0108] Feature vector generation: The preprocessed image is fed into CLIP's image encoder to generate high-dimensional feature vectors. These feature vectors accurately represent the image content and provide a foundation for subsequent feature fusion and analysis.
[0109] 2. Dynamic geometric data feature extraction:
[0110] Preprocessing: Preprocessing the dynamic geometric data of the contact network collected by the sensor, including data calibration, filtering, denoising and normalization operations to ensure the accuracy and stability of the data.
[0111] Calibration: Sensor acquisition time and accuracy correction to ensure that the output of each sensor is consistent with the actual physical quantity.
[0112] Filtering and denoising: Remove high-frequency noise through low-pass filtering to ensure data stability.
[0113] Normalization: Standardize the contact force, current, and vibration data to make their mean 0 and variance 1, thereby avoiding gradient explosion or disappearance during model training.
[0114] Feature extraction model: Different deep learning models are selected for feature extraction based on the type and characteristics of dynamic geometric data. For time series data, a long short-term memory network (LSTM) is used to capture temporal dependencies in the data; for spatially distributed data, a convolutional neural network (CNN) is used to extract spatial features.
[0115] Dynamic Geometry DataX geom ∈R 100×3 , representing the three sensor features (contact force, current, vibration) within 100 time steps.
[0116] The LSTM network consists of two layers, each layer includes an input gate, a forget gate, and an output gate, and the number of hidden units in each layer is 256. The dynamic geometric data is processed by the LSTM network, and the output feature vector is f geom ∈R 256 .
[0117] Feature vector generation: The preprocessed dynamic geometry data is fed into the corresponding deep learning model to generate high-dimensional feature vectors of the data. These feature vectors can accurately represent the spatiotemporal characteristics of the dynamic geometry data, providing a basis for subsequent feature fusion and analysis.
[0118] 3. Defect text feature extraction:
[0119] Preprocessing: Preprocess the recorded defect reports, including text cleaning, word segmentation, stop word removal, and stemming operations, to remove irrelevant information and retain key information.
[0120] Text cleaning: remove punctuation and unnecessary symbols from ledger text, retain key information, and address spelling errors.
[0121] Word segmentation: Use the Jieba word segmentation tool to segment the text into words.
[0122] Feature extraction model: The text encoder (Transformer) of the CLIP model is used to extract text features. These models are pre-trained on large-scale text datasets to learn the semantic information and context of text.
[0123] Feature vector generation: The preprocessed defect text is input into the natural language processing model to generate high-dimensional feature vectors of the text. These feature vectors can accurately represent the semantic content and contextual information of the defect, providing a basis for subsequent feature fusion and analysis.
[0124] Example 4
[0125] This embodiment further elaborates on step S3 based on embodiment 3. Step S3 constructs a ternary contrast loss function to maximize the correct matching similarity for training.
[0126] The ternary contrast loss function includes:
[0127]
[0128] Among them, L(θ) is the ternary contrast loss function; θ is the model parameter; x, y, and z represent images, text, and dynamic geometric data, respectively; f, g, and h represent the embedding functions of images, text, and dynamic geometric data, respectively, and the embedding function maps the original input into a shared feature space; i is the corresponding modal data; j is the positive sample data of the other modality corresponding to i, which matches the modal data i; k is the negative sample data of the other modality corresponding to i, which does not match the modal data i; sim represents the similarity measurement function, which uses cosine similarity to calculate the similarity between two embedding vectors; m is a positive threshold used to control the interval between positive and negative sample pairs.
[0129] During training, we need to construct positive and negative sample pairs so that we can optimize them using the triplet contrastive loss. The selection of positive and negative samples is crucial and impacts the learning performance of the model. The following is a detailed description of how to construct positive and negative sample pairs.
[0130] Positive sample construction:
[0131] Positive pairs: A positive pair consists of images, text, and dynamic geometry data representing the same defect category. For example, if an image shows a "broken suspension cord" defect, the matching text should describe the condition, while the dynamic geometry data should reflect sensor data related to the condition (e.g., contact force, current, etc.). These modalities should be highly similar, thus forming a positive pair.
[0132] Negative sample construction:
[0133] Negative pairs: Negative pairs are images, text, and dynamic geometry data that belong to different defect categories. The purpose of these negative pairs is to force the model to learn how to distinguish between different types of defects. For example, an image may show a "broken suspension cord," but the text may describe a "missing current ring." The dynamic geometry data may be irrelevant to either the image or the text.
[0134] Hard Negative Mining:
[0135] During training, the selection of negative samples is crucial. We usually select negative samples that are very similar to the positive samples, which are called “hard negative samples”.
[0136] For example, if the image and text description match, but the dynamic geometry data does not, such negative samples may be closer in feature space, and the model needs more training to distinguish them. Therefore, the frequency of such negative sample pairs will increase, helping the model learn more detailed distinctions.
[0137] Training process:
[0138] 1. Data input, preprocessing, and feature extraction:
[0139] For images, texts, and sensor data, the above schemes are used for cleaning and feature extraction, and a fully connected layer is used to map images, texts, and dynamic geometry data to the same dimension R. 512 .
[0140] 2. Sample Construction
[0141] Image, text, and dynamic data pairs are constructed using the above positive and negative sample pair construction method.
[0142] 3. Calculate the ternary contrast loss function
[0143] 4. Backpropagation and parameter update
[0144] The gradient calculated by the ternary contrast loss function is back-propagated to each network (image feature extraction network, text feature extraction network, geometric data feature extraction network), thereby updating the parameters of each modality feature extraction network; the Adam optimizer is used to update the gradient, adjust the weights and biases of the model, and gradually optimize the accuracy of feature extraction and similarity calculation.
[0145] 5. Optimize the training process
[0146] Learning rate adjustment: The initial learning rate can be set to 1e-4 and gradually reduced through the learning rate decay strategy to prevent gradient explosion or oscillation during training.
[0147] Batch training: During each training epoch, the model processes a batch of data, typically containing 64 samples. The samples in each batch are then evaluated using a ternary contrastive loss function to calculate the loss for each positive and negative sample pair, and the model is then updated. This helps accelerate model convergence, reduce the risk of overfitting, and improve training stability.
[0148] 6. Hyperparameter Tuning
[0149] The margin m in the ternary contrastive loss controls the minimum distance between positive and negative samples. It's typically set to a small constant, such as 0.2. If the margin m is too small, there may be little difference between positive and negative samples, affecting the model's ability to distinguish them. If the margin m is too large, negative samples may be too similar, making training difficult.
[0150] Positive-Negative Sample Ratio: In each training batch, the ratio of positive and negative sample pairs also has a certain impact on the training results. Balanced sampling is used to adjust the sample distribution and thus improve the effectiveness of training.
[0151] Example 5
[0152] This example further illustrates step S4 based on Example 4, and step S4 is used downstream.
[0153] By optimizing the ternary contrastive loss function, the model learns the matching relationships between images, text, and sensor data. When querying, the user can provide a modality (such as text, image, or sensor data), and the model will use the embedding space learned by the ternary contrastive loss to return data from other modalities that are most relevant to the query modality. This query method can be used not only for defect detection but also for tasks such as cross-modal information retrieval and multimodal anomaly detection. The specific processing flow is as follows:
[0154] Input: Given a query modality (image, text, or dynamic geometric data), for example, a defect text description "broken suspension string" entered by the user.
[0155] Processing: Use the trained encoder to convert the query modality into a feature vector.
[0156] Similarity calculation: Calculate the similarity between the query modality and other modalities in the database, such as calculating the cosine similarity between text and images, or the similarity between images and sensor data.
[0157] Return results: Based on similarity ranking, the most relevant images, texts, or sensor data are returned as query results.
[0158] The above is a detailed description of the embodiments of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without departing from the spirit of the present invention. These equivalents or substitutions are all included in the scope defined by the claims of the present invention.
Claims
1. A method for extracting and processing features of multi-source and multi-dimensional data of a contact network, characterized in that: The following steps are involved: Constructing a contact network data set, wherein the contact network data set includes image data, dynamic geometry data, and defect text data of the contact network; The image feature extraction model, the dynamic geometry feature extraction model and the defect text feature extraction model are used to extract the image features in the image data, the dynamic geometry features in the dynamic geometry data and the defect text features in the defect text data, respectively, to obtain high-dimensional feature vectors of the image data, the dynamic geometry data and the defect text data; Construct image sensor text positive and negative sample pairs, a ternary contrast loss function, and a feature fusion analysis model. Use high-dimensional feature vectors, image sensor text positive and negative sample pairs, and the ternary contrast loss function to train the feature fusion analysis model. The image feature extraction model, dynamic geometric feature extraction model, defect text feature extraction model, and feature fusion analysis model are provided for downstream use.
2. A method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: Constructing the contact network dataset includes: using cameras and sensors to collect contact network image data and contact network dynamic geometry data respectively, and using defect reports to record contact network defect text data. The contact network image data, contact network dynamic geometry data and contact network defect text data form image-dynamic geometry-defect text data triples. In the image-dynamic geometry-defect text data triples, each contact network image data, contact network dynamic geometry data and contact network defect text data are paired by timestamp and spatial coordinates.
3. The method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: Using an image feature extraction model to extract image features from image data and obtain a high-dimensional feature vector of the image data includes the following steps: Preprocessing the contact network image data collected by the camera; Use the image encoder of the CLIP model to build an image feature extraction model; Input the pre-processed catenary image data into the image feature extraction model to obtain the high-dimensional feature vector of the image data; in: The preprocessing of the contact network image data collected by the camera includes: denoising, enhancement, cropping and normalization operations; the denoising includes: applying Gaussian filtering to remove image noise; the cropping and normalization include: cropping the image to 224×224 pixels and normalizing it to the range of [0,1], with the size of each image being 224×224×3, representing RGB color channels; The image feature extraction model uses the image encoder part of the CLIP model, based on a deep learning architecture, and is pre-trained on an image dataset to learn the underlying visual features and high-level semantic features of the image. In the image feature extraction model, ResNet-50 is used as the image encoder. The ResNet-50 architecture uses residual connections. In the image feature extraction model: Input: preprocessed image, size 224×224×3; Convolutional layer: The initial convolutional layer is 7×7 convolution with a stride of 2 and an output size of 112×112×64; Residual module: includes multiple residual units, each of which includes two to three layers of convolution; Fully connected layer: Output a 2048-dimensional feature vector fimg∈R through the final fully connected layer 2048 , which is used for subsequent fusion with other modality data.
4. The method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: The dynamic geometric feature extraction model is used to extract dynamic geometric features from dynamic geometric data to obtain a high-dimensional feature vector of the dynamic geometric data, including the following steps: Preprocessing the dynamic geometric data of the contact network collected by sensors; Build a dynamic geometric feature extraction model using long short-term memory networks and convolutional neural networks; Input the pre-processed catenary dynamic geometry data into the dynamic geometry feature extraction model to obtain the high-dimensional feature vector of the dynamic geometry data; in: The preprocessing of the contact network dynamic geometric data collected by the sensor includes: data calibration, filtering, denoising and normalization operations; the calibration includes: correcting the sensor acquisition time and accuracy; the filtering and denoising include: removing high-frequency noise through low-pass filtering; the normalization includes: standardizing the contact force, current and vibration data to make their mean 0 and variance 1; In the dynamic geometric feature extraction model, for time series data, a long short-term memory network is used to capture the temporal dependencies in the data; for spatial distribution data, a convolutional neural network is used to extract the spatial features of the data; In the dynamic geometric feature extraction model, the dynamic geometric data X geom ∈R 100×3 , represents the three sensor features within 100 time steps: contact force, current, and vibration; The long short-term memory network includes two layers, each layer includes an input gate, a forget gate and an output gate, and the number of hidden units in each layer is 256. The dynamic geometric data is processed by the long short-term memory network, and the output feature vector is f geom ∈R 256 .
5. The method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: Using the defect text feature extraction model, the defect text features in the defect text data are extracted to obtain a high-dimensional feature vector of the defect text data, including the following steps: Preprocess the defect text data recorded in the defect report; Use the text encoder of the CLIP model to build a defect text feature extraction model; Input the preprocessed contact network defect text data into the defect text feature extraction model to obtain the high-dimensional feature vector of the defect text data; in: The preprocessing of the defect text data recorded in the defect report includes: text cleaning, word segmentation, stop word removal and stemming operations; the text cleaning includes: removing punctuation and unnecessary symbols in the ledger text, retaining key information, and processing spelling errors; the word segmentation includes: using the Jieba word segmentation tool to segment the text into words; The defect text feature extraction model learns the semantic information and contextual association of the text by pre-training on a large-scale text dataset.
6. The method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: The ternary contrast loss function includes: Among them, L(θ) is the ternary contrast loss function; θ is the model parameter; x, y, and z represent images, text, and dynamic geometric data, respectively; f, g, and h represent the embedding functions of images, text, and dynamic geometric data, respectively, and the embedding function maps the original input into a shared feature space; i is the corresponding modal data; j is the positive sample data of the other modality corresponding to i, which matches the modal data i; k is the negative sample data of the other modality corresponding to i, which does not match the modal data i; sim represents the similarity measurement function, which uses cosine similarity to calculate the similarity between two embedding vectors; m is a positive threshold used to control the interval between positive and negative sample pairs.
7. The method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: The constructing of the image sensor text positive and negative sample pairs includes constructing positive samples and constructing negative samples; The positive sample construction includes constructing a positive sample pair, wherein the positive sample pair refers to images, texts and dynamic geometry data of the same defect category; The negative sample construction includes constructing negative sample pairs and constructing hard negative samples. The negative sample pairs refer to images, text, and dynamic geometry data all belonging to different defect categories; the hard negative samples refer to two of the three types of data, namely, images, text, and dynamic geometry data, belonging to the same defect category, and the other one belonging to a different defect category.
8. The method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: The training of the feature fusion analysis model includes: For images, texts, and dynamic geometry data, image feature extraction models, defect text feature extraction models, and dynamic geometry feature extraction models are used for cleaning and feature extraction, and a fully connected layer is used to map images, texts, and dynamic geometry data to the same dimension R. 512 ; Construct image, text, and dynamic geometry data pairs through positive and negative sample pair construction; The ternary contrast loss function is calculated using image, text, and dynamic geometry data pairs. The gradient calculated by the ternary contrast loss function is back-propagated to each feature extraction model to update the parameters of each modal feature extraction model. The Adam optimizer is used to update the gradient, adjust the weights and biases of the feature fusion analysis model, and gradually optimize the accuracy of feature extraction and similarity calculation. The training process is optimized, including learning rate adjustment and batch training. In the learning rate adjustment, the initial learning rate is set to 1e-4 and gradually reduced through a learning rate decay strategy. In the batch training, in each training cycle, the model processes a batch of data, each batch includes 64 samples, and the samples in each batch are calculated using a ternary contrastive loss function to calculate the loss of positive and negative sample pairs, and then update the model. Hyperparameter adjustment includes adjusting the interval m in the ternary contrastive loss and adjusting the positive-negative sample ratio. In the interval m adjustment in the ternary contrastive loss, m is set to 0.2; in the positive-negative sample ratio adjustment, balanced sampling is used to adjust the distribution of positive and negative samples.
9. The method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: When the model is used downstream, the feature fusion analysis model learns the matching relationship between image, text, and sensor data through the optimization of the ternary contrast loss function. When querying, the user provides a modal data, and the feature fusion analysis model uses the embedding space learned by the ternary contrast loss to return other modal data that is most relevant to the query modality. This query method is used for defect detection, cross-modal information retrieval, and multimodal anomaly detection tasks.
10. The method for extracting and processing multi-source and multi-dimensional data features of a contact network according to claim 1, characterized in that: The model is provided for downstream use including: Given a query modality data, the query modality data includes image data, dynamic geometry data or defect text data of the contact network; The query modality data is converted into a feature vector using a feature extraction model; Calculate the similarity between the query modality feature vector and other modalities in the database; Based on similarity ranking, the most relevant images, texts, or sensor data are returned as query results.
Citation Information
Patent Citations
Electric power defect image detection method based on image-text question-answer multi-modal model
CN117763107A
Comparative learning unsupervised cross-modal hash retrieval algorithm based on graph attention mechanism
CN119377462A
Overhead line fault detection method based on multi-source three-mode data fusion network
CN119669890A