A cable defect monitoring and risk grading method based on multi-modal deep fusion
By employing a multimodal fusion method combining Mask R-CNN and Transformer, the problems of image distortion and uneven illumination in cable monitoring were solved, enabling accurate identification and risk assessment of internal cable faults and improving the accuracy and robustness of cable health status assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-21
AI Technical Summary
Existing cable monitoring technologies cannot effectively identify early faults that degrade the internal electrical performance of cables, and the deep physical correlation between multimodal data is not fully utilized, resulting in low defect identification accuracy and shallow fusion mechanisms.
Image defect quantification is performed using an instance segmentation model based on Mask R-CNN, and geometric parameters are extracted by combining adaptive illumination enhancement and wavelet threshold denoising. Multi-task decision-making is achieved through Transformer cross-modal attention mechanism and multi-layer perceptron network to realize multi-modal deep fusion.
It improves the ability to detect early and minor faults, enhances the accuracy of defect profile identification, improves the accuracy and robustness of cable health status assessment, and provides a complete risk assessment information chain.
Smart Images

Figure CN121350851B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system condition monitoring and predictive maintenance technology, and in particular to a method for cable defect monitoring and risk classification based on multimodal deep fusion. Background Technology
[0002] As a crucial component of modern power transmission networks, the health of power cables directly impacts the safety and stability of the entire power system. Especially in urban power grids, a large number of cables are laid in concealed and complex environments such as underground cable trenches and tunnels, facing multiple threats over the long term, including mechanical stress, insulation aging, localized overheating, chemical corrosion, and moisture intrusion. These factors inevitably lead to various defects in cables, which evolve over time and may ultimately cause serious power outages. Therefore, developing a technical system capable of effective online monitoring, early defect identification, and accurate health status assessment of cables has become a key technological challenge in the field of power system operation and maintenance. Currently, existing cable monitoring technologies can be mainly categorized into two independent modal-driven paradigms, but both have significant limitations.
[0003] One approach is image-based visual monitoring, which analyzes visible light or infrared images to identify physical damage to the cable surface, such as sheath damage, scratches, and corrosion. While this method is intuitive, its limitations are significant. First, underground environments often suffer from insufficient or uneven lighting, and dense cathodes, resulting in low image contrast and blurred details, severely impacting defect visibility. Second, close-up shots of cylindrical cable surfaces produce significant perspective distortion, misrepresenting the true shape and size of defects. More critically, visual monitoring is limited in its information dimension; it is essentially a morphological diagnosis and cannot directly perceive internal electrical performance degradation, such as critical early insulation defects caused by partial discharge.
[0004] The second approach involves monitoring electrical and physical quantities based on sensor time-series data. This method diagnoses the cable's operating status by analyzing voltage, current, temperature, or partial discharge signals. For example, distributed fiber optic temperature sensing (DTS) can effectively locate hot spots, while partial discharge monitoring is the most effective means of diagnosing internal insulation defects. However, this approach also has inherent technical limitations. Early faults (such as initial insulation aging and weak partial discharges) often exhibit extremely weak characteristics in electrical signals, easily masked by background noise and normal load fluctuations. Traditional signal processing methods struggle to effectively extract these non-stationary, transient fault characteristics. Furthermore, while sensor data can indicate the presence of anomalies, it is often difficult to precisely correlate them with specific physical defect morphologies and locations, resulting in an incomplete diagnosis.
[0005] Chinese invention patent CN120807509A discloses a deep learning-based computer-aided analysis method for medical images. The method includes receiving medical image data to be analyzed and performing metadata parsing and image sequence matching; preprocessing the medical image data to eliminate noise and artifacts and unify image format and pixel values; constructing a lightweight multi-scale feature extraction network to obtain deep semantic information and multi-scale image feature representations of the medical image data; based on the multi-scale image feature representation and a pre-trained deep learning model, identifying or segmenting lesions through a fine-grained lesion identification module, employing an adaptive regularization strategy to enhance the model's generalization ability during the identification process, and applying a multi-modal deep fusion mechanism to fully utilize the complementary information of different modalities of images, thereby generating preliminary analysis results; post-processing and evaluating the preliminary analysis results, and outputting a final computer-aided analysis report containing quantitative indicators, visualization content, and a structured report. The aforementioned invention constructs a lightweight multi-scale feature extraction network, based on multi-scale image feature representation and a pre-trained deep learning model. During the recognition process, an adaptive regularization strategy is employed to enhance the model's generalization ability. A multi-modal deep fusion mechanism is applied to fully utilize the complementary information from different modalities of imagery to generate preliminary analysis results, thus solving the problems of high computational complexity, large memory consumption, long training cycles, and difficulty in achieving efficient real-time processing. However, it still suffers from issues such as poor quality of the acquired raw data, weak early fault features that are difficult to extract effectively, low accuracy in defect contour recognition, underutilization of the deep physical correlation between multi-modal data, and a shallow fusion mechanism.
[0006] In summary, there is currently a lack of a cable defect monitoring and risk classification method based on multimodal deep fusion to solve or partially solve the aforementioned problems.
[0007] Compared to existing technologies, this invention, in terms of feature engineering, effectively addresses the problems of geometric distortion and uneven illumination through preprocessing methods such as cylinder unfolding correction and MSRCR illumination enhancement. It introduces Mask R-CNN instance segmentation, which can extract explicit geometric physical parameters (such as defect area and density) and combine them with two-dimensional time-frequency plot analysis. This significantly improves the quantification accuracy and interpretability of weak faults compared to existing technologies, and enhances the ability to capture transient features. In terms of fusion decision-making, this invention uses a Transformer cross-modal attention mechanism and dual-path fusion to replace simple feature splicing, and constructs a multi-task learning network that simultaneously outputs health index and risk level. This replaces the traditional method of calculating the health index using fixed formulas, significantly enhancing the model's robustness and adaptability under complex conditions. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a cable defect monitoring and risk classification method based on multimodal deep fusion, so as to solve or partially solve the problems of poor quality of acquired raw data, weak early fault features that are difficult to effectively extract, low accuracy of defect contour recognition, underutilization of deep physical correlation between multimodal data, and shallow fusion mechanism.
[0009] The objective of this invention can be achieved through the following technical solutions:
[0010] This invention provides a method for cable defect monitoring and risk classification based on multimodal deep fusion, specifically including the following steps:
[0011] S1. Acquire visual modal data and perform geometric correction and adaptive illumination enhancement processing; acquire sensor temporal modal data and perform standardization processing.
[0012] S2. Extract image defects from the preprocessed visual modality data in step S1, quantize the image defects into geometric parameters, and construct a geometric feature vector from the geometric parameters;
[0013] S3. Extract depth feature vectors representing the deep state of the device from each modal data;
[0014] S4. Concatenate the depth feature vectors of different modalities obtained in step S3, and perform cross-modal fusion on the concatenated vectors to output the final fused feature vector;
[0015] S5. The final fused feature vector obtained in step S4 is concatenated with the geometric feature vector obtained in step S2 to generate a combined vector. The combined vector is then input into a multi-task decision network based on a multilayer perceptron to obtain the diagnostic and quantitative evaluation results of the cable health status, thereby realizing the monitoring of cable defects and the evaluation of its health status.
[0016] As a preferred technical solution, in step S1, the calculation formula for the adaptive illumination enhancement processing is:
[0017]
[0018] In the formula, For the enhanced visual modal data in the first Color channels, coordinates Pixel value at; These are the pixel values of the geometrically corrected visual modal data in the corresponding channels and coordinates. The scale of the Gaussian wrap function used; For the first The weights of each scale satisfy... ; For the first A Gaussian wrapping function of scale * represents a two-dimensional Gaussian kernel; * represents a two-dimensional convolution operation.
[0019] As a preferred technical solution, the standardization processing step of the sensor time-series modal data in step S1 involves using wavelet threshold denoising to filter out high-frequency noise in the signal and performing maximum-minimum normalization processing. The normalization calculation formula is as follows:
[0020]
[0021] In the formula, For the normalized first The value of each data point; For the denoised signal sequence, the first... The value of each data point; and These are the minimum and maximum values in the signal sequence, respectively.
[0022] As a preferred technical solution, the standardized visual modal data is input into an instance segmentation model based on the Mask R-CNN architecture. The model identifies defects in the standardized visual modal data and outputs its category, bounding box, and binary segmentation mask. The geometric parameters are then quantized based on this binary segmentation mask. The mathematical expression for the binary segmentation mask is:
[0023]
[0024] The binary segmentation mask is a two-dimensional matrix with elements... This defines whether each pixel within the bounding box belongs to a defect region, and the matrix size is... , and These are the height and width of the bounding box, respectively. For the mask matrix in coordinates The value at that location, , These are the pixel coordinates within the bounding box. , .
[0025] As a preferred technical solution, the geometric parameters include actual area, length, width, and shape descriptor.
[0026] The formula for calculating the actual area is:
[0027]
[0028] In the formula, This represents the actual physical area of the defect. The actual physical area represented by a single pixel. For the mask matrix in coordinates The value at that location, , These are the pixel coordinates within the bounding box, expressed in square millimeters per pixel. and These are the height and width of the bounding box, respectively.
[0029] The length and width are respectively the segmentation mask. The longer and shorter sides of the smallest bounding rectangle of the defined region;
[0030] The shape descriptor is calculated using density:
[0031]
[0032] In the formula, For density, its range is (0, 1], and the closer the value is to 1, the closer the shape is to a circle. This represents the actual physical area of the defect. Let be the perimeter of the defect profile.
[0033] As a preferred technical solution, the extraction of depth feature vectors for each modality data in step S3 specifically includes:
[0034] For visual modal data, the regions of the segmentation mask are mapped onto the deep feature map of the backbone network of the instance segmentation model based on region of interest alignment, and the features within the region are pooled to obtain a fixed-length visual depth feature vector.
[0035] For sensor time-series modal data, the time-series signal preprocessed in step S1 is converted into a two-dimensional time-spectrum map through continuous wavelet transform. The two-dimensional time-spectrum map is then input into a lightweight convolutional neural network based on the MobileNetV3 architecture to obtain the sensor depth feature vector.
[0036] As a preferred technical solution, the transformation process of converting the time-series signal into a two-dimensional time spectrum through continuous wavelet transform is as follows:
[0037]
[0038] In the formula, The wavelet transform coefficients form the matrix elements of the two-dimensional time-spectrum graph. The input is a normalized one-dimensional time-series signal. The mother wavelet function, specifically the Morlet wavelet. For the complex conjugate of the mother wavelet function, The scaling factor is inversely proportional to the frequency. The shift factor represents the time position. The independent variable is time.
[0039] As a preferred technical solution, the cross-modal fusion in step S4 is based on the Transformer encoder layer to achieve deep fusion of vectors. Through the internal multi-head self-attention mechanism, the internal dependencies between the two features are calculated and modeled. By assigning weights to the features or feature combinations, the final fused feature vector is output. The calculation process of the internal dependencies between the features is as follows:
[0040]
[0041] In the formula, For the input matrix, For querying the matrix, The key matrix, For value matrices, The weight matrix is a learnable linear transformation. The dimensions of the query and key vectors;
[0042] The calculation process for a single attention head in the multi-head self-attention mechanism is as follows:
[0043]
[0044] In the formula, The similarity score between the query and the key. As a scaling factor, The function normalizes the similarity score along the key dimension. For querying the matrix, The key matrix, Let T be a value matrix, and T denotes the transpose.
[0045] As a preferred technical solution, the combined vector is represented as follows:
[0046]
[0047] In the formula, For combined vectors, The final fused feature vector obtained in step S4 The geometric feature vector obtained in step S2, and These are the dimensions of the two vectors, respectively.
[0048] As a preferred technical solution, in step S5, the multi-task decision network simultaneously executes the health index regression branch task and the risk level classification branch task through its internal parallel output branches, and outputs the health index and risk level:
[0049] The health index regression branch consists of a linear output layer of a single neuron, and the calculation formula is as follows:
[0050]
[0051] In the formula, This is a health index, with a value range of [0, 100]. The output feature vectors of the shared hidden layer, ,in and The first The weight matrix and bias vector of a shared hidden layer Its corresponding activation function is The final combined vector is a concatenation of the eigenvectors and the geometric eigenvectors. The number of shared hidden layers; and These are the learnable weight matrix and bias term for the health index regression branch, respectively. The activation function is Sigmoid; the health index regression branch uses the mean squared error loss function. Optimize for a given set of elements The loss for each training batch of samples is calculated as follows:
[0052]
[0053] In the formula, It is the first Predicted health index for each sample It is its corresponding real health index label. The number of samples in the training batch;
[0054] The risk level classification branch consists of a... It consists of a fully connected layer with 10 neurons and a Softmax activation function, where The formula for calculating the number of predefined risk levels is as follows:
[0055]
[0056] In the formula, For each level, the score vector and These are the learnable weight matrix and bias term for the risk level classification branch, respectively. The output feature vectors of the shared hidden layer;
[0057] The Softmax function is calculated as follows:
[0058]
[0059] The Softmax function converts the score vector Convert to a dimensional probability vector , The cable was assessed as being at risk level. The probability, This is the index for the target risk level currently being calculated. This is the traversal index in the denominator used for normalized summation;
[0060]
[0061] In the formula, The final output should represent the clearly defined risk level with the highest probability. This indicates that the cable has been assessed as having a risk level. The probability and risk level classification branch use the cross-entropy loss function. Optimize:
[0062]
[0063] In the formula, It is an indicator variable, if the sample The true category is If it is 1, then it is 1; otherwise it is 0. The model predicts the sample. Category The probability, The number of samples in the training batch;
[0064] The total loss function of the entire multi-task decision network It is a weighted sum of the losses from the two tasks:
[0065]
[0066] In the formula, and These are hyperparameters used to balance the importance of the two tasks. Let the mean squared error loss function be . This is the cross-entropy loss function.
[0067] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0068] 1) Enhanced identifiability of early potential fault features: This invention solves the problems of uneven image illumination and geometric distortion by combining geometric correction and adaptive illumination enhancement processing; by converting sensor time-series modal data into a two-dimensional time-frequency domain and using a convolutional neural network to extract depth features, it solves the problem of difficulty in extracting non-stationary and transient mode features that are highly correlated with faults, and achieves the technical effect of correcting image distortion and improving image quality, enhancing the detection capability of early weak and low-contrast defects, and improving the monitoring sensitivity of early electrical or thermal anomalies.
[0069] 2) Improved accuracy of defect contours: This invention adopts an instance segmentation model based on the Mask R-CNN architecture, which outputs the category, bounding box and accurate binary segmentation mask for each identified defect. By using the mask obtained from segmentation and combining it with the proposed geometric parameter calculation method, the key physical dimensions such as defect area and length are automatically quantified, and two-dimensional geometric parameters are obtained to form a geometric feature vector. This solves the problems of blurred bounding boxes and low recognition accuracy in risk assessment in the prior art, and improves the reliability and scientificity of the assessment results.
[0070] 3) Coordinated reasoning of multimodal information: This invention constructs a cross-modal attention fusion mechanism based on Transformer, and uses its internal self-attention mechanism to deeply fuse and splice feature vectors of different modalities and coordinate information to generate a final fused feature vector containing deep cross-modal correlation information. This solves the problem that a single indicator may cause misjudgment in the prior art, and improves the accuracy and robustness of fault diagnosis results in complex scenarios.
[0071] 4) Quantitative Health Status Assessment and Risk Classification System: This invention utilizes deeply fused multimodal features and a multi-task decision network based on a multilayer perceptron to comprehensively assess the health status and defect risk level of cables. This decision network performs two tasks simultaneously—health index regression and risk level classification—through parallel output branches, ultimately outputting a quantified health index and a clear risk level. This solves the problem of low reliability in predictive maintenance of cables, achieves the technical effect of improving the generalization ability and overall performance of the model, and provides a complete information chain from macro-level status (health index) to specific action instructions (risk level) for operation and maintenance decisions, enhancing the practicality and credibility of the method. Attached Figure Description
[0072] Figure 1 This is a schematic diagram of the present invention. Detailed Implementation
[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0074] Example 1
[0075] To address the problems existing in the aforementioned prior art, this embodiment provides a cable defect monitoring and risk classification method based on multimodal deep fusion. This method introduces image segmentation to achieve accurate defect quantification and combines a cross-modal attention mechanism to enhance the collaborative reasoning ability for multi-source heterogeneous information. Its core lies in constructing a deep learning model capable of effectively addressing complex working conditions and performing objective quantitative assessments from a data perspective. Figure 1 As shown, the acquired cable image data (visual modal data) and sensor time-series modal data are preprocessed and feature extracted according to their respective modalities to obtain depth feature vectors for each modality. Subsequently, feature fusion is performed on the obtained modal feature vectors, and health status assessment and risk level classification are carried out based on the fusion results. The specific steps include:
[0076] S1. Multimodal Data Preprocessing. This step standardizes the raw cable image data and sensor time-series modal data to eliminate environmental interference and provide high-quality input for subsequent feature extraction. For the input cable image data, a cylindrical surface unfolding algorithm is first used for geometric correction to eliminate perspective distortion caused by the cylindrical curved surface structure of the cable. Subsequently, an image enhancement algorithm based on multi-scale Retinex theory is used for adaptive illumination enhancement to suppress the influence of uneven illumination and output a high-quality normalized image. For the input sensor time-series modal data, wavelet thresholding is used to filter out high-frequency noise in the signal, and max-min normalization is performed to eliminate the influence of dimensions. This preprocessing step aims to eliminate systematic biases and random noise introduced by the acquisition environment and the characteristics of the equipment itself, transforming the raw data into a consistent and robust high-quality input, thus laying a solid data foundation for subsequent deep feature extraction, fusion, and decision analysis. This step specifically includes two parallel processing flows for image modalities and sensor time-series modalities.
[0077] 1. Image data normalization preprocessing
[0078] To address the two major technical challenges of geometric distortion and uneven illumination commonly found in cable image data, this invention designs a serial processing flow that includes geometric correction and illumination enhancement.
[0079] First, geometric correction is performed to eliminate perspective distortion caused by the cylindrical curved surface structure of the cable and its projection in the two-dimensional image. This distortion distorts the true shape and size of defects in the image, severely affecting the accuracy of subsequent identification and quantization. This invention employs a cylindrical surface unfolding algorithm to achieve this correction. The first stage of this algorithm uses the Hough Transform technique to accurately identify the straight lines of the cable's contour boundary in the image. Subsequently, through geometric analysis of the detected straight lines, the equation of the cable's central axis and radius are calculated. After obtaining these geometric parameters, the second stage of the algorithm establishes the original image pixel coordinate system based on the inverse transformation model of cylindrical projection. Unfolded two-dimensional plane coordinate system The mapping relationship between them. Specifically, for any pixel in the original image... The corresponding cylindrical coordinates can be calculated. and height along the axis Then, coordinate transformation is performed using the following formula:
[0080]
[0081] in: These are the planar coordinates of the unfolded image; The calculated cable radius; Original pixels The angle corresponding to the center of the cable cross-section; Original pixels The projection position along the central axis of the cable. By traversing all relevant pixels and interpolating, an orthophoto image is finally generated, which is geometrically similar to a slice and flattened image along the cable axis. .
[0082] Secondly, after geometric correction, adaptive illumination enhancement is performed to address common problems in underground laying environments such as insufficient lighting, uneven lighting distribution, heavy shadows, and localized highlights. This invention employs a multi-scale Retinex theory-based image enhancement algorithm (MSRCR) to handle this problem. Retinex theory assumes that an image can be decomposed into illumination and reflection components, where the reflection component represents the inherent properties of the object. The MSRCR algorithm enhances the reflection component, which represents object details, by operating in the logarithmic domain, estimating and removing the slowly varying illumination component. For the first... Each color channel is located at coordinates pixel value at Its enhanced output value It is composed of a weighted combination of Gaussian surround function filtering results at multiple different scales, and its core calculation formula is as follows:
[0083]
[0084] in: To enhance the image in the first... Color channels, coordinates Pixel value at; This represents the pixel values of the original (geometrically corrected) image in the corresponding channel and coordinates. The scale of the Gaussian wrap function used; For the first The weights of each scale satisfy... ; For the first A Gaussian wrapping function of scale * represents a two-dimensional Gaussian kernel; * represents a two-dimensional convolution operation.
[0085] This step ultimately outputs a normalized image with uniform lighting and clear details. .
[0086] 2. Standardization preprocessing of sensor time-series data
[0087] For the collected one-dimensional sensor time-series data such as voltage, current, and temperature, it is necessary to perform noise reduction and normalization processing.
[0088] First, signal denoising is performed. This invention employs wavelet thresholding denoising. This method first processes the original signal... Multi-level wavelet decomposition is performed to obtain detail coefficients at different scales. Then, a soft thresholding function is applied to these detail coefficients to effectively suppress noise while preserving signal edge features. Finally, the denoised signal is reconstructed using inverse wavelet transform.
[0089] Secondly, after denoising, data normalization is performed to eliminate the influence of different physical quantities' dimensions. This invention employs the Min-Max Normalization method to normalize the signal sequence. Each data point in It is linearly mapped to the interval [-1, 1]. Its normalization formula is:
[0090]
[0091] in: For the normalized first The value of each data point; The first (denoised) signal in the original (denoised) signal sequence The value of each data point; and These are the signal sequences. The minimum and maximum values in.
[0092] This step ensures that sensor data from different sources are equally important when fed into subsequent models, avoiding model training bias caused by large differences in numerical ranges.
[0093] S2. Image Defect Segmentation and Quantization. This step accurately extracts the contours and geometric parameters of defects from the preprocessed image, achieving a leap from qualitative identification to quantitative analysis. The normalized image obtained in step S1 is input into an instance segmentation model based on the Mask R-CNN architecture. This model simultaneously locates, classifies, and segments defects in the image at the pixel level, outputting the category, bounding box, and precise binary segmentation mask for each identified defect. Subsequently, based on this binary segmentation mask, a series of image morphological calculations are performed to accurately quantize the two-dimensional geometric parameters of each defect, including the actual area, length, width, and shape descriptor, and these parameters are used to construct a geometric feature vector. This step constitutes the core technology for upgrading from traditional defect identification to precise geometric quantization, providing a solid and structured data foundation for subsequent deep feature extraction and risk assessment.
[0094] The core of this step lies in applying an instance segmentation model based on the Mask R-CNN architecture. Unlike object detection models that can only provide rectangular bounding boxes, the instance segmentation model can simultaneously locate, classify, and segment multiple defect instances in an image at the pixel level. Specifically, the normalized image obtained in step S1... The input is fed into a pre-trained Mask R-CNN model. This model contains a backbone network for extracting deep image features, a Region Proposal Network (RPN) for generating candidate regions, and a detection head that performs classification, bounding box regression, and mask prediction for each candidate region. For each identified defect instance in the image, the model ultimately outputs three key pieces of information: the semantic category of the defect, a precise rectangular bounding box, and a binary segmentation mask of the same size as the bounding box. The mask is a two-dimensional matrix with dimensions of . ,in and These represent the height and width of the bounding box, respectively. Elements in the mask matrix. The mathematical expression for determining whether each pixel within the bounding box belongs to a defect region is as follows:
[0095]
[0096] Where: the binary segmentation mask is a two-dimensional matrix, with elements... This defines whether each pixel within the bounding box belongs to a defect region, and the matrix size is... , and These are the height and width of the bounding box, respectively. For the mask matrix in coordinates The value at that location, , These are the pixel coordinates within the bounding box. , .
[0097] This binary segmentation mask depicts the true outline of the defect on a two-dimensional plane with pixel-level precision, and serves as the direct basis for the quantization of all subsequent geometric parameters.
[0098] Obtain the binary segmentation mask for each defect. Subsequently, this invention uses a series of image morphological calculations to precisely quantify the two-dimensional geometric parameters of the defect. These calculations are all performed on a mask matrix and specifically include the following aspects:
[0099] First, calculate the area of the defect. The actual area of the defect. By counting all pixels with a value of 1 in the mask matrix and multiplying by the physical area represented by a unit pixel, which is predetermined by camera calibration or image resolution. The calculation can be expressed by the following formula:
[0100]
[0101] in: The actual physical area of the defect; The actual physical area represented by a single pixel. For the mask matrix in coordinates The value at that location, , These are the pixel coordinates within the bounding box, expressed in square millimeters per pixel. and These represent the height and width of the bounding box, respectively.
[0102] Secondly, the main dimensions of the defect, namely its length and width, are calculated. This is done by calculating the segmentation mask. The minimum area bounding box of the defined region is used to determine this. The algorithm finds a rotated rectangle that completely encloses the defect contour and has the smallest area; the longer side of this rectangle is defined as the length of the defect. The shorter side is defined as the width of the defect. .
[0103] Next, to further describe the complexity of the defect morphology, other shape descriptors are calculated. The perimeter of the defect profile is calculated. This refers to the total length of the boundary pixels of the mask region. Based on the area and perimeter, a dimensionless shape parameter, density, can be calculated. , used to measure the roundness or regularity of a shape, is defined as follows:
[0104]
[0105] in: For density, its range is (0, 1], and the closer the value is to 1, the closer the shape is to a circle; and These are the calculated defect area and perimeter, respectively.
[0106] Through the above series of calculations, a result is generated for each detected defect, including its area. ,length ,width and density Geometric eigenvectors of multiple quantitative indicators This vector objectively and quantitatively describes the two-dimensional physical form of the defect, providing crucial input information for subsequent risk assessment that differs from traditional methods.
[0107] S3. Multimodal Deep Feature Extraction. This step extracts abstract feature representations that characterize the deep state of the device from each modality of data. For the visual modality, Region of Interest (ROI) Alignment is used to accurately map the segmentation mask region obtained in step S2 back onto the deep feature map of the backbone network of the instance segmentation model, and pooling is performed on the features within this region to generate a fixed-length visual depth feature vector. For the sensor temporal modality, the preprocessed temporal signal in step S1 is converted into a two-dimensional temporal spectrogram through Continuous Wavelet Transform (CWT), and then this spectrogram is input into a lightweight convolutional neural network based on the MobileNetV3 architecture to extract sensor depth feature vectors that characterize the dynamic patterns of the signal. This step aims to transform the standardized, information-rich input data into a highly condensed latent feature space that is more conducive to subsequent fusion and decision-making through a deep neural network. This step processes the visual modality and the sensor temporal modality in parallel.
[0108] For visual modalities, the goal of deep feature extraction is to generate a feature vector that not only contains semantic information about the defect but also is highly correlated with the precise spatial extent of the defect. To achieve this goal, this invention utilizes the deep feature map generated by the instance segmentation model in step S2 and performs feature extraction using a precise feature pooling technique, namely Region of Interest Align (ROI Align). This technique uses the binary segmentation mask obtained in step S2, which represents the precise contour of the defect, to extract features. It accurately maps back to the last convolutional feature map output by the backbone network of the instance segmentation model. Above. ROI Align avoids quantization rounding operations through bilinear interpolation, thus achieving lossless alignment between the mask boundaries and the feature map. After alignment, only the mask is aligned. The feature vectors within the covered feature map region are aggregated. This invention uses average pooling as the aggregation function, calculating the mean of all feature vectors within the region to ultimately generate a fixed-length, highly condensed visual depth feature vector. This vector not only contains high-level semantic information such as the defect category and texture, but its generation process is also strictly limited to the precise pixel range of the defect, effectively eliminating interference from background information.
[0109] For sensor time-series modes, the goal of deep feature extraction is to capture non-stationary, transient dynamic modes related to potential faults from one-dimensional time-series signals. To achieve this goal, this invention uses a method combining time-frequency analysis and convolutional neural networks for feature extraction. First, the sensor time-series signal preprocessed in step S1 is... Time-frequency analysis is performed using Continuous Wavelet Transform (CWT). CWT can provide localized information of a signal at different time and frequency scales, and is particularly suitable for analyzing non-stationary signals. The transformation process is defined by the following equation, which transforms a one-dimensional signal into a two-dimensional time-frequency spectrum. :
[0110]
[0111] in: These are wavelet transform coefficients, forming the matrix elements of the two-dimensional time-spectrum graph; The input is a normalized one-dimensional time-series signal; The Morlet wavelet is selected in this invention to balance the time-frequency resolution, serving as the mother wavelet function. The complex conjugate of the mother wavelet function; This is a scaling factor, inversely proportional to frequency; The shift factor represents the time position. The independent variable is time.
[0112] This transformation ultimately generates a two-dimensional time-frequency spectrum. This figure visually illustrates the distribution of signal energy in time and frequency. Secondly, the time-frequency spectrum is presented. The image is treated as a special type of image and input into a lightweight convolutional neural network based on the MobileNetV3 architecture. This network is specially trained so that its convolutional layers can automatically learn and extract key patterns from the time-spectral image, such as vertical high-frequency pulse bands corresponding to partial discharge and specific harmonic energy clusters related to overheating. Through forward propagation, the network's global average pooling layer outputs a fixed-length sensor depth feature vector that characterizes the overall dynamic pattern of the signal. This vector encodes the time-frequency domain information most relevant to fault diagnosis in the original time-series signal in a compact form.
[0113] S4. Cross-modal attention fusion. This step performs deep interaction and information synergy on the different modal features obtained from step S3 to generate a more discriminative unified representation. The visual depth feature vector and the sensor depth feature vector are concatenated, and the concatenated vector is input into a cross-modal fusion module based on a Transformer encoder layer. This module, through its internal self-attention mechanism, calculates and models the interdependencies within and between the two feature sets. By dynamically and data-drivenly assigning higher weights to more important features or feature combinations, it ultimately outputs a final fused feature vector containing deep cross-modal correlation information. This step aims to overcome the limitations of traditional feature concatenation methods in fully exploring deep intermodal correlations by introducing an attention mechanism to achieve synergistic information enhancement.
[0114] Specifically, the module receives the visual depth feature vector from step S3. and sensor depth feature vector As input, where Let be the dimension of the feature vectors. To enable the model to distinguish and process these two types of features from different sources, a learnable modality type embedding is first appended to each feature vector. Then, the two enhanced feature vectors are concatenated along the sequence dimension to form an input matrix. This input matrix is then fed into the Transformer encoder layer for deep fusion processing. The core component of this encoder layer is the multi-head self-attention mechanism, which computes the dependencies between features by linearly mapping the input features to three different representation spaces: query, key, and value. For the input matrix... Its query matrix Key matrix Sum matrix It is generated by the following formula:
[0115]
[0116] in: For the input matrix, For querying the matrix, The key matrix, For value matrices, The weight matrix is a learnable linear transformation. The dimensions of the query and key vectors.
[0117] Subsequently, the weighted feature representation is obtained by calculating the scaled dot product attention. The calculation process of a single attention head is defined by the following formula:
[0118]
[0119] in: The similarity score between the query and the key was calculated; It is a scaling factor used to prevent the gradient from vanishing due to an excessively large dot product result; The function normalizes the similarity scores along the key dimension, generating attention weights that represent the importance of each element in the value vector. For querying the matrix, The key matrix, Let T be a value matrix, and T denotes the transpose.
[0120] Multi-head attention mechanisms execute in parallel. The model performs independent attention calculations, concatenates the outputs of each head, and then performs a final linear transformation to obtain the aggregated output. Through the self-attention mechanism, the model can automatically and data-drivenly learn the complex dependencies between visual features and sensor features within their respective dimensions, as well as between cross-modal feature dimensions. Following the self-attention sub-layer, the Transformer encoder layer includes a feed-forward network sub-layer, along with corresponding residual connections and layer normalization to promote stable training and deep representation of the model. After processing by this fusion module, the output is a matrix of the same size as the input. Finally, average pooling is performed on the two feature vectors in the output matrix to generate a fixed-length final fused feature vector that contains deep cross-modal correlation information. Compared to simple feature concatenation, this vector has stronger representational power and robustness, providing high-quality input for subsequent health status assessment.
[0121] S5. Health Status Assessment and Risk Classification. This step transforms the fused features into a final diagnostic conclusion and quantitative assessment of the cable's health status. The final fused feature vector obtained in step S4 is concatenated with the geometric feature vector obtained in step S2, and this combined vector is input into a multi-task decision network based on a multilayer perceptron (MLP). This decision network, through its parallel output branches, simultaneously performs two tasks: health index regression and risk level classification, ultimately outputting a quantitative health index and a clear risk level, providing a scientific and reliable basis for predictive cable maintenance. This step aims to construct a decision model capable of integrating multi-dimensional information to output objective, clear, and practically guiding assessment results. The core of this step lies in applying a multi-task learning decision network based on a multilayer perceptron (MLP). This network can improve the model's generalization ability and overall performance by sharing low-level feature representations and simultaneously performing two related tasks: health index regression and risk level classification.
[0122] Specifically, to ensure that the decision-making model can utilize the most comprehensive information, this invention first constructs a combined input vector. This vector is the final fused feature vector obtained in step S4, which contains deep cross-modal correlation information. Compared with the geometric feature vectors describing the two-dimensional physical morphology of the defect obtained in step S2 It is assembled by splicing along the feature dimension, where and Let be the dimensions of the two vectors. The combined vector can be represented as:
[0123]
[0124] This combined vector The data is then fed into a multi-task decision network. This network consists of several shared fully connected hidden layers and a separate task-specific output layer. The shared hidden layers are responsible for learning a higher-level abstract feature representation that can serve two tasks simultaneously. This process can be formally represented as:
[0125]
[0126] in and They are the first The weight matrix and bias vector of a shared hidden layer It is its corresponding activation function. The final combined vector is a concatenation of the eigenvectors and the geometric eigenvectors. The number of shared hidden layers. Task-specific output layers are located in... Based on this, it includes the following two parallel output branches:
[0127] 1. Health Index (HI) Regression Branch
[0128] This branch consists of a linear output layer with a single neuron, and its goal is to represent shared features. It reverts to a standardized continuous numerical range to output a quantified health index. This index is used to comprehensively reflect the overall health status of the cable. Its calculation can be expressed as:
[0129]
[0130] in: The health index has a value range of [0, 100]. The output feature vectors of the shared hidden layer; and These are the learnable weight matrix and bias term for this regression branch, respectively; The Sigmoid activation function is defined as follows: The output value is constrained to the interval (0, 1), and then multiplied by 100 for scaling. During model training, this branch uses the Mean Squared Error (MSE) loss function. Optimize. For a given... The loss for each training batch of samples is calculated as follows:
[0131]
[0132] in: It is the first Predicted health index for each sample It is its corresponding real health index label. This represents the number of samples in the training batch.
[0133] 2. Risk Grade Classification Branches
[0134] This branch consists of a containing It consists of a fully connected layer with 10 neurons and a Softmax activation function, where The number of predefined risk levels. Its goal is to represent shared features. This is mapped to a probability distribution of a predefined set of risk levels. This branch first computes the score vector for each level through a linear transformation. :
[0135]
[0136] in and These are the learnable weight matrix and bias term for this classification branch, respectively. This is the output feature vector of the shared hidden layer. Then, the Softmax function converts the score vector... Convert to a dimensional probability vector , This indicates that the cable has been assessed as having a risk level. The probability, This is the index for the target risk level currently being calculated. The Softmax function is calculated as follows, using the traversal index in the denominator for normalized summation:
[0137]
[0138] The final output is a clear risk level. The level that is identified as having the highest probability, i.e. During the model training phase, this branch employs the cross-entropy loss function. Optimize:
[0139]
[0140] in It is an indicator variable (if the sample) The true category is If it is 1, then it is 1; otherwise it is 0. The model predicts the sample. Category The probability, This represents the number of samples in the training batch.
[0141] The total loss function of the entire multi-task decision network It is a weighted sum of the losses from the two tasks:
[0142]
[0143] in and It is a hyperparameter used to balance the importance of two tasks. Let the mean squared error loss function be . Let be the cross-entropy loss function. By minimizing this total loss function, the model can simultaneously learn two tasks: health index regression and risk level classification, thereby achieving a comprehensive and quantitative assessment of the cable's health status.
[0144] Example 2
[0145] Based on the foregoing embodiments, this embodiment provides an electronic device, including a processor, an internal bus, a network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the aforementioned method for cable defect monitoring and risk classification based on multimodal deep fusion. Of course, in addition to software implementation, this invention does not exclude other implementation methods, such as logic devices or a combination of hardware and software, that is, the execution entity of the processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0146] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for cable defect monitoring and risk classification based on multimodal deep fusion, characterized in that, Specifically, the following steps are included: S1. Acquire visual modal data and perform geometric correction and adaptive illumination enhancement processing; acquire sensor temporal modal data and perform standardization processing. S2. Extract image defects from the preprocessed visual modality data in step S1, quantize the image defects into geometric parameters, and construct a geometric feature vector from the geometric parameters; The standardized visual modal data is input into an instance segmentation model based on the Mask R-CNN architecture. The model identifies defects in the standardized visual modal data and outputs its category, bounding box, and binary segmentation mask. The geometric parameters are then quantized based on this binary segmentation mask. The mathematical expression for the binary segmentation mask is as follows: The binary segmentation mask is a two-dimensional matrix with elements... This defines whether each pixel within the bounding box belongs to a defect region, and the matrix size is... , and These are the height and width of the bounding box, respectively. For the mask matrix in coordinates The value at that location, , These are the pixel coordinates within the bounding box. , ; S3. Extract depth feature vectors representing the deep state of the device from each modal data; For visual modal data, the regions of the segmentation mask are mapped onto the deep feature map of the backbone network of the instance segmentation model based on region of interest alignment, and the features within the region are pooled to obtain a fixed-length visual depth feature vector. For sensor time-series modal data, the time-series signal preprocessed in step S1 is converted into a two-dimensional time-spectrum map through continuous wavelet transform. The two-dimensional time-spectrum map is then input into a lightweight convolutional neural network based on the MobileNetV3 architecture to obtain the sensor depth feature vector. S4. Concatenate the depth feature vectors of different modalities obtained in step S3, and perform cross-modal fusion on the concatenated vectors to output the final fused feature vector; S5. The final fused feature vector obtained in step S4 is concatenated with the geometric feature vector obtained in step S2 to generate a combined vector. The combined vector is then input into a multi-task decision network based on a multilayer perceptron to obtain the diagnostic and quantitative evaluation results of the cable health status, thereby realizing the monitoring of cable defects and the evaluation of its health status.
2. The cable defect monitoring and risk classification method based on multimodal deep fusion according to claim 1, characterized in that, In step S1, the calculation formula for the adaptive illumination enhancement processing is: In the formula, For the enhanced visual modal data in the first Color channels, coordinates Pixel value at; These are the pixel values of the geometrically corrected visual modal data in the corresponding channels and coordinates. The scale of the Gaussian wrap function used; For the first The weights of each scale satisfy... ; For the first A Gaussian wrapping function of scale * represents a two-dimensional Gaussian kernel; * represents a two-dimensional convolution operation.
3. The cable defect monitoring and risk classification method based on multimodal deep fusion according to claim 1, characterized in that, The standardization process for the sensor time-series modal data in step S1 involves using wavelet threshold denoising to filter out high-frequency noise in the signal and performing max-min normalization. The normalization calculation formula is as follows: In the formula, For the normalized first The value of each data point; For the denoised signal sequence, the first... The value of each data point; and These are the minimum and maximum values in the signal sequence, respectively.
4. The cable defect monitoring and risk classification method based on multimodal deep fusion according to claim 1, characterized in that, The geometric parameters include the actual area, length, width, and shape descriptor. The formula for calculating the actual area is: In the formula, This represents the actual physical area of the defect. The actual physical area represented by a single pixel. For the mask matrix in coordinates The value at that location, , These are the pixel coordinates within the bounding box, expressed in square millimeters per pixel. and These are the height and width of the bounding box, respectively. The length and width are respectively the segmentation mask. The longer and shorter sides of the smallest bounding rectangle of the defined region; The shape descriptor is calculated using density: In the formula, For density, its range is (0, 1], and the closer the value is to 1, the closer the shape is to a circle. This represents the actual physical area of the defect. Let be the perimeter of the defect profile.
5. The cable defect monitoring and risk classification method based on multimodal deep fusion according to claim 1, characterized in that, The transformation process of the time-series signal into a two-dimensional time spectrum through continuous wavelet transform is as follows: In the formula, The wavelet transform coefficients form the matrix elements of the two-dimensional time-spectrum graph. The input is a normalized one-dimensional time-series signal. The mother wavelet function, specifically the Morlet wavelet. For the complex conjugate of the mother wavelet function, The scaling factor is inversely proportional to the frequency. The shift factor represents the time position. The independent variable is time.
6. The cable defect monitoring and risk classification method based on multimodal deep fusion according to claim 1, characterized in that, In step S4, cross-modal fusion is based on the Transformer encoder layer to achieve deep fusion of vectors. Through the internal multi-head self-attention mechanism, the internal dependencies between the two features are calculated and modeled. By assigning weights to features or feature combinations, the final fused feature vector is output. The calculation process of the internal dependencies between features is as follows: In the formula, For the input matrix, For querying the matrix, The key matrix, For value matrices, The weight matrix is a learnable linear transformation. The dimension of the feature vector. The dimensions of the query and key vectors; The calculation process for a single attention head in the multi-head self-attention mechanism is as follows: In the formula, The similarity score between the query and the key. As a scaling factor, The function normalizes the similarity score along the key dimension. For querying the matrix, The key matrix, Let T be a value matrix, and T denotes the transpose.
7. The cable defect monitoring and risk classification method based on multimodal deep fusion according to claim 1, characterized in that, The combined vector is represented as follows: In the formula, For combined vectors, The final fused feature vector obtained in step S4 The geometric feature vector obtained in step S2, and These are the dimensions of the two vectors, respectively.
8. The cable defect monitoring and risk classification method based on multimodal deep fusion according to claim 1, characterized in that, In step S5, the multi-task decision network simultaneously executes the health index regression branch task and the risk level classification branch task through its internal parallel output branches, outputting the health index and risk level: The health index regression branch consists of a linear output layer of a single neuron, and the calculation formula is as follows: In the formula, This is a health index, with a value range of [0, 100]. The output feature vectors of the shared hidden layer, ,in and The first The weight matrix and bias vector of a shared hidden layer Its corresponding activation function is The final combined vector is a concatenation of the eigenvectors and the geometric eigenvectors. The number of shared hidden layers; and These are the learnable weight matrix and bias term for the health index regression branch, respectively. The activation function is Sigmoid; the health index regression branch uses the mean squared error loss function. Optimize for a given set of elements The loss for each training batch of samples is calculated as follows: In the formula, It is the first Predicted health index for each sample It is its corresponding real health index label. The number of samples in the training batch; The risk level classification branch consists of a... It consists of a fully connected layer with 10 neurons and a Softmax activation function, where The formula for calculating the number of predefined risk levels is as follows: In the formula, For each level, the score vector and These are the learnable weight matrix and bias term for the risk level classification branch, respectively. The output feature vectors of the shared hidden layer; The Softmax function is calculated as follows: The Softmax function converts the score vector Convert to a dimensional probability vector , The cable was assessed as being at risk level. The probability, This is the index for the target risk level currently being calculated. This is the traversal index in the denominator used for normalized summation; In the formula, The final output should represent the clearly defined risk level with the highest probability. This indicates that the cable has been assessed as having a risk level. The probability and risk level classification branch use the cross-entropy loss function. Optimize: In the formula, It is an indicator variable, if the sample The true category is If it is 1, then it is 1; otherwise it is 0. The model predicts the sample. Category The probability, The number of samples in the training batch; The total loss function of the entire multi-task decision network It is a weighted sum of the losses from the two tasks: In the formula, and These are hyperparameters used to balance the importance of the two tasks. Let the mean squared error loss function be . This is the cross-entropy loss function.
Citation Information
Patent Citations
Medical image computer-aided analysis method based on deep learning
CN120807509A
Composite material processing surface defect detection and evaluation method based on deep learning
CN115661071A
Robot inspection system and method based on task instruction
CN119942278A
Cable fault monitoring and health state assessment method based on multi-modal fusion
CN120334805A
Vehicle surrounding structure fault prediction maintenance method and system based on machine learning
CN120493716A