Cable defect diagnosis method and device based on asynchronous contrast learning, equipment and medium

CN122365406BActive Publication Date: 2026-08-18CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610827455.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-18
Estimated Expiration
2046-06-09

AI Technical Summary

Technical Problem

[0005]本申请目的在于提供一种基于异步对比学习的电缆缺陷诊断方法、装置、设备及介质,旨在解决如何在不依赖严格同步采样和强空间配准的条件下,对存在降质或缺失的电缆带电检测多模态异步数据进行稳定表征与鲁棒融合,从而实现高精度的电缆缺陷诊断的技术问题

Benefits of technology

首先,利用时间容差窗口将不同时刻采集的声学数据、红外图像和可见光图像进行组合归并得到异步多模态样本组,解决了多源传感器感知数据难以严格同步对齐的问题;随后,对各模态进行特征提取并将其投影到共享隐空间中得到跨模态标准化表征,消除了不同模态间的物理底层差异,实现了异构特征在同一维度下的标准化对齐;接着,基于该表征进行假负样本抑制以构建包含正负样本的训练样本集,并据此计算双向异步跨模态对比损失,避免了对比学习中因误判相似样本为负样本而导致的特征表征混乱,提升了特征一致性学习的准确性;同时,结合各模态的质量指标和基于数据获取状态生成的有效模态标志位来计算质量感知注意力权重,使得融合过程能够动态关注高质量数据并有效屏蔽缺失模态带来的结构性干扰;进一步地,对跨模态标准化表征进行降质模拟处理和维度统一映射,并利用前述权重对映射特征进行加权融合得到多模态融合特征,增强了模型在面对现场数据残缺或信号衰减时的抗干扰能力;最后,根据双向异步跨模态对比损失和多模态融合特征对缺陷诊断模型进行联合优化训练并输出最终结果,兼顾了跨模态特征的全局对齐与具体的异常分类任务,综合上述流程,本申请能够在不依赖严格同步采样和强空间配准的条件下,对存在降质或缺失的电缆带电检测多模态异步数据进行稳定表征与鲁棒融合,从而实现高精度的电缆缺陷诊断。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365406B_ABST
    Figure CN122365406B_ABST
Patent Text Reader

Abstract

The application discloses a cable defect diagnosis method and device based on asynchronous contrast learning, equipment and medium, relates to the technical field of cable fault diagnosis, and the method comprises the steps of: merging cable acoustic, infrared and visible light data according to a time tolerance window to obtain an asynchronous multi-modal sample group; extracting sample features, projecting into a hidden space to obtain cross-modal standardized representations; based on this, inhibiting false negative samples to construct a training set to calculate a bidirectional asynchronous contrast loss; performing degradation simulation and dimension mapping on the standardized representations to obtain mapping features, and combining quality indicators and effective flag bits to calculate attention weights; obtaining multi-modal fusion features by weighting and fusing the mapping features according to the weights; and training a diagnosis model using the contrast loss and the fusion features to output defect results. The application can stably represent and robustly fuse degraded or missing multi-modal asynchronous data without relying on strict synchronous sampling and strong spatial registration conditions, and realizes high-precision cable defect diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cable fault diagnosis technology, and in particular to cable defect diagnosis methods, devices, equipment and media based on asynchronous comparative learning. Background Technology

[0002] During long-term operation, power cables are prone to defects such as main insulation deterioration, partial discharge, and joint overheating. In order to ensure the safe operation of the power system and achieve early defect identification, it is necessary to use a combination of multi-source heterogeneous sensing methods, such as acoustic detection, infrared thermal imaging detection, and visible light visual detection, to jointly characterize and monitor the cable condition on the engineering site.

[0003] Current engineering practices primarily employ various sensors to perform live-line testing on cable equipment. In terms of technical implementation, existing methods mostly focus on single-modal analysis of acoustic, thermal, or visual information, or fuse features acquired from different sensors through simple vector concatenation, using this as the basis for assessing the cable's operating status.

[0004] Due to differences in sampling frequencies, triggering mechanisms, and installation locations among various sensors, existing methods struggle to achieve strict synchronization and accurate spatial registration of multimodal data. Simple feature stitching fails to uncover deep correlations between multi-source data, resulting in insufficient robustness of the model when faced with noise interference, image blurring, or sensor failure. Furthermore, conventional contrastive learning in asynchronous scenarios is prone to false negatives, weakening the feature representation effect. Therefore, how to stably represent and robustly fuse multimodal asynchronous data from cable live-line detection, which may be degraded or missing, without relying on strict synchronous sampling and strong spatial registration, to achieve high-precision cable defect diagnosis, has become an urgent problem to be solved. Summary of the Invention

[0005] The purpose of this application is to provide a cable defect diagnosis method, device, equipment and medium based on asynchronous contrastive learning. It aims to solve the technical problem of how to stably characterize and robustly fuse multimodal asynchronous data of cable live-line detection with degradation or missing data without relying on strict synchronous sampling and strong spatial registration, so as to achieve high-precision cable defect diagnosis.

[0006] To achieve the above objectives, this application proposes a cable defect diagnosis method based on asynchronous contrastive learning, the method comprising: Based on the time tolerance window, the acoustic data, infrared images, and visible light images of the cable equipment to be tested are merged to obtain an asynchronous multimodal sample group; Modal feature extraction and shared latent space projection are performed on the acoustic data, infrared images, and visible light images in the asynchronous multimodal sample group to obtain cross-modal normalized representations; Based on the cross-modal normalized representation, false negative sample suppression is performed to construct a training sample set containing positive sample pairs and candidate negative sample sets, and bidirectional asynchronous cross-modal contrastive loss is calculated based on the training sample set. The cross-modal standardized representation is subjected to degraded simulation processing and dimensionality unification mapping to obtain the mapped features. The quality-perceived attention weights are calculated based on the quality indicators and effective modality flags of each modality in the asynchronous multimodal sample group. The effective modality flags are generated based on the acquisition status of each modality within the time tolerance window. The mapped features are weighted and fused according to the quality-perceived attention weights to obtain multimodal fusion features; The defect diagnosis model is jointly optimized and trained based on the bidirectional asynchronous cross-modal contrast loss and the multimodal fusion features, and the cable defect diagnosis result is output through the trained defect diagnosis model.

[0007] In one embodiment, the step of performing spurious negative sample suppression based on the cross-modal normalized representation to construct a training sample set containing positive sample pairs and a candidate negative sample set, and calculating a bidirectional asynchronous cross-modal contrastive loss based on the training sample set includes: The cross-modal normalized representations belonging to the same asynchronous multimodal sample group are combined to obtain positive sample pairs, and any one of the cross-modal normalized representations in the positive sample pair is determined as an anchor sample. Based on the anchor sample, obtain the device identifier and reference timestamp of the remaining cross-modal standardized representations, and remove samples whose device identifier is the same as the device identifier of the anchor sample and whose time interval between the reference timestamp and the reference timestamp of the anchor sample is less than a preset time threshold, to obtain an initial negative sample set. Calculate the working condition vector distance and cluster center distance between the anchor sample in the positive sample pair and each sample in the initial negative sample set; Samples whose working condition vector distance is greater than a preset working condition distance threshold and whose cluster center distance is greater than a preset clustering threshold are retained in the candidate negative sample set; The bidirectional asynchronous cross-modal contrast loss is obtained by performing an inner product operation on the positive sample pairs, the candidate negative sample set, and the preset temperature hyperparameter.

[0008] In one embodiment, the step of merging the acoustic data, infrared image, and visible light image of the cable device to be tested according to a time tolerance window to obtain an asynchronous multimodal sample group includes: Obtain the reference time of the cable equipment to be tested, and determine the preset time tolerance window corresponding to the reference time; Retrieve acoustic data, infrared images, and visible light images acquired within the preset time tolerance window, and read the timestamps of the acoustic data, infrared images, and visible light images; Calculate the time difference between the timestamp and the reference time, and calculate the asynchronous pairing confidence based on the time difference; Determine whether there is message data of the corresponding modality within the preset time tolerance window, and generate the valid modality flag bit according to the existence status and content integrity of the message data; The acoustic data, the infrared image, and the visible light image, which are within the preset time tolerance window and contain the message data, are correlated to obtain an asynchronous multimodal sample group.

[0009] In one embodiment, the step of performing modal feature extraction and shared latent space projection on the acoustic data, infrared image, and visible light image in the asynchronous multimodal sample group to obtain a cross-modal normalized representation includes: The acoustic data in the asynchronous multimodal sample group is converted into a two-dimensional time-frequency tensor by short-time Fourier transform, and the regions of interest are cropped for the infrared image and the visible light image. The two-dimensional time-frequency tensor, the cropped infrared image, and the cropped visible light image are input into the corresponding depth encoder to calculate high-level features and obtain modal high-level features. In the case where the effective modality flag indicates that a modality is missing, the high-level features of the modality are filled using a preset placeholder vector; The filled high-level features of the modality are mapped into a latent space of uniform dimension through a projection head network; The mapped features are normalized to obtain a cross-modal standardized representation.

[0010] In one embodiment, the step of calculating the quality-perceived attention weights based on the quality indices and effective modality flags of each modality in the asynchronous multimodal sample group includes: The Laplacian variance of the visible light image is calculated as a visual quality index, and the signal-to-noise ratio of the acoustic data is calculated as an acoustic quality index. The average temperature difference between the target area and the background area in the infrared image is calculated as a thermal quality indicator. The visual quality index, the acoustic quality index, and the thermal quality index are concatenated to obtain a multimodal quality vector; The multimodal quality vector is input into the attention weight calculation network to obtain the initial fusion weights corresponding to each modality; The initial fusion weights are multiplied by the effective modality flag of the corresponding modality to obtain the effective initial weights; The effective initial weights are normalized to obtain the quality-aware attention weights.

[0011] In one embodiment, the step of performing degradation simulation processing and dimensionality unification mapping on the cross-modal normalized representation to obtain mapped features, and then weighting and fusing the mapped features according to the quality-perceived attention weights to obtain multimodal fused features includes: Random noise injection or masking is performed on the cross-modal normalized representation with a preset discard probability to obtain degraded features; The degraded features are transformed into uniform-dimensional mapped features through a linear mapping layer; Multiply the quality-perceived attention weights corresponding to each modality with the mapping features to obtain the weighted mapping features corresponding to each modality; The weighted mapping features corresponding to each modality are added together to obtain the multimodal fusion features.

[0012] In one embodiment, the step of jointly optimizing and training the defect diagnosis model based on the bidirectional asynchronous cross-modal contrast loss and the multimodal fusion features, and outputting the cable defect diagnosis result through the trained defect diagnosis model includes: The multimodal fusion features are input into the defect classifier to obtain the predicted probability that the cable equipment to be inspected belongs to each preset defect category; Calculate the cross-entropy classification loss based on the predicted probabilities and the actual defect labels; A joint loss function is constructed based on the bidirectional asynchronous cross-modal contrastive loss and the cross-entropy classification loss; Based on the joint loss function, the network parameters of the defect diagnosis model are updated through the backpropagation algorithm to obtain the trained defect diagnosis model; The target asynchronous data is input into the trained defect diagnosis model to obtain the cable defect diagnosis results.

[0013] Furthermore, to achieve the above objectives, this application also proposes a cable defect diagnosis device based on asynchronous contrastive learning, the device comprising: The sample merging module is used to merge the acoustic data, infrared images, and visible light images of the cable equipment to be tested according to the time tolerance window to obtain asynchronous multimodal sample groups. The feature projection module is used to extract modal features and project shared latent space onto the acoustic data, infrared image, and visible light image in the asynchronous multimodal sample group to obtain cross-modal normalized representations. The contrastive learning module is used to perform spurious negative sample suppression based on the cross-modal normalized representation, so as to construct a training sample set containing positive sample pairs and candidate negative sample sets, and calculate bidirectional asynchronous cross-modal contrastive loss based on the training sample set. The weight calculation module is used to perform degradation simulation processing and dimension unification mapping on the cross-modal standardized representation to obtain the mapping features, and calculate the quality-perceived attention weights based on the quality indicators and effective modality flags of each modality in the asynchronous multimodal sample group. The feature fusion module is used to perform weighted fusion of the mapped features according to the quality-aware attention weights to obtain multimodal fusion features; The model diagnosis module is used to jointly optimize and train the defect diagnosis model based on the bidirectional asynchronous cross-modal contrast loss and the multimodal fusion features, and output the cable defect diagnosis result through the trained defect diagnosis model.

[0014] Furthermore, to achieve the above objectives, this application also proposes a cable defect diagnosis device based on asynchronous contrastive learning, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the cable defect diagnosis method based on asynchronous contrastive learning as described above.

[0015] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the cable defect diagnosis method based on asynchronous contrastive learning as described above.

[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the cable defect diagnosis method based on asynchronous contrastive learning as described above.

[0017] One or more technical solutions proposed in this application have at least the following technical effects: First, an asynchronous multimodal sample set is obtained by combining acoustic data, infrared images, and visible light images acquired at different times using a time tolerance window, thus solving the problem of difficult strict synchronization and alignment of multi-source sensor data. Then, features of each modality are extracted and projected into a shared latent space to obtain a cross-modal standardized representation, eliminating the underlying physical differences between different modalities and achieving standardized alignment of heterogeneous features in the same dimension. Next, false negative sample suppression is performed based on this representation to construct a training sample set containing both positive and negative samples, and a bidirectional asynchronous cross-modal contrastive loss is calculated accordingly. This avoids feature representation confusion caused by misclassifying similar samples as negative samples in contrastive learning, improving the accuracy of feature consistency learning. Simultaneously, the quality perception is calculated by combining the quality indicators of each modality and the effective modal flag generated based on the data acquisition state. By understanding the attention weights, the fusion process can dynamically focus on high-quality data and effectively shield structural interference caused by missing modalities. Furthermore, the cross-modal standardized representations are subjected to degradation simulation processing and dimensionality unification mapping, and the mapped features are weighted and fused using the aforementioned weights to obtain multimodal fusion features, which enhances the model's anti-interference ability when faced with incomplete field data or signal attenuation. Finally, the defect diagnosis model is jointly optimized and trained based on bidirectional asynchronous cross-modal contrast loss and multimodal fusion features, and the final result is output, taking into account both the global alignment of cross-modal features and specific anomaly classification tasks. Combining the above processes, this application can stably represent and robustly fuse multimodal asynchronous data of cable live-line detection with degradation or missing data without relying on strict synchronous sampling and strong spatial registration, thereby achieving high-precision cable defect diagnosis. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating an embodiment of the cable defect diagnosis method based on asynchronous contrastive learning in this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the cable defect diagnosis method based on asynchronous contrastive learning in this application; Figure 3 This is a schematic diagram of the module structure of the cable defect diagnosis device based on asynchronous contrastive learning according to an embodiment of this application; Figure 4 This is a schematic diagram of the hardware operating environment involved in the cable defect diagnosis method based on asynchronous contrastive learning in the embodiments of this application.

[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0024] It should be noted that the executing entity of this application embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or cable defect diagnosis system capable of realizing the above functions. The following uses a cable defect diagnosis system as an example to describe this embodiment and the following embodiments.

[0025] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0026] Based on this, embodiments of this application provide a cable defect diagnosis method based on asynchronous contrastive learning, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the cable defect diagnosis method based on asynchronous contrastive learning in this application.

[0027] In this embodiment, the cable defect diagnosis method based on asynchronous contrastive learning includes steps S10 to S60: Step S10: Merge the acoustic data, infrared image and visible light image of the cable equipment to be tested according to the time tolerance window to obtain an asynchronous multimodal sample group; It should be noted that the time tolerance window can be a time interval set to allow for deviations due to differences in the acquisition time of different modal data. An asynchronous multimodal sample group refers to a collection of multi-source data that are merged into the same time tolerance window and reflect the operating status of the same cable equipment within a similar time period.

[0028] Understandably, the process involves obtaining a reference timestamp of the cable equipment under test, and then defining a time interval before and after this timestamp as a time tolerance window. The length of this window can be set to 30 seconds, based on the fact that the time difference between acoustic, infrared, and visible light acquisition devices passing the same monitoring point during routine equipment inspections is typically within minutes. Acoustic data, infrared images, and visible light images falling within the time tolerance window are retrieved. Acoustic data, infrared images, and visible light images belonging to the same equipment identifier and falling within the time tolerance window are combined to form an asynchronous multimodal sample group. This step solves the problem of strict synchronization of multi-source data, improves the utilization rate of multi-source heterogeneous data in asynchronous acquisition environments, and provides a data foundation for subsequent multimodal information fusion.

[0029] Step S20: Modal feature extraction and shared latent space projection are performed on the acoustic data, infrared image and visible light image in the asynchronous multimodal sample group to obtain cross-modal normalized representation; It should be noted that cross-modal normalized representation refers to multimodal feature vectors that, after dimensionality unification and scaling, can be mapped onto the same feature space for direct measurement and comparison. Shared latent space refers to a high-dimensional feature space constructed through network mapping, where the high-dimensional features of different modalities possess comparable mathematical properties.

[0030] Understandably, the acoustic data is converted into a time-frequency image using a short-time Fourier transform. The time-frequency image, infrared image, and visible light image are then input into their respective Vision Transformer networks for feature extraction, outputting high-level modal feature vectors. A multilayer perceptron network maps these high-level feature vectors to a shared latent space of uniform dimension, and L2 regularization is used to scale the mapped feature vectors, outputting a cross-modal standardized representation. This step eliminates the underlying differences in physical properties and data structure between different modalities, enabling acoustic, thermal, and visual data to be aligned within a unified dimension, facilitating subsequent consistency learning.

[0031] Step S30: Perform spurious negative sample suppression based on the cross-modal normalized representation to construct a training sample set containing positive sample pairs and candidate negative sample sets, and calculate bidirectional asynchronous cross-modal contrastive loss based on the training sample set; It should be noted that spurious negative sample suppression refers to the process of removing samples that are similar to the anchor sample in actual physical state or belong to the same working environment during the negative sample sampling process of contrastive learning. Bidirectional asynchronous cross-modal contrastive loss can be a loss function value that measures the similarity error between the mapping of features from two different modalities.

[0032] Understandably, the process involves pairing cross-modal standardized representations within the same asynchronous multimodal sample group as positive sample pairs; extracting features from other sample groups as initial negative samples; and extracting operating condition data (including operating voltage and ambient temperature) at the corresponding acquisition times of the anchor sample and the initial negative sample. The Euclidean distance between the anchor sample and the initial negative sample's operating condition data is calculated. If the Euclidean distance is greater than a preset operating condition distance threshold, the initial negative sample is included in the candidate negative sample set. The preset operating condition distance threshold can range from 1.0 to 1.5, based on the statistical mean of historical normal operating condition fluctuations. The feature inner product between the positive sample pairs and the candidate negative sample set is calculated using cosine similarity, and the bidirectional asynchronous cross-modal contrastive loss is calculated using the cross-entropy loss function. This step avoids incorrectly rejecting samples with similar physical states or operating conditions as negative samples, thus improving the stability of cross-modal feature learning and the discriminative ability of the representation results.

[0033] Step S40: Perform degradation simulation processing and dimension unification mapping on the cross-modal standardized representation to obtain the mapping features, and calculate the quality-perceived attention weights based on the quality indicators and effective modality flags of each modality in the asynchronous multimodal sample group. The effective modality flags are generated based on the acquisition status of each modality within the time tolerance window. It should be noted that degradation simulation processing refers to randomly applying masks or noise to the input features during the training phase to simulate signal attenuation or partial physical obstruction phenomena in industrial field data acquisition equipment. Mapped features refer to a set of multimodal features that have undergone dimensionality unification transformation through a linear mapping layer, eliminating differences in the number of feature channels and possessing the same feature space dimension. Quality metrics are numerical values ​​reflecting the clarity, signal-to-noise ratio, or feature saliency of each modality's data. Effective modality flags can be binary values ​​used to identify whether a certain modality's data was successfully acquired within the current time tolerance window. Quality-aware attention weights are fusion ratio coefficients assigned to different modalities based on the quality and presence of modal data.

[0034] Understandably, random numbers following a uniform distribution are generated. If the random number is less than the degradation probability, the values ​​of some dimensions in the cross-modal normalization representation are set to zero or Gaussian noise is added. The degradation probability is preset between 0.15 and 0.25, based on historical packet loss rate statistics of industrial field communication equipment. A fully connected layer is used to transform the features after degradation simulation to the same channel dimension, resulting in mapped features.

[0035] The Laplacian variance of the visible light image is calculated as its corresponding quality metric; the signal-to-noise ratio of the acoustic data is calculated as its corresponding quality metric; and the mean temperature difference between the target and background regions in the infrared image is calculated as its corresponding quality metric. It is determined whether the corresponding modality data packet was successfully acquired within the time tolerance window. If successfully acquired and the packet is not empty, the valid modality flag is set to 1; otherwise, the valid modality flag is set to 0. The acquired quality metrics are concatenated into a feature vector and input into the multilayer perceptron to obtain the initial weights. The initial weights are multiplied by the corresponding valid modality flags to obtain the quality perception attention weights. This step enhances the model's robustness in extreme situations such as incomplete, obstructed, or communication-abnormal field equipment acquisition, and enables the subsequent fusion process to dynamically focus on high-quality, valid data modalities, reducing the interference of low-quality data or missing modalities on the overall diagnostic results.

[0036] Step S50: The mapped features are weighted and fused according to the quality-perceived attention weights to obtain multimodal fusion features; It should be noted that multimodal fusion features refer to the comprehensive feature expression obtained by weighted integration of feature vectors from multiple sensor modalities.

[0037] Understandably, the process involves multiplying each element of the mapped feature by its corresponding quality-aware attention weight, and then summing the modal-weighted mapped features to obtain the multimodal fusion feature. This step ensures that the fused feature effectively retains key diagnostic information.

[0038] Step S60: Perform joint optimization training on the defect diagnosis model based on the bidirectional asynchronous cross-modal contrast loss and the multimodal fusion features, and output the cable defect diagnosis result through the trained defect diagnosis model.

[0039] It should be noted that the defect diagnosis model can be a deep learning network that receives multimodal fusion features and outputs the classification probability of whether the cable has abnormal states such as partial discharge and insulation degradation.

[0040] Understandably, the multimodal fusion features are input into the fully connected classification layer within the defect diagnosis model, outputting the predicted probabilities of each preset defect category. The classification cross-entropy loss is calculated by combining the predicted probabilities with the actual defect labels. This loss is then added to the bidirectional asynchronous cross-modal comparison loss using predetermined weighting coefficients to form the total network loss. The Adam optimization algorithm is used for backpropagation based on the total network loss to update the network parameters of the defect diagnosis model. After training convergence, the real-time extracted multimodal fusion features are input into the trained defect diagnosis model, outputting the defect category corresponding to the highest predicted probability as the cable defect diagnosis result. This step balances cross-modal consistency learning with the specific anomaly classification task, enabling the network to directly serve on-site anomaly monitoring while aligning multi-source features, thus improving the accuracy of defect category identification.

[0041] This embodiment effectively solves the problem of difficult-to-synchronize data utilization from multi-source sensors in industrial settings due to different sampling frequencies and spatial misalignment by introducing a time tolerance mechanism and a contrastive learning strategy. Simultaneously, by employing degradation simulation and a quality-aware attention mechanism, the model can still dynamically allocate weights and extract key correlation features even in harsh environments with missing modalities or external noise pollution. This embodiment improves the utilization rate of multimodal sensing data under asynchronous conditions, enhances the robustness of the diagnostic model against interference, and thus achieves stable and high-precision diagnosis of minor cable defects, effectively reducing maintenance costs caused by false alarms or missed alarms.

[0042] As an example, the step of merging the acoustic data, infrared image, and visible light image of the cable device under test according to the time tolerance window to obtain an asynchronous multimodal sample group includes: obtaining a reference time of the cable device under test and determining a preset time tolerance window corresponding to the reference time; retrieving the acoustic data, infrared image, and visible light image collected within the preset time tolerance window, and reading the timestamps of the acoustic data, infrared image, and visible light image; calculating the time difference between the timestamp and the reference time, and calculating the asynchronous pairing confidence based on the time difference; determining whether there is message data of the corresponding mode within the preset time tolerance window, and generating the valid mode flag bit according to the existence status and content integrity of the message data; associating the acoustic data, infrared image, and visible light image that contain the message data within the preset time tolerance window to obtain an asynchronous multimodal sample group.

[0043] It should be noted that the reference time refers to the time axis reference point selected during data merging and processing. The timestamp can be the specific system time at which the underlying sensing device captures the physical signal. Asynchronous pairing confidence refers to the weighted numerical value used to evaluate the reliability of matching between asynchronously acquired multi-source data within the same window in the time dimension. Taking acoustic and thermal modes as an example, the formula for calculating asynchronous pairing confidence is as follows: in, and They represent the sample groups respectively. The actual acquisition time of acoustic and thermal data. This represents the time decay coefficient. Using the same method, the asynchronous pairing confidence scores for acoustic and visual inputs can also be obtained separately. Asynchronous pairing confidence of thermal and visual modalities .

[0044] Message data refers to the data carrier containing sensing content and communication packet headers sent by the sensor node to the receiving end. Existence status can be the verification result of whether a certain type of modality data is retained in the data buffer queue within a specific time period. Content integrity refers to the healthy state of the received data carrier, meaning it is not empty and the data format is not corrupted. In this example, the time tolerance window can be a time span centered on a reference time, set to account for differences in sampling frequencies of different sensors. In this example, the valid modality flag can be a binary label used to characterize whether a certain modality has usable sensing data within the current merging period. In this example, the asynchronous multimodal sample group refers to a multi-source data set merged within the same time tolerance window and reflecting similar operating states of the same cable equipment.

[0045] Understandably, the process begins by receiving cable equipment status information from the field monitoring terminal to extract the reference time for the current analysis task. This reference time is then extended forward and backward by 15 seconds, forming a preset time tolerance window with a total span of 30 seconds. The window length is based on the fact that the physical delay time for polling and collecting acoustic, thermal, and optical signals by the field inspection robot or fixed sensors is typically less than half a minute. Next, the system searches the local database or real-time cache queue for acoustic data, infrared images, and visible light images that fall within the preset time tolerance window span. Simultaneously, the corresponding communication packet headers are parsed to extract the timestamps of when the acoustic data, infrared images, and visible light images were actually captured by the sensors.

[0046] Then, the timestamps of each extracted modality are subtracted from the reference time to obtain the time difference. The calculated time difference is then substituted into a confidence function that exhibits negative exponential decay to calculate the asynchronous pairing confidence. The larger the time difference in this function, the lower the output confidence, reflecting that data with longer time intervals have lower reliability in representing the same physical transient.

[0047] Furthermore, the receiving queues of the acoustic, thermal, and visual modalities are traversed to check whether the corresponding message data has been actually received within the preset time tolerance window. When the message data of the corresponding modality is found, the file size of the message data is checked to see if it is greater than zero and the file structure is complete. If the existence status and the content integrity check pass, it is determined that there is usable sensing data for the modality and the corresponding valid modality flag is set to 1. Otherwise, if no message data is received or the message data is corrupted, the corresponding valid modality flag is set to 0.

[0048] Finally, the acoustic data, infrared image, and visible light image corresponding to the effective modal flag bit being assigned a value of 1 are bound together with the device identifier to establish a mapping relationship between the corresponding modal data. This combination constructs an asynchronous multimodal sample group that reflects the current operating status of the cable equipment, providing a preliminary screening and time-aligned foundation for subsequent cross-modal feature fusion.

[0049] This example overcomes the sampling asynchrony limitations caused by different hardware triggering mechanisms of multiple sensors in the field by combining a time tolerance window with timestamp verification, thus improving the efficiency of organizing heterogeneous data in the time dimension. By combining message data integrity verification and existence status judgment, a flag representing data availability is dynamically generated, filtering out damaged or missing invalid signals at the source and reducing the interference of underlying communication faults on upper-layer diagnostic models. Overall, this merging method builds a reliable data association foundation, providing high-confidence multimodal samples for subsequent cable health status assessment and improving the adaptability of defect diagnosis in complex industrial environments.

[0050] As an example, the step of extracting modal features and projecting them into a shared latent space from the acoustic data, infrared image, and visible light image in the asynchronous multimodal sample group to obtain a cross-modal standardized representation includes: converting the acoustic data in the asynchronous multimodal sample group into a two-dimensional time-frequency tensor using a short-time Fourier transform, and cropping the regions of interest in the infrared image and the visible light image; inputting the two-dimensional time-frequency tensor, the cropped infrared image, and the cropped visible light image into the corresponding depth encoder for high-level feature calculation to obtain modal high-level features; filling the modal high-level features with a preset placeholder vector when the effective modal flag indicates a missing modality; mapping the filled modal high-level features to a latent space of a unified dimension using a projection head network; and normalizing the mapped features to obtain a cross-modal standardized representation.

[0051] It should be noted that a two-dimensional time-frequency tensor refers to matrix data that contains both time and frequency dimensions after a one-dimensional time series signal is transformed through frequency domain conversion. Region of interest (ROI) cropping can be an image processing operation that uses object detection algorithms or edge contour information to extract pixel regions containing key parts of an image, such as the cable body and intermediate joints. A depth encoder is a neural network model structure used to extract high-dimensional features from input data. Preset placeholder vectors can be multi-dimensional arrays with all elements equal to zero, used to maintain the consistency of the neural network input structure and prevent forward propagation interruption when data is missing. A projection head network is a network component composed of multilayer perceptrons used to linearly map modal features of different dimensions to the same latent space dimension.

[0052] Understandably, the process involves first acquiring acoustic data from asynchronous multimodal sample groups, extracting their time-frequency features using short-time Fourier transform, generating two-dimensional time-frequency tensors with corresponding channel number, height, and width, and simultaneously using the YOLO target detection algorithm to identify cable joints and body regions in infrared and visible light images. These are then used as boundary coordinates for region-of-interest cropping, eliminating irrelevant background pixels.

[0053] Next, a deep encoder group including an acoustic encoder, a thermal encoder, and a visual encoder is constructed. These encoders all adopt a vision transformer architecture based on block embedding and self-attention mechanisms. The two-dimensional time-frequency tensor, the cropped infrared image, and the cropped visible light image are input into the corresponding modal deep encoders, respectively. After forward computation by a multi-head self-attention layer, the high-level features of each modality are extracted. Specifically, the two-dimensional time-frequency tensor is treated as a single-channel or pseudo-color image, divided into multiple time-frequency patches according to a preset size, and linearly projected to adapt to the sequence input format of the vision transformer.

[0054] Read the valid modality flags corresponding to each modality. When the flag values ​​indicate that data for a specific modality is missing, generate a preset placeholder vector that is consistent with the high-level feature dimension of that modality. All elements in this vector are zero. Use this zero vector directly as the feature representation of the missing modality to avoid errors in subsequent feature splicing or fusion networks due to local data incompleteness.

[0055] Furthermore, the high-level features or preset placeholder vectors of each modality are uniformly input into the projection head network. After linear transformation by the multilayer perceptron, they are mapped to a latent space with a uniform number of feature channels. Finally, the L2 norm is used to normalize the feature vectors in the latent space, scale and constrain the numerical distribution range of the feature vectors, and output a cross-modal standardized representation.

[0056] This example uses short-time Fourier transform and region-of-interest (ROI) clipping to filter out environmental background noise and highlight key features of the cable, reducing the computational load of subsequent feature extraction. A visual transformer architecture is introduced to achieve deep global semantic mining of data from different modalities, and a placeholder zero-vector padding mechanism addresses the common data gap problem in asynchronous industrial scenarios, ensuring the structural coherence of the model. Combining dimensionality reduction mapping and normalization with a projection head network, heterogeneous multi-source sensor data is transformed into the same scaling space, eliminating the barriers between the underlying physical dimensions of sound, light, and heat, providing a measurable and structurally unified representation foundation for subsequent cross-modal consistent comparative learning.

[0057] As an example, the step of calculating the quality-perceived attention weights based on the quality indices and effective modality flags of each modality in the asynchronous multimodal sample group includes: calculating the Laplacian variance of the visible light image as a visual quality index, and calculating the signal-to-noise ratio of the acoustic data as an acoustic quality index; calculating the average temperature difference between the target region and the background region in the infrared image as a thermal quality index; concatenating the visual quality index, the acoustic quality index, and the thermal quality index to obtain a multimodal quality vector; inputting the multimodal quality vector into an attention weight calculation network to obtain the initial fusion weights corresponding to each modality; multiplying the initial fusion weights with the effective modality flags of the corresponding modality to obtain effective initial weights; and normalizing the effective initial weights to obtain the quality-perceived attention weights.

[0058] It should be noted that visual quality metrics can be scalar values ​​reflecting the sharpness of image texture and edges. Laplacian variance refers to the variance of the gradients of all pixels after applying the second derivative of the Laplacian operator to the image pixels. Visual quality metrics (sharpness evaluation metrics) are expressed as follows: in, This indicates the total number of pixels in the visible light image. Represents the Laplace operator. This represents the mean of the Laplace response. This represents the traversal index of each pixel in the image. Indicates the first The grayscale matrix representation of the visible light images of each sample group.

[0059] Acoustic quality metrics are numerical parameters used to measure the relative strength of effective monitoring information in an acoustic signal compared to background noise, such as signal-to-noise ratio (SNR). Thermal quality metrics can be physical quantities that characterize the degree of thermal contrast between a heat source in an infrared image and its surrounding normal environment, such as average temperature difference. A multimodal quality vector is a one-dimensional array composed of scalar quality evaluation values ​​from different sensors arranged in a fixed order. In this example, the attention weight calculation network is a lightweight neural network consisting of a multilayer perceptron and a normalized exponential function (Softmax) layer, used to output normalized probability distribution characteristics. Initial fusion weights can be the attention weight parameters initially assigned to each modality by the network. Effective initial weights refer to the weight values ​​retained only for the modalities with actual sensing data after existence state verification.

[0060] Understandably, the process begins with extracting visible light images from the asynchronous multimodal sample group and converting them to grayscale. The Laplacian operator is then used to calculate the second-order derivatives of the pixels in the entire image, and the variance of these second-order derivatives is taken as a visual quality indicator. A larger variance indicates a clearer visible light image captured on-site. Next, the original signal frequency bands of the acoustic data are extracted, and the ratio of the effective high-frequency discharge pulse or vibration signal power to the background noise power is calculated to obtain the signal-to-noise ratio, which is used as an acoustic quality indicator. Finally, the pixel temperature values ​​of the heat-generating target area and the surrounding background area in the infrared image are extracted separately, and the difference between their mean values ​​is calculated to obtain the average temperature difference, which is used as a thermal quality indicator.

[0061] Then, the calculated visual quality index, acoustic quality index, and thermal quality index are concatenated in a fixed order of sound, heat, and light to construct a multimodal quality vector containing three-dimensional features. The multimodal quality vector is input into the attention weight calculation network, which uses an internal multilayer perceptron to perform a nonlinear mapping of the quality correlation between the various modalities. Finally, the output value is transformed into a probability distribution form through the softmax layer at the end to obtain the initial fusion weights corresponding to each modality.

[0062] Furthermore, the effective modality flags generated based on the message acquisition status within the time tolerance window are read. The initial fusion weights corresponding to acoustic, thermal, and visual modes are multiplied element-wise by their respective effective modality flags (values ​​of 0 or 1) to obtain the effective initial weights. At this point, the weights of missing modalities are directly masked because the multiplier is zero. The effective initial weights are then normalized by dividing the effective initial weight of each modality by the sum of the effective initial weights of all modalities and adding the result to the zero-prevention constant. The zero-prevention constant can be 1e-5 to prevent the denominator from being zero. The output is the quality-aware attention weight used to guide subsequent multi-source feature fusion.

[0063] This example objectively and quantitatively evaluates the usability of various types of sensory data by extracting Laplacian variance, signal-to-noise ratio, and mean temperature difference, reducing the interference of blurry images or high-noise signals on the overall health status assessment. By combining a lightweight attention network with a flag masking mechanism, the fusion process can dynamically and adaptively allocate the attention weight of each feature based on real-time data quality. This effectively masks structural interference information introduced by missing modalities, improves the model's robustness under complex conditions such as packet loss in on-site communication or physical degradation of sensors, and ensures the rationality of information selection during the multi-source heterogeneous feature fusion stage.

[0064] As an example, the step of performing degradation simulation processing and dimensional unification mapping on the cross-modal standardized representation to obtain mapping features, and then weighting and fusing the mapping features according to the quality-aware attention weights to obtain multimodal fusion features includes: injecting random noise or masking the cross-modal standardized representation with a preset dropout probability to obtain degraded features; converting the degraded features into mapping features with a unified dimension through a linear mapping layer; multiplying the quality-aware attention weights corresponding to each modality with the mapping features to obtain weighted mapping features corresponding to each modality; and summing the weighted mapping features corresponding to each modality to obtain multimodal fusion features.

[0065] It should be noted that the preset discard probability refers to a threshold value set manually during model training to artificially interfere with and disrupt the features of the input data. Random noise injection can be an operation that superimposes a sequence of random numerical values ​​following a Gaussian distribution onto the original feature tensor. Masking refers to the operation of forcibly replacing the values ​​of some dimensions in the feature tensor with zero values. Degraded features refer to feature vectors used to simulate harsh communication and perception environments after interference and disruption processing. A linear mapping layer refers to a basic network structure composed of fully connected neural networks used for dimensionality transformation. Mapped features can be a set of multimodal features with the same channel dimension after dimensionality scaling. Weighted mapped features refer to single-modal feature vectors assigned different attention weight coefficients.

[0066] Understandably, the process begins by generating a random number that follows a uniform distribution from zero to one. This random number is then compared to a preset drop probability, which can be set to 0.2. This value is based on simulating a typical industrial scenario where sensor communication packet loss is around 20% or there is local physical obstruction. When the generated random number is less than the preset drop probability, Gaussian noise with a mean of zero and a variance of one is superimposed on the cross-modal normalized representation, or elements in some dimensions of its feature tensor are forced to zero to perform masking. This simulates signal degradation or loss in the field, resulting in degradation features. Conversely, if the random number is greater than or equal to the preset drop probability, the original numerical distribution of the cross-modal normalized representation remains unchanged, and it is directly used as a degradation feature in subsequent networks.

[0067] The degraded features corresponding to each mode are input into the internally constructed linear mapping layer for matrix multiplication. This eliminates the differences in the number of channels formed in the early stage of acoustic, infrared and visible light feature extraction, and transforms the originally different deep features into the same feature space dimension, resulting in mapping features with consistent internal structure, which provides an aligned matrix basis for subsequent numerical addition operations.

[0068] Next, the quality perception attention weights of the corresponding modalities calculated in the previous step are obtained. The quality perception attention weights of each modality are multiplied by the corresponding mapping features using dot multiplication. Based on the quality of the multi-source data, the feature expressions are amplified or suppressed at the numerical level to obtain the weighted mapping features corresponding to each modality. Then, the weighted mapping features corresponding to acoustic, infrared and visible light are matrix-added to integrate the complementary state information of various sensing devices and obtain the final multimodal fusion features used for defect assessment.

[0069] This example effectively simulates common signal degradation and local failure phenomena in real industrial environments during model training by introducing random noise injection and masking mechanisms. This allows the network to adapt to imperfect sensing conditions in advance, improving the diagnostic model's generalization ability and robustness when faced with incomplete sensing data. By combining the weighted calculation of the linear mapping layer and the quality-sensing attention weights, not only are the underlying dimensional barriers between multi-source heterogeneous data eliminated, but the fusion weight of each modal feature is also dynamically adjusted according to real-time signal quality. This effectively shields the negative interference of degraded modes on the overall evaluation, improves the rationality of multi-modal data information integration, and thus ensures high accuracy in cable defect identification under complex working conditions.

[0070] As an example, the steps for constructing the defect diagnosis model include: constructing a deep encoder based on a visual transformer structure based on block embedding and self-attention mechanism; constructing a projection head network based on a multilayer perceptron; constructing a linear mapping layer based on a fully connected linear layer; constructing an attention weight calculation network based on a multilayer perceptron and a Softmax activation function; constructing a defect classifier based on a fully connected classification layer; and combining the deep encoder, the projection head network, the linear mapping layer, the attention weight calculation network, and the defect classifier to obtain the defect diagnosis model.

[0071] It should be noted that patch embedding refers to the preprocessing operation of dividing input data into fixed-size local blocks and flattening them to map into a sequence of feature vectors. Self-attention mechanisms can be computational modules within a network used to calculate the dependencies between elements at different positions in the sequence and assign corresponding feature weights. The visual transformer architecture refers to a deep neural network architecture that primarily relies on self-attention mechanisms to capture global contextual information of the input data, namely the VisionTransformer (ViT) architecture. A multilayer perceptron is a feedforward artificial neural network consisting of at least one hidden layer with fully connected neurons in adjacent layers. The Softmax activation function is a mathematical function that maps the output values ​​of multiple neurons to a probability distribution feature between zero and one, with a sum of one. A fully connected linear layer can be a network layer where neurons are connected to all nodes in the previous layer and a linear weighted summation operation is performed. A fully connected classification layer is a network structure located at the end of the network used to reduce the dimensionality of high-dimensional feature vectors, map them to a preset number of categories, and output classification scores. A defect classifier is a network component that receives fused multi-source features and determines the type of fault in cable equipment. A defect diagnosis model can be a deep learning system that integrates multiple network components such as feature extraction, dimension transformation, weight allocation, and category determination.

[0072] Understandably, firstly, a deep encoder is constructed using a visual transformer structure. Specifically, a block embedding module is set at the front end of the network to divide the input feature map into a sequence of fixed-size patches. Then, multiple multi-head self-attention modules are stacked afterward to extract global contextual features of each modality without relying on traditional convolutional operations.

[0073] Next, a projection head network is constructed using a multilayer perceptron comprising input, hidden, and output layers, enabling the network to perform nonlinear mapping. This is used to transform high-level features of different dimensions into a shared latent space during the contrastive learning stage. Simultaneously, a linear mapping layer is constructed using fully connected linear layers without nonlinear activation functions, relying on pure matrix multiplication to adjust the feature channel dimensions after the degradation simulation. Then, an attention weight calculation network is constructed by cascading a multilayer perceptron and a softmax activation function. The multilayer perceptron is responsible for learning the hidden layer correlation features within the multimodal quality vector, while the softmax activation function at the end is responsible for converting these feature values ​​into normalized probability weights corresponding to each modality.

[0074] Furthermore, a fully connected classification layer with the same number of nodes as the total number of preset defect categories such as cable insulation degradation and partial discharge is placed at the prediction end of the network to construct a defect classifier that outputs specific fault probabilities. The constructed deep encoder, projection head network, linear mapping layer, attention weight calculation network, and defect classifier are then combined and interfaced with tensors and encapsulated in code according to the data forward propagation flow order to obtain a complete end-to-end defect diagnosis model.

[0075] This example employs a visual transformer architecture to construct a deep encoder, enhancing the model's ability to capture global correlation information from multi-source perceptual data and overcoming the limitations of traditional local receptive fields in processing complex industrial images. A projection head and dimension mapping network are constructed using a multilayer perceptron and fully connected linear layers, respectively, ensuring smoothness and computational efficiency during feature transformations across different spaces. Combining an attention network with a softmax activation function and a fully connected classifier achieves a direct mapping from data quality assessment to the final anomaly category. Overall, this modular approach improves the clarity and scalability of the network structure, enhances the coherence of heterogeneous data processing, and provides a robust underlying model for complex cable live-line detection scenarios.

[0076] As an example, the step of jointly optimizing and training the defect diagnosis model based on the bidirectional asynchronous cross-modal contrastive loss and the multimodal fusion features, and outputting the cable defect diagnosis result through the trained defect diagnosis model includes: inputting the multimodal fusion features into a defect classifier to obtain the predicted probability that the cable device to be detected belongs to each preset defect category; calculating the cross-entropy classification loss based on the predicted probability and the real defect label; constructing a joint loss function based on the bidirectional asynchronous cross-modal contrastive loss and the cross-entropy classification loss; updating the network parameters of the defect diagnosis model through a backpropagation algorithm based on the joint loss function to obtain the trained defect diagnosis model; and inputting the target asynchronous data into the trained defect diagnosis model to obtain the cable defect diagnosis result.

[0077] It should be noted that in this example, the true defect label refers to the actual classification code representing the actual physical health status of the cable equipment, pre-labeled by human experts. The cross-entropy classification loss can be a numerical metric used to measure the difference between the predicted probability distribution of the model output and the distribution of the true defect labels. The joint loss function is a global optimization objective mathematical expression formed by combining the measurement error of the contrastive learning task with the measurement error of the classification task in a predetermined proportion. The target asynchronous data can be the multi-source sensing data to be detected, collected in real time by field monitoring equipment after the model is deployed and before it has undergone synchronization and alignment processing.

[0078] Understandably, the calculated multimodal fusion features are first input into the fully connected network layer inside the defect classifier. Matrix multiplication operations within the fully connected network layer are then used for feature dimensionality reduction, outputting a predicted probability vector that includes various preset defect categories such as normal cable operation, insulation main layer degradation, partial discharge, and overheating of intermediate joints. The predicted probability vector is represented as follows: in, Indicates the first The predicted probability vector of each sample group belonging to each preset cable defect category (such as insulation degradation, partial discharge, etc.). This represents the total number of defect categories. This represents a defect classifier network (usually a fully connected classification layer). Indicates the first The multimodal fusion features were obtained by performing quality reduction simulation, dimensional unification, and quality-perceived attention weighting on the sample groups.

[0079] Next, the true defect labels corresponding to the current multimodal fusion features are extracted from the dataset, and the cross-entropy classification loss between the predicted probability vector and the true defect labels is calculated. The cross-entropy classification loss is expressed as follows: in, For batch size during training, Indicates the first The weight coefficients for each defect category, when no category reweighting is required, let ; Indicates the first The sample at the th Real labels in each category The classifier predicts the first... The sample belongs to the first The probability values ​​of each category (corresponding to the output components of the predicted probability vector).

[0080] The bidirectional asynchronous cross-modal contrastive loss calculated from the preceding network layers is obtained. Preset weight coefficients are assigned to the cross-entropy classification loss and the bidirectional asynchronous cross-modal contrastive loss, respectively. The value range of the weight coefficients is between 0.1 and 1.0. The purpose of setting this value range is to balance the contribution ratio of the feature representation alignment task and the specific defect classification task to the gradient of model parameter update. The weighted bidirectional asynchronous cross-modal contrastive loss and the cross-entropy classification loss are added together to construct a joint loss function that guides the optimization of the global network.

[0081] Furthermore, with minimizing the output value of the joint loss function as the network optimization objective, the Adam optimizer combined with the backpropagation algorithm is used to calculate the gradient information of the parameters of each node in the network. The network parameters of each component within the defect diagnosis model are then updated end-to-end along the gradient descent direction according to a set learning rate. After multiple iterations on the training set until the joint loss function converges and stabilizes, the trained defect diagnosis model is obtained and saved for online deployment.

[0082] In the practical application phase at the engineering site, the system receives the target asynchronous data uploaded by the field sensor nodes, inputs the target asynchronous data into the trained defect diagnosis model to perform forward inference operations, and extracts the classification category with the largest value from the probability vector output by the fully connected layer at the end of the model as the cable defect diagnosis result of the cable equipment to be inspected.

[0083] This example constructs a joint loss function, enabling the model to focus not only on the accuracy of the final defect classification during optimization but also on the consistency alignment of the underlying cross-modal feature space, effectively improving the discriminative ability of multi-source heterogeneous feature representations. The backpropagation algorithm is used to update the parameters of the overall diagnostic network end-to-end, avoiding the error accumulation problem caused by the independent training of each component and enhancing the collaborative effect between sub-network modules. By directly inputting the target asynchronous data into the trained model for inference, efficient online diagnosis under non-strictly synchronous conditions in industrial fields is achieved, improving the reliability and practical value of early warning of cable defects.

[0084] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in Embodiment 1 above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the cable defect diagnosis method based on asynchronous contrastive learning according to this application. Step S30 of the cable defect diagnosis method based on asynchronous contrastive learning includes steps S31 to S35: Step S31: Combine the cross-modal normalized representations belonging to the same asynchronous multimodal sample group to obtain positive sample pairs, and determine any one of the cross-modal normalized representations in the positive sample pairs as anchor samples; Step S32: Based on the anchor sample, obtain the device identifier and reference timestamp of the remaining cross-modal normalized representations, and remove samples whose device identifier is the same as the device identifier of the anchor sample and whose time interval between the reference timestamp and the reference timestamp of the anchor sample is less than a preset time threshold, to obtain an initial negative sample set. Step S33: Calculate the working condition vector distance and cluster center distance between the anchor sample in the positive sample pair and each sample in the initial negative sample set; Step S34: Samples whose working condition vector distance is greater than a preset working condition distance threshold and whose cluster center distance is greater than a preset clustering threshold are retained in the candidate negative sample set. Step S35: Perform inner product operation based on the positive sample pair, the candidate negative sample set, and the preset temperature hyperparameter to obtain the bidirectional asynchronous cross-modal contrast loss.

[0085] It should be noted that in this example, a positive sample pair refers to a matching data pair composed of different modal features extracted from the same asynchronous multimodal sample group. Anchor samples refer to target feature vectors used as benchmark features to measure similarity with other samples during the contrastive learning process. The initial negative sample set can be a group of mismatched feature vectors initially screened after excluding interference data from the same source and similar time periods. The operating condition vector distance refers to the numerical difference between two samples in attributes such as operating voltage or load current, calculated using the Euclidean distance algorithm. The cluster center distance can be the spatial span of the cluster centers of the two feature vectors in the latent space. The candidate negative sample set refers to the effective feature vector set that can be used to construct the reverse thrust of contrastive learning after multi-dimensional difference judgment. In this example, the bidirectional asynchronous cross-modal contrastive loss refers to the loss function value that simultaneously considers the differences in bidirectional mapping feature distributions between different modalities. The preset temperature hyperparameter can be a scalar parameter used to smooth the probability distribution and adjust the model's attention to difficult negative samples.

[0086] Understandably, the process begins by pairwise combining the cross-modal normalized representations extracted from the same asynchronous multimodal sample group to form positively correlated pairs. One of these pairs is then randomly selected as the anchor sample. Using the anchor sample as a benchmark, the device identifier and reference timestamp associated with the remaining cross-modal normalized representations in the dataset are read. To avoid misclassifying highly similar data as negative samples, the device identifiers and reference timestamps of the remaining samples are compared with the corresponding information of the anchor sample. Samples with the same device identifier as the anchor sample but whose reference timestamp difference is less than a preset time threshold are removed. The preset time threshold can be set to 600 seconds, based on the assumption that the physical thermal inertia and partial discharge state of the cable typically do not undergo drastic changes within ten minutes. The remaining samples are then aggregated to obtain the initial negative sample set.

[0087] Next, the operating data corresponding to the anchor point samples and each sample in the initial negative sample set are obtained, and the cluster centers of each sample are extracted using the K-Means clustering algorithm. The operating vector distance between the anchor point sample in the positive sample pair and each sample in the initial negative sample set is calculated, and the distance between their cluster centers is also calculated.

[0088] The calculated distance values ​​are compared against a threshold. Samples whose working condition vector distance is greater than a preset working condition distance threshold and whose cluster center distance is greater than a preset clustering threshold are selected and retained, and added to the candidate negative sample set. The preset working condition distance threshold can be set to 1.0, and the preset clustering threshold can be set to 0.5. These two parameters are set based on the extreme value statistics of historical normal operation fluctuation range, which ensures that the retained negative samples do have significant differences in physical state.

[0089] Furthermore, cosine similarity is used to calculate the inner product between the anchor sample and the matching samples in the positive sample pair, as well as the inner product between the anchor sample and each sample in the candidate negative sample set. Combined with a preset temperature hyperparameter set to 0.07, the calculated inner products are substituted into the InfoNCE loss function for comparison, resulting in a bidirectional asynchronous cross-modal contrastive loss. The preset temperature hyperparameter is set to 0.07 because this value effectively differentiates the model's gradients for positive and negative sample distributions.

[0090] This embodiment does not include all Instead of directly treating cross-modal samples as negative samples, a filtered set of candidate negative samples is constructed for anchor samples. For anchor samples, only samples that do not belong to the same asynchronous sample group and do not belong to the same device's nearest time period are retained as candidate samples; further filtering is then performed based on label differences, operating condition differences, or clustering differences. If a sample has a defect category label, then when the sample group... With sample group satisfy At that time, the sample group As candidate negative samples, among which... and They represent the sample groups respectively. and sample group Defect category labels.

[0091] When labels are unavailable or incomplete, construct the condition vector for the sample group: in, They represent the sample groups respectively. The corresponding operating voltage, current, ambient temperature, and load rate. After normalizing the operating condition vector, a sample group is defined. With sample group The working distance between them is: in, and These represent the normalized operating condition vectors. When the following conditions are met... At that time, determine the sample group With sample group Significant differences exist in the operating conditions, and these are used as candidate negative samples. This is the preset working condition distance threshold.

[0092] Furthermore, cluster analysis can be performed on the samples. Let the sample groups be... and sample group The cluster centers are respectively and Then the cluster center distance is defined as: When satisfied At that time, determine the sample group With sample group These belong to a cluster of states with significant differences, and are considered as candidate negative samples. A preset clustering distance threshold is defined. Taking acoustic anchor points and thermal modes as examples, their candidate negative sample set is denoted as... The above screening mechanism can effectively suppress the false negative sample problem caused by samples with similar states being mistakenly selected as negative samples in asynchronous industrial data scenarios, thereby improving the stability and representation quality of cross-modal contrastive learning.

[0093] Taking acoustic and thermal modes as an example, from an acoustic perspective... To thermodynamics The contrast loss in direction is defined as: in, This represents the set of sample group indices in the current batch that simultaneously possess both acoustic and thermal modes. This indicates the number of sample groups in the set. Temperature is a hyperparameter used to adjust the model's attention to difficult negative samples and the similarity distribution; This represents the confidence weighting factor between acoustic and thermal pairings. The closer the sample pairs are in time, the larger the weighting factor, and the greater its role in training. Indicates the first Normalized representation of acoustic modes of each sample group in the shared latent space (anchor sample). Indicates the first Normalized representation of the thermal modes of each sample group in the shared latent space (positive samples). This indicates the calculation of the cosine similarity between two cross-modal feature vectors; This represents the thermal modal characteristics of candidate negative samples (or positive samples themselves) in the candidate negative sample set.

[0094] Contrast loss can be constructed in the same way from the thermal to the acoustic directions. Thus, a two-way contrast loss between acoustic and thermal modes is obtained: Following the same principle, acoustic and visual two-way contrast loss can also be constructed separately. and thermal and visual two-way contrast loss By introducing asynchronous pairing confidence weighting, spurious negative sample suppression, and bidirectional contrastive constraints, this embodiment can learn more stable cross-modal consistency representations under non-strict synchronization conditions, rather than relying solely on conventional unidirectional or coarse-grained contrastive optimization.

[0095] This embodiment effectively identifies and eliminates false negative samples with similar physical states by combining multiple conditions such as time interval, operating condition differences, and cluster distance for negative sample screening. This avoids the representation confusion caused by the network erroneously rejecting similar features during training. The use of a preset temperature hyperparameter to calculate the bidirectional contrast loss enhances the model's ability to distinguish difficult negative samples. This allows the network to bring homogeneous multimodal features closer together in the latent space while more reasonably distancing feature vectors representing different devices or different operating states. This improves the accuracy and anti-interference capability of heterogeneous sensing data in cross-modal representation alignment.

[0096] This application also provides a cable defect diagnosis device based on asynchronous contrastive learning, please refer to... Figure 3 The cable defect diagnosis device based on asynchronous contrastive learning includes: The sample merging module 10 is used to merge the acoustic data, infrared images and visible light images of the cable equipment to be tested according to the time tolerance window to obtain asynchronous multimodal sample groups. Feature projection module 20 is used to extract modal features and project shared latent space onto the acoustic data, infrared image and visible light image in the asynchronous multimodal sample group to obtain cross-modal normalized representation; The contrastive learning module 30 is used to perform spurious negative sample suppression based on the cross-modal normalized representation, so as to construct a training sample set containing positive sample pairs and candidate negative sample sets, and calculate bidirectional asynchronous cross-modal contrastive loss based on the training sample set. The weight calculation module 40 is used to perform degradation simulation processing and dimension unification mapping on the cross-modal standardized representation to obtain the mapping features, and calculate the quality-perceived attention weights based on the quality indicators and effective modal flags of each modality in the asynchronous multimodal sample group. The feature fusion module 50 is used to perform weighted fusion of the mapped features according to the quality-aware attention weights to obtain multimodal fusion features; The model diagnosis module 60 is used to jointly optimize and train the defect diagnosis model based on the bidirectional asynchronous cross-modal contrast loss and the multimodal fusion features, and output the cable defect diagnosis result through the trained defect diagnosis model.

[0097] The cable defect diagnosis device based on asynchronous contrastive learning provided in this application employs the cable defect diagnosis method based on asynchronous contrastive learning in the above embodiments. It solves the technical problem of how to stably characterize and robustly fuse multimodal asynchronous data from live-line testing of cables with degraded or missing data without relying on strict synchronous sampling and strong spatial registration, thereby achieving high-precision cable defect diagnosis. Compared with the prior art, the beneficial effects of the cable defect diagnosis device based on asynchronous contrastive learning provided in this application are the same as those of the cable defect diagnosis method based on asynchronous contrastive learning provided in the above embodiments, and other technical features in the cable defect diagnosis device based on asynchronous contrastive learning are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0098] This application provides a cable defect diagnosis device based on asynchronous contrastive learning. The cable defect diagnosis device based on asynchronous contrastive learning includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the cable defect diagnosis method based on asynchronous contrastive learning in the above embodiment 1.

[0099] The following is for reference. Figure 4The diagram illustrates a structural schematic of a cable defect diagnosis device based on asynchronous contrastive learning, suitable for implementing embodiments of this application. The cable defect diagnosis device based on asynchronous contrastive learning in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Android Devices), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The cable defect diagnosis device based on asynchronous contrastive learning shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0100] like Figure 4 As shown, the cable defect diagnosis device based on asynchronous contrastive learning may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in ROM 1002 (Read Only Memory) or the program loaded from storage device 1003 into RAM 1004 (Random Access Memory). RAM 1004 also stores various programs and data required for the operation of the oil depot fire simulation scenario construction device that integrates dual-model inversion. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. I / O interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the asynchronous contrastive learning-based cable defect diagnosis device to exchange data wirelessly or via wired communication with other devices. Although the figure shows an asynchronous contrastive learning-based cable defect diagnosis device with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0101] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0102] The cable defect diagnosis device based on asynchronous contrastive learning provided in this application, employing the cable defect diagnosis method based on asynchronous contrastive learning in the above embodiments, can solve the technical problem of how to stably characterize and robustly fuse multimodal asynchronous data of cable live-line detection with degradation or missing data without relying on strict synchronous sampling and strong spatial registration, thereby achieving high-precision cable defect diagnosis. Compared with the prior art, the beneficial effects of the cable defect diagnosis device based on asynchronous contrastive learning provided in this application are the same as those of the cable defect diagnosis method based on asynchronous contrastive learning provided in the above embodiments, and other technical features in this cable defect diagnosis device based on asynchronous contrastive learning are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0103] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0104] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0105] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the cable defect diagnosis method based on asynchronous contrastive learning in the above embodiments.

[0106] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash Memory), optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0107] The aforementioned computer-readable storage medium may be included in a cable defect diagnosis device based on asynchronous contrastive learning; or it may exist independently and not be assembled into a cable defect diagnosis device based on asynchronous contrastive learning.

[0108] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the asynchronous contrastive learning-based cable defect diagnosis device, cause the asynchronous contrastive learning-based cable defect diagnosis device to: merge the acoustic data, infrared image, and visible light image of the cable device to be tested according to a time tolerance window to obtain an asynchronous multimodal sample group; perform modal feature extraction and shared latent space projection on the acoustic data, infrared image, and visible light image in the asynchronous multimodal sample group to obtain a cross-modal normalized representation; and perform spurious negative sample suppression based on the cross-modal normalized representation to construct a training sample set containing positive sample pairs and candidate negative sample sets, and... The bidirectional asynchronous cross-modal contrastive loss is calculated based on the training sample set; the cross-modal standardized representation is subjected to degradation simulation processing and dimensionality unification mapping to obtain the mapping features; and the quality-aware attention weights are calculated based on the quality indicators and effective modality flags of each modality in the asynchronous multimodal sample set, wherein the effective modality flags are generated based on the acquisition state of each modality within the time tolerance window; the mapping features are weighted and fused based on the quality-aware attention weights to obtain multimodal fusion features; the defect diagnosis model is jointly optimized and trained based on the bidirectional asynchronous cross-modal contrastive loss and the multimodal fusion features, and the cable defect diagnosis results are output through the trained defect diagnosis model.

[0109] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0111] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0112] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described cable defect diagnosis method based on asynchronous contrastive learning. This solves the technical problem of how to stably characterize and robustly fuse multimodal asynchronous data from cable live-line testing that is degraded or missing, without relying on strict synchronous sampling and strong spatial registration, thereby achieving high-precision cable defect diagnosis. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the cable defect diagnosis method based on asynchronous contrastive learning provided in the above embodiments, and will not be repeated here.

[0113] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the cable defect diagnosis method based on asynchronous contrastive learning as described above.

[0114] The computer program product provided in this application solves the technical problem of how to stably characterize and robustly fuse multimodal asynchronous data of cable live-line detection that has suffered from quality degradation or missing data, without relying on strict synchronous sampling and strong spatial registration, thereby achieving high-precision cable defect diagnosis. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the cable defect diagnosis method based on asynchronous contrastive learning provided in the above embodiments, and will not be repeated here.

[0115] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A cable defect diagnosis method based on asynchronous contrastive learning, characterized in that, The method includes: Based on the time tolerance window, the acoustic data, infrared images, and visible light images of the cable equipment to be tested are merged to obtain an asynchronous multimodal sample group; Modal feature extraction and shared latent space projection are performed on the acoustic data, infrared images, and visible light images in the asynchronous multimodal sample group to obtain cross-modal normalized representations; Based on the cross-modal normalized representation, false negative sample suppression is performed to construct a training sample set containing positive sample pairs and candidate negative sample sets, and bidirectional asynchronous cross-modal contrastive loss is calculated based on the training sample set. The cross-modal standardized representation is subjected to degradation simulation processing to obtain degradation features. The degradation features are then subjected to dimensional uniform mapping to obtain mapping features. Quality-aware attention weights are calculated based on the quality indicators and effective modality flags of each modality in the asynchronous multimodal sample group. The effective modality flag is set to 1 when the corresponding modality's message data exists within the time tolerance window, the file size of the message data is greater than zero, and the file structure is complete; otherwise, it is set to 0. The mapped features are weighted and fused according to the quality-perceived attention weights to obtain multimodal fusion features; The defect diagnosis model is jointly optimized and trained based on the bidirectional asynchronous cross-modal contrast loss and the multimodal fusion features, and the cable defect diagnosis result is output through the trained defect diagnosis model.

2. The method as described in claim 1, characterized in that, The step of performing spurious negative sample suppression based on the cross-modal normalized representation to construct a training sample set containing positive sample pairs and a candidate negative sample set, and calculating the bidirectional asynchronous cross-modal contrastive loss based on the training sample set includes: The cross-modal normalized representations belonging to the same asynchronous multimodal sample group are combined to obtain positive sample pairs, and any one of the cross-modal normalized representations in the positive sample pair is determined as an anchor sample. Based on the anchor sample, obtain the device identifier and reference timestamp of the remaining cross-modal standardized representations, and remove samples whose device identifier is the same as the device identifier of the anchor sample and whose time interval between the reference timestamp and the reference timestamp of the anchor sample is less than a preset time threshold, to obtain an initial negative sample set. Calculate the operating condition vector distance and cluster center distance between the anchor sample in the positive sample pair and each sample in the initial negative sample set. The operating condition vector distance refers to the numerical difference between two samples in operating voltage, load current, ambient temperature and load rate calculated by the Euclidean distance algorithm. Samples whose working condition vector distance is greater than a preset working condition distance threshold and whose cluster center distance is greater than a preset clustering threshold are retained in the candidate negative sample set; The bidirectional asynchronous cross-modal contrast loss is obtained by performing an inner product operation on the positive sample pairs, the candidate negative sample set, and the preset temperature hyperparameter.

3. The method as described in claim 1, characterized in that, The step of merging the acoustic data, infrared images, and visible light images of the cable equipment under test according to the time tolerance window to obtain an asynchronous multimodal sample group includes: Obtain the reference time of the cable equipment to be tested, and determine the preset time tolerance window corresponding to the reference time; Retrieve acoustic data, infrared images, and visible light images acquired within the preset time tolerance window, and read the timestamps of the acoustic data, infrared images, and visible light images; Calculate the time difference between the timestamp and the reference time, and calculate the asynchronous pairing confidence based on the time difference; Determine whether there is message data of the corresponding modality within the preset time tolerance window, and generate the valid modality flag bit according to the existence status and content integrity of the message data; The acoustic data, the infrared image, and the visible light image, which are within the preset time tolerance window and contain the message data, are correlated to obtain an asynchronous multimodal sample group.

4. The method as described in claim 1, characterized in that, The step of extracting modal features and projecting shared latent space onto the acoustic data, infrared images, and visible light images in the asynchronous multimodal sample group to obtain cross-modal normalized representations includes: The acoustic data in the asynchronous multimodal sample group is converted into a two-dimensional time-frequency tensor by short-time Fourier transform, and the regions of interest are cropped for the infrared image and the visible light image. The two-dimensional time-frequency tensor, the cropped infrared image, and the cropped visible light image are input into the corresponding depth encoder to calculate high-level features and obtain modal high-level features. In the case where the effective modality flag indicates that a modality is missing, the high-level features of the modality are filled using a preset placeholder vector; The filled high-level features of the modality are mapped into a latent space of uniform dimension through a projection head network; The mapped features are normalized to obtain a cross-modal standardized representation.

5. The method as described in claim 1, characterized in that, The step of calculating the quality-perceived attention weight based on the quality index and effective modality flag bits of each modality in the asynchronous multimodal sample group includes: The Laplacian variance of the visible light image is calculated as a visual quality index, and the signal-to-noise ratio of the acoustic data is calculated as an acoustic quality index. The average temperature difference between the target area and the background area in the infrared image is calculated as a thermal quality indicator. The visual quality index, the acoustic quality index, and the thermal quality index are concatenated to obtain a multimodal quality vector; The multimodal quality vector is input into the attention weight calculation network to obtain the initial fusion weights corresponding to each modality; The initial fusion weights are multiplied by the effective modality flag of the corresponding modality to obtain the effective initial weights; The effective initial weights are normalized to obtain the quality-aware attention weights.

6. The method as described in claim 1, characterized in that, The steps of performing degradation simulation processing on the cross-modal normalized representation to obtain degraded features, and performing dimensionality-unified mapping on the degraded features to obtain mapped features include: Random noise injection or masking is performed on the cross-modal normalized representation with a preset discard probability to obtain degraded features; The degraded features are transformed into uniform-dimensional mapped features through a linear mapping layer.

7. The method according to any one of claims 1 to 6, characterized in that, The step of jointly optimizing and training the defect diagnosis model based on the bidirectional asynchronous cross-modal contrast loss and the multimodal fusion features, and outputting the cable defect diagnosis result through the trained defect diagnosis model includes: The multimodal fusion features are input into the defect classifier to obtain the predicted probability that the cable equipment to be inspected belongs to each preset defect category; Calculate the cross-entropy classification loss based on the predicted probabilities and the actual defect labels; A joint loss function is constructed based on the bidirectional asynchronous cross-modal contrastive loss and the cross-entropy classification loss; Based on the joint loss function, the network parameters of the defect diagnosis model are updated through the backpropagation algorithm to obtain the trained defect diagnosis model; The target asynchronous data is input into the trained defect diagnosis model to obtain the cable defect diagnosis results.

8. A cable defect diagnosis device based on asynchronous contrastive learning, characterized in that, The device employs the cable defect diagnosis method based on asynchronous contrastive learning as described in any one of claims 1 to 7, and the device comprises: The sample merging module is used to merge the acoustic data, infrared images, and visible light images of the cable equipment to be tested according to the time tolerance window to obtain asynchronous multimodal sample groups. The feature projection module is used to extract modal features and project shared latent space onto the acoustic data, infrared image, and visible light image in the asynchronous multimodal sample group to obtain cross-modal normalized representations. The contrastive learning module is used to perform spurious negative sample suppression based on the cross-modal normalized representation, so as to construct a training sample set containing positive sample pairs and candidate negative sample sets, and calculate bidirectional asynchronous cross-modal contrastive loss based on the training sample set. The weight calculation module is used to perform degradation simulation processing on the cross-modal standardized representation to obtain degradation features, perform dimensional unification mapping on the degradation features to obtain mapping features, and calculate quality-perceived attention weights based on the quality indicators and effective modality flags of each modality in the asynchronous multimodal sample group. The feature fusion module is used to perform weighted fusion of the mapped features according to the quality-aware attention weights to obtain multimodal fusion features; The model diagnosis module is used to jointly optimize and train the defect diagnosis model based on the bidirectional asynchronous cross-modal contrast loss and the multimodal fusion features, and output the cable defect diagnosis result through the trained defect diagnosis model.

9. A cable defect diagnosis device based on asynchronous contrastive learning, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the cable defect diagnosis method based on asynchronous contrastive learning as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the cable defect diagnosis method based on asynchronous contrastive learning as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Defect diagnosis method and system based on multi-modal data cooperative training

    CN120670964A

  • System and method for using three dimensional infrared imaging to provide detailed anatomical structure maps

    US20100172567A1