Fault detection method and device of power transmission line and nonvolatile storage medium
By collecting and fusing images, infrared thermal images, and text data of transmission lines, and combining them with cross-modal geometric constraints, a target fault detection model was used to solve the problem of low accuracy in transmission line fault detection results, achieving high-precision fault detection and intelligent inspection.
Patent Information
- Application Number
- CN202511170657.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-12-05
AI Technical Summary
Existing technologies have low accuracy in detecting transmission line faults, manual inspection is inefficient and costly, and automated equipment is susceptible to environmental interference, resulting in large errors in detection results.
By collecting image data, infrared thermal image data, and text data of transmission lines, and combining cross-modal fusion features and target fault detection models with cross-modal geometric constraints, the fault type and location of transmission lines can be identified, achieving high-precision fault detection.
It improves the accuracy and reliability of transmission line fault detection, realizes high-precision, real-time six-dimensional location information detection of transmission line faults, and enhances the intelligence and precision of inspection and maintenance.
Smart Images

Figure CN121069093A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power systems, and more specifically, to a method, apparatus, and non-volatile storage medium for detecting faults in transmission lines. Background Technology
[0002] The safe operation of transmission lines is crucial for ensuring the stability of power supply and the reliability of the power grid. However, transmission lines are exposed to the elements for extended periods, making them susceptible to various faults such as insulator cracks, hardware corrosion, conductor wear, and foreign object suspension. These faults can not only reduce power transmission efficiency but also trigger serious power accidents or even grid collapse, causing enormous losses and damage. Therefore, regular fault detection of transmission lines is essential for maintaining the safe and stable operation of the power grid.
[0003] In related technologies, fault detection of transmission lines is typically performed using manual visual inspection or simple automated equipment. Manual visual inspection is not only inefficient and costly, but the accuracy of the results also depends heavily on the inspector's skill. Using simple automated equipment for fault detection is susceptible to environmental interference and has poor environmental adaptability, leading to significant errors in the detection results. Therefore, related technologies suffer from the technical problem of low accuracy in fault detection results for transmission lines.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a method, apparatus, and non-volatile storage medium for detecting faults in transmission lines, in order to at least solve the technical problem of low accuracy of fault detection results for transmission lines in related technologies.
[0006] According to one aspect of the embodiments of this application, a fault detection method for a transmission line is provided, comprising: acquiring current image data, current infrared thermal image data, and current text data of the transmission line, wherein the current infrared thermal image data is used to describe the current temperature distribution of the transmission line; determining current cross-modal fusion features of the transmission line based on the current image data, current infrared thermal image data, and current text data, wherein the current cross-modal fusion features are used to fuse information from the current image data, current infrared thermal image data, and current text data; and obtaining a first fault detection result of the transmission line based on the current cross-modal fusion features using a fault type identification module included in a target fault detection model, wherein the first fault detection result includes at least: The fault type and initial fault location are defined. The initial fault location describes the two-dimensional location information of the transmission line fault. The target fault detection model is used to determine the fault type and fault location of the transmission line. Based on the current cross-modal geometric constraints of the transmission line, the current cross-modal fusion features, and the initial fault location included in the first fault detection result, the fault location identification module included in the target fault detection model is used to obtain the second fault detection result of the transmission line. The second fault detection result describes the six-dimensional location information of the transmission line fault, which includes three-dimensional location information and three-dimensional pose information. The current cross-modal geometric constraints are used to ensure the consistency of the current image data, current infrared thermal image data, and current text data in the spatial coordinate system.
[0007] According to another aspect of the embodiments of this application, a fault detection device for a transmission line is provided, comprising: a data acquisition module, configured to acquire current image data, current infrared thermal image data, and current text data of the transmission line, wherein the current infrared thermal image data is used to describe the current temperature distribution of the transmission line; a current cross-modal fusion feature determination module, configured to determine current cross-modal fusion features of the transmission line based on the current image data, current infrared thermal image data, and current text data, wherein the current cross-modal fusion features are used to fuse information from the current image data, current infrared thermal image data, and current text data; and a first fault detection result determination module, configured to obtain a first fault detection result of the transmission line based on the current cross-modal fusion features and using a fault type identification module included in a target fault detection model, wherein... The first fault detection result includes at least: fault type and initial fault location. The initial fault location is used to describe the two-dimensional location information of the transmission line fault. The target fault detection model is used to determine the fault type and fault location of the transmission line. The second fault detection result determination module is used to obtain the second fault detection result of the transmission line based on the current cross-modal geometric constraints of the transmission line, the current cross-modal fusion features, and the initial fault location included in the first fault detection result, using the fault location identification module included in the target fault detection model. The second fault detection result is used to describe the six-dimensional location information of the transmission line fault. The six-dimensional location information includes three-dimensional location information and three-dimensional pose information. The current cross-modal geometric constraints are used to ensure the consistency of the current image data, the current infrared thermal image data, and the current text data in the spatial coordinate system.
[0008] According to another aspect of the embodiments of this application, a non-volatile storage medium is provided, which stores a plurality of instructions, any one of which is adapted to be loaded by a processor for a fault detection method for a power transmission line.
[0009] According to another aspect of the embodiments of this application, an electronic device is provided, including: one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement any one of the fault detection methods for power transmission lines.
[0010] According to another aspect of the embodiments of this application, a computer program product is provided, which, when executed on a data processing device, is a program adapted to perform the steps of a fault detection method for power transmission lines.
[0011] In this embodiment, current image data, current infrared thermal image data, and current text data of the transmission line are collected, wherein the current infrared thermal image data is used to describe the current temperature distribution of the transmission line; based on the current image data, current infrared thermal image data, and current text data, current cross-modal fusion features of the transmission line are determined, wherein the current cross-modal fusion features are used to fuse information from the current image data, current infrared thermal image data, and current text data; based on the current cross-modal fusion features, a fault type identification module included in the target fault detection model is used to obtain a first fault detection result of the transmission line, wherein the first fault detection result includes at least: fault type and initial fault location. The initial fault location is used to describe the two-dimensional location information of the transmission line fault. The target fault detection model is used to determine the fault type and location of the transmission line. Based on the current cross-modal geometric constraints of the transmission line, the current cross-modal fusion features, and the initial fault location included in the first fault detection result, the fault location identification module included in the target fault detection model is used to obtain the second fault detection result of the transmission line. The second fault detection result is used to describe the six-dimensional location information of the transmission line fault, including three-dimensional location information and three-dimensional pose information. The current cross-modal geometric constraints are used to ensure the consistency of the current image data, current infrared thermal image data, and current text data in the spatial coordinate system. This achieves the goal of determining the cross-modal fusion features of the transmission line by fusing image data, infrared thermal image data, and text data, and combining the cross-modal geometric constraints of the transmission line with the target fault detection model to determine the fault detection result of the transmission line. This improves the accuracy of the fault detection result of the transmission line, thereby solving the technical problem of low accuracy in fault detection results of transmission lines in related technologies. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0013] Figure 1 This is a flowchart of a fault detection method for a power transmission line according to an embodiment of this application;
[0014] Figure 2 This is a flowchart of an optional fault detection method for transmission lines provided according to an embodiment of this application;
[0015] Figure 3 This is a schematic diagram of an optional fault detection device for a power transmission line according to an embodiment of this application. Detailed Implementation
[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0018] For ease of description, the following explains some of the nouns or terms used in the embodiments of this application:
[0019] The VLM-Gym-Transmission environment is a reinforcement learning platform built in a virtual environment for researching and developing visual language models. Gym refers to OpenAI Gym, a reinforcement learning environment used to test and develop learning algorithms for intelligent agents. Transmission refers to the process by which a visual language model transmits or conveys instructions by observing the environment (visual information) and understanding the instructions (textual information). In this environment, the model predicts or performs actions based on the observed visual scene and given textual instructions to complete specific tasks, such as moving to a target location, recognizing objects, or performing a series of operations.
[0020] ViT-B / 16, or Vision Transformer (ViT), is a novel architecture for image classification. "B" stands for Base, and "16" indicates that the image is segmented into 16x16 blocks for processing. ViT-B / 16 can extract rich semantic information from images, enabling understanding and analysis of image content.
[0021] Swin Transformer is a variant of the Transformer architecture for processing image data. It improves the standard Transformer model by introducing a "window multi-head self-attention" mechanism, making it more suitable for image recognition tasks.
[0022] ViT-G (Vision Transformer–Global) is a large-scale, high-performance computer vision model based on the Vision Transformer (ViT) architecture.
[0023] RGB images are a color image format based on the three primary colors of red, green, and blue. R stands for Red, G for Green, and B for Blue.
[0024] According to an embodiment of this application, a method embodiment for fault detection of transmission lines is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0025] Figure 1 This is a flowchart of a fault detection method for power transmission lines according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0026] Step S102: Collect current image data, current infrared thermal image data, and current text data of the transmission line, wherein the current infrared thermal image data is used to describe the current temperature distribution of the transmission line.
[0027] It is understood that collecting current image data, current infrared thermal image data, and current text data of the transmission line to be inspected is necessary. The current infrared thermal image data is used to describe the current temperature distribution of the transmission line. The current image data is used to obtain the current target image features of the transmission line, the current infrared thermal image data is used to obtain the current target temperature features of the transmission line, and the current text data, such as equipment ledgers and historical maintenance records, is used to obtain the current text features of the transmission line. By collecting image data, infrared thermal image data, and text data of the transmission line, multi-source information coverage can be ensured, enabling a comprehensive assessment of the transmission line's condition from multiple perspectives, including visual, temperature, and historical records, thereby improving the accuracy and reliability of fault detection results.
[0028] In one optional embodiment, acquiring current image data of the transmission line includes: determining the acquisition strategy of the image acquisition device based on the current environmental data of the transmission line using an acquisition strategy optimization model, wherein the acquisition strategy includes at least the acquisition path and acquisition pose of the image acquisition device, and the acquisition strategy optimization model pre-learns the correlation between the environmental data of the transmission line and the acquisition strategy; and acquiring current image data according to the acquisition strategy.
[0029] It is understandable that before acquiring current image data of a transmission line, the acquisition strategy of the image acquisition equipment needs to be determined. Based on the current environmental data of the transmission line, a pre-trained acquisition strategy optimization model is used to determine the acquisition strategy of the image acquisition equipment, including the acquisition path and acquisition pose. This acquisition strategy optimization model pre-learns the correlation between the environmental data of the transmission line and the acquisition strategy. Following this acquisition strategy, images of the transmission line are acquired, obtaining the current image data of the transmission line. Through the acquisition strategy optimization model and the current environmental data of the transmission line, the optimal acquisition strategy for the image acquisition equipment is dynamically generated, ensuring that high-quality image data can be captured even in complex environments, providing high-quality image data for subsequent fault detection.
[0030] Optionally, the above acquisition strategy may include not only the acquisition path and acquisition pose of the image acquisition device, but also the magnification of the acquisition object by the acquisition device.
[0031] Optionally, the data acquisition strategy optimization model can be trained and optimized by constructing a transmission line inspection environment. The inspection environment modeling can be achieved by constructing a VLM-Gym-Transmission environment to simulate transmission line inspection scenarios (such as different weather conditions, lighting conditions, and drone altitudes), and defining the reward function R for the data acquisition strategy optimization model. The reward function R can be determined as follows:
[0032] R=λ1R det +λ2R pose +λ3R reg
[0033] Among them, R det To detect mAP (mean Average Precision) reward, R pose For the pose estimation ADD-S (Average Distance of Model Points after Symmetry Transformation) error reward, R reg The reward is the characteristic rank regularization reward, where λ1 = 0.6, λ2 = 0.3, and λ3 = 0.1.
[0034] Optionally, the acquisition strategy and optimization model can be optimized as follows: The GRPO (Greedy Randomized Path Optimization) algorithm is used to sample 10 action hypotheses in parallel (e.g., "move left 10cm", "magnify by 2 times"), and the advantage function A is used to... s,i,t Normalization is used to screen for high-value actions. Advantage function A s,i,t It can be determined in the following way:
[0035]
[0036] Where s is the current state; a i For the i-th action, the clip mechanism is used. Limit the policy update magnitude to avoid overfitting; R(s,a) i Given the current state s, perform action a. i The "benefits" (such as improved fault detection accuracy and improved data acquisition efficiency); mean(R(s,a1,...,a) 10 )) for all candidate actions (a1~a 10 The average return represents the benchmark return level; std(R(s,a1,...,a) 10 The standard deviation of the returns for all candidate actions represents the degree of return volatility. The logic of the above formula is to convert the action value into the degree of deviation from the average level by dividing (current action return - average return) by the standard deviation of returns.
[0037] Optionally, if A s,i,t >0 indicates action a i The returns are better than the average return level, and the action value is higher; if A s,i,t <0 indicates action a i The returns are lower than the average, and the action value is even lower. Based on the above value function, high-value actions (such as "shift left by 10cm" and "magnify by 2x") are selected to optimize the acquisition strategy. In the fault detection scenario of transmission lines, the acquisition strategy optimization model traverses all candidate actions and determines the corresponding advantage function value. This advantage function value considers not only the return of the candidate action itself, but also the relationship with other candidate actions. It can be used to measure the benefit of the acquisition equipment executing candidate actions for fault detection of transmission lines, avoiding blindly selecting high-return but practically useless candidate actions as the acquisition strategy of the acquisition equipment.
[0038] Step S104: Based on the current image data, current infrared thermal image data, and current text data, determine the current cross-modal fusion features of the transmission line, wherein the current cross-modal fusion features are used to fuse information from the current image data, current infrared thermal image data, and current text data;
[0039] It is understandable that cross-modal data fusion is performed on the collected current image data, current infrared thermal image data, and current text data of the transmission line to obtain the current cross-modal fusion features of the transmission line, which are used to fuse the information of the current image data, current infrared thermal image data, and current text data. Through cross-modal feature fusion technology, information from image data, infrared thermal image data, and text data is effectively combined, thereby overcoming the limitations of single-modal data processing and improving the accuracy and reliability of fault detection.
[0040] In an optional embodiment, before obtaining the first fault detection result of the transmission line by using the fault type identification module included in the target fault detection model based on the current cross-modal fusion features, the method further includes: acquiring training image data, training infrared thermal image data, and training text data of the transmission line; obtaining training cross-modal fusion features of the transmission line based on the training image data, training infrared thermal image data, and training text data; and training the initial fault detection model using the training cross-modal fusion features and the training cross-modal geometric constraints of the transmission line to obtain the target fault detection model.
[0041] It is understandable that before utilizing the current cross-modal fusion features of the transmission line and employing the fault type identification module included in the target fault detection model to obtain the first fault detection result of the transmission line, it is necessary to construct the aforementioned target fault detection model. This involves acquiring training image data, training infrared thermal image data, and training text data of the transmission line. Cross-modal data fusion is then performed on the aforementioned training image data, training infrared thermal image data, and training text data to obtain the training cross-modal fusion features of the transmission line. The initial fault detection model is then trained using the training cross-modal fusion features and the training cross-modal geometric constraints of the transmission line to obtain the target fault detection model. By collecting and processing a large amount of multimodal training data, the initial fault detection model can be specifically trained, enabling it to fully learn and understand the image features, temperature features, and semantic features of the transmission line equipment, thereby improving the recognition accuracy and generalization ability of the target fault detection model.
[0042] In one optional embodiment, training cross-modal fusion features of a transmission line are obtained based on training image data, training infrared thermal image data, and training text data, including: extracting features from the training image data to obtain training target image features of the transmission line; extracting features from the training infrared thermal image data to obtain training target temperature features of the transmission line; extracting features from the training text data to obtain training target text features of the transmission line; and determining training cross-modal fusion features based on the training target image features, training target temperature features, and training target text features.
[0043] It is understandable that feature extraction is performed separately on the training image data, training infrared thermal image data, and training text data of the transmission line to obtain the training target image features, training target temperature features, and training target text features of the transmission line. These training target image features, training target temperature features, and training target text features are then fused to obtain the training cross-modal fusion features of the transmission line. By integrating image features, temperature features, and text features, the training cross-modal fusion features can more comprehensively reflect the state of the transmission line and improve the target fault detection model's ability to detect minute and hidden defects.
[0044] In one optional embodiment, feature extraction is performed on the training image data to obtain training target image features of the transmission line, including: extracting features from the training image data to obtain initial training image features of the transmission line; performing Fourier transform on the initial training image features to obtain frequency domain image features of the transmission line; segmenting the frequency domain image features based on a preset frequency band segmentation threshold to obtain frequency domain image sub-features corresponding to multiple frequency bands; using inverse Fourier transform on the frequency domain image sub-features corresponding to multiple frequency bands to obtain time domain image sub-features corresponding to multiple frequency bands; and determining the training target image features of the transmission line based on the time domain image sub-features corresponding to multiple frequency bands.
[0045] It is understandable that feature extraction is performed on the training image data of transmission lines to obtain the training target image features of the transmission lines. Feature extraction is performed on the training image data of the transmission lines to obtain the initial training image features. Fourier transform is used to perform frequency domain transformation on the initial training image features to obtain the frequency domain image features of the transmission lines. The frequency domain image features are then segmented according to a preset frequency band segmentation threshold to obtain frequency domain image sub-features corresponding to multiple frequency bands. Inverse Fourier transform is used to perform time domain transformation on the frequency domain image sub-features corresponding to the above multiple frequency bands to obtain time domain image sub-features corresponding to the above multiple frequency bands. Based on the time domain image sub-features corresponding to the above multiple frequency bands, the training target image features of the transmission lines are obtained. Through frequency domain transformation and Fourier transform analysis, the feature representation capability of the training image data can be improved, providing high-quality feature input for subsequent cross-modal feature fusion and initial fault detection model training, thereby enhancing the accuracy of the target fault detection model and its applicability in complex environments.
[0046] Optionally, the CLIP (Contrastive Language-Image Pre-training) visual encoder (ViT-B / 16) can be used to extract multi-scale features (32×32×768) and (16×16×1024) from the inspection images (i.e., training image data). Infrared thermal images (i.e., training infrared thermal image data) can be extracted using a thermal infrared convolution branch (kernel size 5×5) to extract features of abnormal temperature regions (i.e., training target temperature features). Equipment ledger text (i.e., training text data) can be used to generate semantic vectors (dimension 512) through the CLIP text encoder, and then fused with visual features through a cross-modal attention mechanism to obtain the training cross-modal fusion feature F. cross F cross It can be determined in the following way:
[0047] F cross =MultiHeadAttn(F vis ,F txt )+F vis
[0048] Among them, F vis For visual features, F txt For text features, feature interactions are enhanced through residual connections.
[0049] Optionally, an FDConv (Frequency Domain Convolutional Layer) layer can be embedded in the CLIP visual encoder to decompose the convolutional kernel into three frequency bands: low, medium, and high (with preset frequency band segmentation thresholds of 0.1 and 0.5, respectively). Simultaneously, frequency domain feature decoupling can be achieved through Fourier transform. The convolutional kernels W for the low, medium, and high frequency bands... s ,s∈{low,mid,high} can be determined in the following way:
[0050] W s =iDFT(M s ·DFT(W)), s∈{low,mid,high}
[0051] Among them, M s Here, W is the frequency band mask, W is the original convolution kernel, low represents the low frequency band, mid represents the mid frequency band, and high represents the high frequency band. Inverse Fourier transform can be used to generate spatial kernels for each frequency band, capturing the frequency domain image sub-features corresponding to multiple frequency bands of the transmission line, including contours (frequency domain image sub-features for low frequencies), textures (frequency domain image sub-features for mid frequencies), and edge details (frequency domain image sub-features for high frequencies).
[0052] In one optional embodiment, an initial fault detection model is trained using cross-modal fusion features and cross-modal geometric constraints of the transmission line to obtain a target fault detection model. This includes: determining a loss function for the initial fault detection model based on the detection loss and pose loss of the transmission line, wherein the detection loss is used to quantify the error of the first detection result, and the pose loss is used to quantify the error of the second detection result; adjusting the initial parameter values of the target parameters of the initial fault detection model based on the function value corresponding to the loss function and a preset loss function threshold, wherein the target parameters are the parameters of the adapter included in the initial fault detection model, and the adapter is used to fuse image data, infrared thermal image data, and text data; when the function value corresponding to the loss function is less than or equal to the preset loss function threshold, the parameter value of the target parameter is determined as the target parameter value; and determining the target fault detection model based on the target parameter value.
[0053] It is understandable that the loss function of the initial fault detection model is determined based on the detection loss and pose loss of the transmission line. The aforementioned detection loss is used to quantify the error of the first detection result, and the pose loss is used to quantify the error of the second detection result. Using the training cross-modal fusion features and training cross-modal geometric constraints as input, the difference between the function value of the loss function and the preset loss function threshold is calculated, and the initial parameter values of the target parameters of the initial fault detection model are adjusted according to this difference. The target parameters of the initial fault detection model refer to the parameters of the adapter included in the initial fault detection model, which is used to fuse image data, infrared thermal imaging data, and text data. The target parameters of the initial fault detection model are continuously adjusted according to the above process until the function value corresponding to the loss function is less than or equal to the preset loss function threshold. The parameter value at this point is determined as the target parameter value of the initial fault detection model, and the target fault detection model is determined based on this target parameter value. By optimizing the adapter parameters, the target fault detection model can better fuse image data, infrared thermal imaging data, and text data information, enhancing the understanding and recognition capabilities of complex transmission line inspection scenarios.
[0054] Optionally, a Mona Adapter can be inserted after each Swin Transformer block of the CLIP visual encoder, containing a scale normalization layer and a multi-branch convolutional group. The scale normalization layer can adjust the distribution of the input feature x0 using learnable parameters (i.e., target parameters) s1 and s2 to obtain the adjusted feature x. norm = s1·LayerNorm(x0) + s2·x0. Multi-branch convolutional groups can extract multi-scale geometric features through parallel 3×3, 5×5, and 7×7 depthwise convolutions, and aggregate them through 1×1 convolutions. in, For depthwise convolution kernels, w pwThe kernel is a point convolution, and the original features are preserved through skip connections, f dw The output features are the fusion results of multi-scale depthwise convolutional branches, aggregating geometric features of different sizes, where i is the i-th layer and f is the output feature. pw This is the fusion result of the point convolution branches, supplemented with local detail features, where x is the input feature value. For convolution operations, This is a weighted convolution operation.
[0055] Optionally, the target fault detection model can be optimized by freezing the CLIP backbone network and updating only the adapter parameters (approximately 2.56% of the total parameters). In transmission line fault detection tasks, this method can reduce FLOPs (Floating Point Operations per Second) from 1820G to 1150G, and reduce the number of parameters by 28%.
[0056] In an optional embodiment, before adjusting the initial parameter values of the target parameters of the initial fault detection model based on the function value corresponding to the loss function and a preset loss function threshold, the method further includes: inputting the trained cross-modal fusion features into the initial fault detection model to obtain a third fault detection result of the transmission line; inputting the trained cross-modal fusion features into the teacher model respectively to obtain a fourth fault detection result of the transmission line, wherein the teacher model is used to guide the learning process of the initial fault detection model; determining the detection error between the initial fault detection model and the teacher model based on the third fault detection result and the fourth fault detection result; and determining the initial parameter values based on the detection error and a preset error threshold.
[0057] It is understandable that the cross-modal fusion features of the transmission line training are input into the initial fault detection model and the teacher model, respectively, to obtain the third and fourth fault detection results of the transmission line. The teacher model is a higher-performing and more thoroughly trained model used to guide the learning process of the initial fault detection model. The third and fourth fault detection results are compared and analyzed to determine the detection error between the initial fault detection model and the teacher model. Based on the detection error and a preset error threshold, the initial parameter values of the initial fault detection model are determined. Through the guidance of the teacher model, the learning process of the initial fault detection model can be accelerated, the training cost reduced, and the initial fault detection model can reach the expected performance level in a shorter time.
[0058] Optionally, based on the detection error between the initial fault detection model and the teacher model, the initial parameter values of the initial fault detection model can be determined as follows: Compare the third and fourth fault detection results to quantify the difference between the two models and determine the detection error between the initial fault detection model and the teacher model in defect detection. Compare the detection error with a preset error threshold. If the detection error exceeds the preset error threshold, it indicates that the performance of the initial fault detection model still has room for improvement. Then, based on the detection error, use contrastive learning and knowledge distillation techniques to adjust the parameter values of the target parameters of the initial fault detection model (i.e., the parameter values of the adapter parameters of the initial fault detection model) to narrow the performance gap with the teacher model. Iteratively optimize the parameter values of the target parameters according to the above process until the detection error is less than or equal to the preset error threshold. This indicates that the performance of the initial fault detection model meets the requirements, and the parameter values of the target parameters at this point are determined as the initial parameter values of the initial fault detection model.
[0059] Optionally, a multi-task loss function L is constructed. L can be determined as follows:
[0060] L = L det +L pose +L reg
[0061] FocalLoss+DIoULoss
[0062]
[0063] Among them, L det To detect loss; L pose For pose loss, L2 loss is used in conjunction with neural implicit field rendering error; L regFeature rank regularization loss forces the feature matrix to have a full rank of ≥95% to avoid dimensionality collapse; L2 loss is Euclidean loss (i.e., L2 loss), which measures the accuracy of the prediction by calculating the "distance" between two vectors (such as pose or implicit field); the smaller the "distance" value, the more accurate the prediction; FocalLoss is a loss function used to solve fault detection problems with "few fault samples and many background samples," such as detecting small faults in a large area of normal conductors (background) of a transmission line. This loss function can make the fault detection model focus more on faults that are difficult to distinguish. Sample; DIoULoss is an improved bounding box loss function that considers not only the overlap between the predicted and ground truth boxes but also the positional difference between them, making fault box prediction more accurate; y is the ground truth label, used to mark whether a fault exists in the current region, 1 if a fault exists, 0 otherwise; γ is the focus coefficient of FocalLoss, which controls the degree of attention the fault detection model pays to samples that are more difficult to detect (difficult-to-distinguish faults / background), the larger the value, the more attention it pays to samples that are more difficult to detect; R pred For the predicted fault area bounding box, such as the predicted location range of a broken conductor strand; R gt The original fault region bounding boxes are manually labeled and used to determine the accuracy of the predicted fault region bounding boxes; N is the number of samples used in the pose loss calculation; pose pred The fault pose predicted by the fault detection model includes 3D position (x, y, z coordinates) and 3D attitude (rotation angles around the x, y, z axes), used to describe the fault's location and pose in space; pose gt λ represents the actual fault pose, used to determine the accuracy of the predicted fault pose; λ is a weighting coefficient used to balance the contributions of the pose loss term and the neural implicit field loss term to the total loss; SDF(x) is the "neural implicit field" predicted by the fault detection model, used to describe the distance from each point in space to the fault surface using a mathematical function, assisting in 3D pose and shape prediction; SDF gt (x) represents the actual neural implicit field, calculated based on the 3D shape of the actual fault, and is used to determine whether the predicted "neural implicit field" is accurate.
[0064] Optionally, a teacher model can be used to generate "defect description-image region" pairs (i.e., third fault detection results), and a distillation loss function L can be constructed through comparative learning. distill Initialize the target parameters of the adapter in the target fault detection model. distill It can be determined in the following way:
[0065]
[0066] Among them, f visFor example, image feature vectors, such as the visual features of fault areas in image data; f txt τ is the text feature vector, such as the semantic vector describing defects like "broken strands in the conductor" or "dirty insulator" in text data; τ is the temperature coefficient, used to control the "smoothness" of the probability distribution and adjust the ease of knowledge transfer.
[0067] Step S106: Based on the current cross-modal fusion features, the fault type identification module included in the target fault detection model is used to obtain the first fault detection result of the transmission line. The first fault detection result includes at least: fault type and initial fault location. The initial fault location is used to describe the two-dimensional location information of the fault in the transmission line. The target fault detection model is used to determine the fault type and fault location of the transmission line.
[0068] It is understandable that by inputting the current cross-modal fusion features into the target fault detection model, the fault type identification module in the target fault detection model analyzes the aforementioned current cross-modal fusion features to obtain the first fault detection result of the transmission line. This first fault detection result includes at least the fault type of the transmission line and the initial fault location using two-dimensional location information to describe the fault. Through in-depth mining of the current cross-modal fusion features by the target fault detection model, multiple fault types in the transmission line can be accurately identified, thereby enabling timely detection and handling of potential safety hazards and ensuring the safe and stable operation of the transmission line.
[0069] Step S108: Based on the current cross-modal geometric constraints of the transmission line, the current cross-modal fusion features, and the initial fault location included in the first fault detection result, the fault location identification module included in the target fault detection model is used to obtain the second fault detection result of the transmission line. The second fault detection result is used to describe the six-dimensional location information of the transmission line fault. The six-dimensional location information includes three-dimensional location information and three-dimensional pose information. The current cross-modal geometric constraints are used to ensure the consistency of the current image data, the current infrared thermal image data, and the current text data in the spatial coordinate system.
[0070] It is understandable that by utilizing the fault location identification module included in the target fault detection model, the current cross-modal geometric constraints, current cross-modal fusion features, and initial fault location included in the first fault detection result of the transmission line are identified and analyzed to obtain a second fault detection result that describes the six-dimensional location information of the transmission line fault. The aforementioned six-dimensional location information includes three-dimensional location information and three-dimensional pose information. The current cross-modal geometric constraints are used to ensure the consistency of the current image data, current infrared thermal image data, and current text data in the spatial coordinate system. Based on the cross-modal geometric constraints and cross-modal fusion features of the transmission line, high-precision, real-time six-dimensional location information detection of transmission line faults is achieved, improving the intelligence and accuracy of transmission line inspection and maintenance.
[0071] In one optional embodiment, based on the current cross-modal geometric constraints of the transmission line, the current cross-modal fusion features, and the initial fault location included in the first fault detection result, a second fault detection result of the transmission line is obtained using a fault location identification module included in the target fault detection model. This includes: acquiring the current depth image and current point cloud data of the transmission line, wherein the current depth image is used to determine the three-dimensional structure and spatial layout of the transmission line; determining the current cross-modal geometric constraints based on the current depth image, the current point cloud data, and the geometric model of the transmission line; and obtaining the second fault detection result based on the current cross-modal geometric constraints, the current cross-modal fusion features, and the initial fault location using a fault location identification module included in the target fault detection model.
[0072] It is understood that acquiring current depth images and point cloud data of the transmission line is used to determine its three-dimensional structure and spatial layout. Based on the current depth images, point cloud data, and the geometric model of the transmission line, the current cross-modal geometric constraints of the transmission line are determined. The fault location identification module included in the target fault detection model is used to identify and analyze the current cross-modal geometric constraints, current cross-modal fusion features, and initial fault location, obtaining the second fault detection result of the transmission line. Through the acquisition of depth image data and point cloud data, the construction of cross-modal geometric constraints, and the precise positioning by the fault location identification module, high-precision, three-dimensional spatial positioning of transmission line faults is achieved, improving the intelligence level and efficiency of transmission line inspection and maintenance.
[0073] Optionally, Figure 2 This is a flowchart of an optional fault detection method for transmission lines provided according to an embodiment of this application, such as... Figure 2The diagram shows the flowchart of the target fault detection model training process. First, an adapter is inserted into the pre-trained model (i.e., the initial fault detection model), and the target parameters of the adapter are initialized using enhanced multimodal data of the transmission line (including image data, infrared thermal image data, and text data). Feature extraction and fusion are performed using the shared feature backbone of the pre-trained model to obtain training cross-modal fusion features. The pre-trained model is then trained using these training cross-modal fusion features. The target parameters of the pre-trained model are updated based on the detection loss and pose loss obtained from the pre-trained model. When the convergence condition is met, i.e., the loss function value is less than or equal to a preset loss function threshold, the target fault detection model is output; otherwise, the target parameters continue to be updated.
[0074] Through the above steps S102 to S108, the cross-modal fusion characteristics of the transmission line can be determined by fusing image data information, infrared thermal image data information, and text data information. Combined with the cross-modal geometric constraints of the transmission line, the fault detection result of the transmission line can be determined using the target fault detection model. This achieves the technical effect of improving the accuracy of the fault detection result of the transmission line, thereby solving the technical problem of low accuracy of fault detection results of transmission lines in related technologies.
[0075] Based on the above embodiments and optional embodiments, this application proposes an optional implementation method for fault detection of transmission lines. This optional implementation method can be understood as a method for fine-tuning a large transmission line inspection model that integrates cross-modal reinforcement learning and multi-cognitive adapters.
[0076] In the field of intelligent power transmission inspection, the accurate perception and decision-making capabilities of the large model (a model that integrates a target fault detection model for fault detection and an acquisition strategy optimization model for determining the acquisition strategy of image acquisition equipment) are the core to achieving unmanned operation. However, existing methods have the following shortcomings: (1) fragmented multimodal features. Traditional convolutional networks only process single-modal visual data, making it difficult to integrate complementary information from inspection images (image data), infrared thermal imaging data, and text data such as equipment ledgers, resulting in a high rate of missed detection for small defects (such as microcracks); (2) insufficient dynamic decision-making ability. Rule-based decision-making models (i.e., the acquisition strategy optimization model in large models) cannot adapt to complex scenarios (such as deformation under different lighting conditions and changes in the perspective of UAVs), resulting in large defect localization errors (average error > 20 cm); (3) low parameter update efficiency. Full fine-tuning of large models (such as ViT-G) requires hundreds of GB of video memory, making it difficult to deploy in real time on edge devices, and the generalization ability to new scenarios is weak (such as a 30% decrease in the detection accuracy of new composite insulators for transmission lines); (4) lack of geometric constraints. Relying solely on 2D (two-dimensional) image pose estimation cannot accurately recover the three-dimensional spatial relationship of transmission lines. In complex occlusion scenarios (such as tree branches occluding conductors), 6D (six-dimensional) constraints are also insufficient. The pose estimation error (ADD-S error > 150 mm) increased significantly.
[0077] To address the aforementioned issues, this application proposes a method for fine-tuning a large-scale transmission line inspection model that integrates cross-modal reinforcement learning and multi-cognitive adapters. The objectives of this method are as follows: (1) Improve cross-modal accuracy perception by aligning image features with text features using CLIP and introducing frequency domain dynamic convolution to enhance multi-scale feature representation, thereby improving the perception of small targets (such as <10px) in transmission lines. 2 (1) Detection accuracy of pin defects (mAP@0.5 improved to 92%); (2) Robust decision optimization: Based on VLM-Gym, a power transmission inspection simulation environment is constructed. The acquisition strategy is optimized by GRPO algorithm to reduce the impact of light changes (such as backlight and fog) on defect identification (error fluctuation <8%); (3) Efficient parameter adaptation: The Mona adapter structure is adopted, and only 3% of the model parameters (i.e., the target parameters of the target fault detection model) are updated to achieve lightweight fine-tuning. The memory usage is reduced by 75%, which meets the requirements of real-time detection of edge devices (frame rate ≥15FPS); (4) Geometric-semantic joint modeling: Combining Neural Implicit Fields (NIF) with the three-dimensional geometric model of the power transmission line, cross-modal geometric constraints are introduced to improve the accuracy of 6D pose estimation (ADD-S error reduced to 45mm).
[0078] To address the challenges of strong electromagnetic interference, concealed equipment defect features, and difficulties in multi-source data fusion in power transmission line inspection scenarios, an end-to-end unified framework integrating Cross-Modal Reinforcement Learning (CRL), a Multi-Cognitive Adapter (Mona Adapter), and CLIP is adopted. Through dynamic interaction of frequency domain features, cross-modal policy optimization, and lightweight parameter adaptation, the contradiction between perception accuracy and real-time decision-making in large-scale models under complex power environments is resolved, exhibiting strong environmental robustness and industrial-grade deployment capabilities. The fine-tuning method for the large-scale power transmission line inspection model integrating CRL and the Multi-Cognitive Adapter includes a cross-modal feature fusion module, a reinforcement learning decision-making module, a Multi-Cognitive Adapter module, and a joint optimization module. These four modules are described in detail below.
[0079] The cross-modal feature fusion module is used to extract and fuse features from multi-source data to obtain training cross-modal fusion features for transmission lines.
[0080] First, multi-source data encoding is performed. The CLIP visual encoder (ViT-B / 16) extracts multi-scale features (32×32×768) and (16×16×1024) from the inspection images (i.e., training image data). Infrared thermal images (i.e., training infrared thermal image data) are processed using a thermal infrared convolution branch (kernel size 5×5) to extract features of abnormal temperature regions (i.e., training target temperature features). For equipment ledger text (i.e., training text data), a semantic vector (dimension 512) is generated using the CLIP text encoder and fused with the visual features through a cross-modal attention fusion mechanism to obtain the training cross-modal fusion feature F. cross F cross The method for determining the value is the same as in the above embodiments, and will not be repeated here.
[0081] Next is frequency domain dynamic convolution (FDConv). An FDConv layer is embedded in the CLIP visual encoder, decomposing the convolution kernel into three frequency bands: low, medium, and high (with preset segmentation thresholds of 0.1 and 0.5 respectively). Frequency domain features are decoupled through Fourier transform. The convolution kernels W for the low, medium, and high frequency bands... s The method for determining s∈{low,mid,high} is the same as in the above embodiments, and will not be repeated here.
[0082] The reinforcement learning decision-making module is used to model the inspection environment of transmission lines and optimize the data acquisition strategy.
[0083] Transmission line inspection environment modeling refers to constructing a VLM-Gym-Transmission environment to simulate transmission line inspection scenarios (such as different climates, lighting conditions, and drone altitudes), and defining the reward function R for the data acquisition strategy optimization model. The method for determining the reward function R is the same as in the above embodiment, and will not be repeated here.
[0084] Optimize the acquisition strategy and model. Use the GRPO algorithm to sample 10 action hypotheses in parallel (e.g., "shift left 10cm", "magnify by 2x"), and then use the advantage function A... s,i,t Normalization is used to screen for high-value actions. Advantage function A s,i,t The method for determining the value is the same as in the above embodiments, and will not be repeated here.
[0085] The multi-cognition adapter module is used to optimize adapter parameters, thereby optimizing the target fault detection model.
[0086] Adapter architecture design. A MonaAdapter is inserted after each Swin Transformer block of the CLIP visual encoder, containing a scale normalization layer and a multi-branch convolutional group. The scale normalization layer adjusts the distribution of the input feature x0 through learnable parameters (i.e., target parameters) s1 and s2 to obtain the adjusted feature x. norm =s1·LayerNorm(x0)+s2·x0; The multi-branch convolutional group extracts multi-scale geometric features through parallel 3×3, 5×5, and 7×7 depth convolutions, and aggregates them through 1×1 convolutions. in, For depthwise convolution kernels, w pw The kernel is a point convolution, and the original features are preserved through skip connections, f dw The output features are the fusion results of multi-scale depthwise convolutional branches, aggregating geometric features of different sizes, where i is the i-th layer and f is the output feature. pw This is the fusion result of the point convolution branches, supplemented with local detail features, where x is the input feature value. For convolution operations, This is a weighted convolution operation.
[0087] Target parameters were efficiently fine-tuned. Only adapter parameters (approximately 2.56% of total parameters) were updated, and the CLIP backbone network was frozen. In transmission line fault detection tasks, FLOPs were reduced from 1820G to 1150G, and the number of parameters was reduced by 28%.
[0088] The joint optimization module is used to optimize the target fault detection model based on the loss function.
[0089] The method for determining the multi-task loss function L is the same as in the above embodiments, and will not be repeated here.
[0090] Cross-modal knowledge distillation. A "defect description-image region" pair (i.e., the third fault detection result) is generated using a teacher model. Through comparative learning, a distillation loss function L is constructed. distill Initialize the target parameters of the adapter in the target fault detection model. distill The method for determining the value is the same as in the above embodiments, and will not be repeated here.
[0091] The pre-trained model and adapter settings for a large-scale power transmission inspection model fine-tuning method integrating cross-modal reinforcement learning and multi-cognitive adapters include updating the target parameters of the model (i.e., the target fault detection model), setting the adapter structure, and initializing the target parameters. The target parameter update method adopts large-scale models such as CLIP and Swin Transformer, freezing the backbone parameters and only activating the multi-cognitive adapter for updates. The adapter structure includes frequency-domain dynamic convolution, scale normalization layers, and multi-branch convolutional groups. The frequency-domain dynamic convolution decomposes the convolution kernel into low, medium, and high frequency bands to capture image contour, texture, and detail features; the scale normalization layer enhances the adaptability of visual signals by adjusting the input feature distribution; and the multi-branch convolutional groups extract multi-scale geometric features through parallel convolutions of different sizes. The target parameters are initialized by generating image-text pairs using a teacher model and initializing the adapter through contrastive learning to reduce cross-modal feature differences.
[0092] A fine-tuning method for a large-scale power transmission line inspection model, integrating cross-modal reinforcement learning and multi-cognitive adapters, can handle data from various modalities, including RGB images, infrared thermal images, and equipment text data. By simulating blurring or edge enhancement, the model's robustness to interference is improved. Simultaneously, by randomizing the pose of the image acquisition equipment, training image data of the transmission line from different perspectives is generated, expanding the diversity of training samples. Finally, rare defect images are generated based on text descriptions, alleviating the difficulty of identifying defects in small sample sizes.
[0093] The reinforcement learning optimization process for fine-tuning a large-scale power transmission line inspection model integrating cross-modal reinforcement learning and multi-cognitive adapters is as follows: A simulation environment for the power transmission line is constructed. Multiple action hypotheses are sampled in parallel, high-value actions are selected through a dominance function, and the adapter parameters are updated using the GRPO algorithm to avoid overfitting.
[0094] A fine-tuning method for a large-scale power transmission line inspection model integrating cross-modal reinforcement learning and multi-cognitive adapters employs a shared feature backbone, dual-task branches, and joint loss optimization to determine the fault type and location of transmission lines. The shared feature backbone involves dividing the feature extraction process into shallow (edge texture), mid-level (medium-sized target), and deep (semantic feature) extraction layers to obtain multi-scale information and enhance feature diversity. The dual-task branches include fault type detection and fault location detection tasks. The fault type detection task utilizes the fault type identification module included in the target fault detection model, generating multi-scale feature maps through a neck network and predicting the initial fault location and fault type using an anchor box mechanism. The fault location detection task utilizes the fault location identification module included in the fault detection model, modeling 3D geometry using neural implicit fields and optimizing pose estimation with synthetic data to improve illumination invariance. Joint loss optimization is performed on the target fault detection model using the first fault detection result obtained from the fault type detection task and the second fault detection result obtained from the fault location detection task.
[0095] The fine-tuning method for a large-scale power transmission inspection model that integrates cross-modal reinforcement learning and multi-cognitive adapters can simultaneously optimize detection loss (target matching), pose loss (geometric constraints), and feature regularization, promoting cross-task information complementarity.
[0096] The fine-tuning method for a large-scale power transmission line inspection model integrating cross-modal reinforcement learning and multi-cognitive adapters employs lightweight deployment and online optimization. Lightweight deployment refers to compressing the model, retaining only adapter parameters (approximately 3% of the total parameters), adapting it to edge devices, reducing memory usage by more than 75%, and achieving an inference frame rate of ≥15 FPS. Online optimization refers to real-time acquisition of new data (including image data, infrared thermal imaging data, and text data) to fine-tune the adapter, adapting to changes in the appearance of equipment in power transmission lines (such as seasonal and lighting effects).
[0097] A fine-tuning method for a large-scale power transmission line inspection model integrating cross-modal reinforcement learning and multi-cognitive adapters is proposed. This method achieves frequency domain decoupling and adaptive fusion of image and text features through CLIP and FDConv, improving detection accuracy for small targets and low-texture defects. A power transmission line inspection simulation environment is constructed based on VLM-Gym, and the multi-modal decision-making strategy is optimized through the GRPO algorithm to enhance adaptability to complex scenes. A Mona Adapter is introduced, achieving efficient parameter fine-tuning through multi-branch convolution and scale normalization, reducing memory usage by more than 75%. Finally, by combining neural implicit fields with a 3D model of the power transmission line, cross-modal geometric loss is introduced to improve the illumination invariance and robustness of 6D pose estimation.
[0098] The above-mentioned optional implementation methods achieve at least the following effects: by collecting image data, infrared thermal imaging data, and text data of transmission lines, it is possible to ensure the coverage of multi-source information, and to comprehensively evaluate the status of transmission lines from multiple perspectives such as vision, temperature, and historical records, thereby improving the accuracy and reliability of fault detection results; by optimizing the acquisition strategy model and the current environmental data of the transmission lines, the optimal acquisition strategy of the image acquisition device is dynamically generated, ensuring that high-quality image data can be captured even in complex environments, providing high-quality image data for subsequent fault detection; by optimizing the adapter parameters, the target fault detection model can better integrate image data, infrared thermal imaging data, and text data information, enhancing the understanding and recognition capabilities of complex transmission line inspection scenarios.
[0099] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0100] This embodiment also provides a fault detection device for transmission lines, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the terms "module" and "device" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.
[0101] According to an embodiment of this application, an apparatus embodiment for implementing a fault detection method for transmission lines is also provided. Figure 3 This is a schematic diagram of a fault detection device for a power transmission line according to an embodiment of this application, as shown below. Figure 3 As shown, the fault detection device for the above-mentioned transmission line includes a data acquisition module 302, a current cross-modal fusion feature determination module 304, a first fault detection result determination module 306, and a second fault detection result determination module 308. The device will be described below.
[0102] The data acquisition module 302 is used to acquire current image data, current infrared thermal image data, and current text data of the transmission line, wherein the current infrared thermal image data is used to describe the current temperature distribution of the transmission line.
[0103] The current cross-modal fusion feature determination module 304 is connected to the data acquisition module 302 and is used to determine the current cross-modal fusion features of the transmission line based on the current image data, the current infrared thermal image data, and the current text data. The current cross-modal fusion features are used to fuse the information of the current image data, the current infrared thermal image data, and the current text data.
[0104] The first fault detection result determination module 306 is connected to the current cross-modal fusion feature determination module 304. It is used to obtain the first fault detection result of the transmission line based on the current cross-modal fusion features and the fault type identification module included in the target fault detection model. The first fault detection result includes at least: fault type and initial fault location. The initial fault location is used to describe the two-dimensional location information of the fault in the transmission line. The target fault detection model is used to determine the fault type and fault location of the transmission line.
[0105] The second fault detection result determination module 308, connected to the first fault detection result determination module 306, is used to obtain the second fault detection result of the transmission line based on the current cross-modal geometric constraints of the transmission line, the current cross-modal fusion features, and the initial fault location included in the first fault detection result, using the fault location identification module included in the target fault detection model. The second fault detection result is used to describe the six-dimensional location information of the transmission line fault, which includes three-dimensional location information and three-dimensional pose information. The current cross-modal geometric constraints are used to ensure the consistency of the current image data, the current infrared thermal image data, and the current text data in the spatial coordinate system.
[0106] This application provides a fault detection device for transmission lines. By setting up a data acquisition module 302, a current cross-modal fusion feature determination module 304, a first fault detection result determination module 306, and a second fault detection result determination module 308, the device achieves the goal of determining the cross-modal fusion features of the transmission line by fusing image data information, infrared thermal image data information, and text data information. Combined with the cross-modal geometric constraints of the transmission line, and using a target fault detection model, the device determines the fault detection result of the transmission line, thereby improving the accuracy of the fault detection result of the transmission line and solving the technical problem of low accuracy of fault detection results for transmission lines in related technologies.
[0107] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0108] It should be noted that the data acquisition module 302, the current cross-modal fusion feature determination module 304, the first fault detection result determination module 306, and the second fault detection result determination module 308 mentioned above correspond to steps S102 to S108 in the embodiments. The instances and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in a computer terminal.
[0109] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the embodiments, and will not be repeated here.
[0110] The fault detection device for the aforementioned transmission line may also include a processor and a memory. The data acquisition module 302, the current cross-modal fusion feature determination module 304, the first fault detection result determination module 306, and the second fault detection result determination module 308 are all stored as program units in the memory. The processor executes the aforementioned program units stored in the memory to realize the corresponding functions.
[0111] The processor contains a core that retrieves the corresponding program unit from memory. One or more cores may be configured. Memory may include non-persistent memory in computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.
[0112] This application provides a non-volatile storage medium storing a program that, when executed by a processor, implements a fault detection method for power transmission lines.
[0113] This application provides an electronic device including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: acquiring current image data, current infrared thermal image data, and current text data of a transmission line, wherein the current infrared thermal image data is used to describe the current temperature distribution of the transmission line; determining current cross-modal fusion features of the transmission line based on the current image data, current infrared thermal image data, and current text data, wherein the current cross-modal fusion features are used to fuse information from the current image data, current infrared thermal image data, and current text data; and obtaining a first fault detection result for the transmission line based on the current cross-modal fusion features using a fault type identification module included in a target fault detection model. The first fault detection result includes at least the fault type and initial fault location. The initial fault location describes the two-dimensional location information of the transmission line fault. The target fault detection model is used to determine the fault type and fault location of the transmission line. Based on the current cross-modal geometric constraints of the transmission line, the current cross-modal fusion features, and the initial fault location included in the first fault detection result, the fault location identification module included in the target fault detection model is used to obtain the second fault detection result of the transmission line. The second fault detection result describes the six-dimensional location information of the transmission line fault, including three-dimensional location information and three-dimensional pose information. The current cross-modal geometric constraints are used to ensure the consistency of the current image data, current infrared thermal image data, and current text data in the spatial coordinate system. The device in this paper can be a server, PC, etc.
[0114] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program that initializes with the following method steps: acquiring current image data, current infrared thermal image data, and current text data of a transmission line, wherein the current infrared thermal image data is used to describe the current temperature distribution of the transmission line; determining current cross-modal fusion features of the transmission line based on the current image data, current infrared thermal image data, and current text data, wherein the current cross-modal fusion features are used to fuse information from the current image data, current infrared thermal image data, and current text data; and obtaining a first fault detection result of the transmission line by using a fault type identification module included in a target fault detection model based on the current cross-modal fusion features, wherein the first fault detection... The results include at least: fault type and initial fault location, where the initial fault location describes the two-dimensional location information of the transmission line fault, and the target fault detection model is used to determine the fault type and fault location of the transmission line; based on the current cross-modal geometric constraints of the transmission line, the current cross-modal fusion features, and the initial fault location included in the first fault detection result, the fault location identification module included in the target fault detection model is used to obtain the second fault detection result of the transmission line, where the second fault detection result describes the six-dimensional location information of the transmission line fault, which includes three-dimensional location information and three-dimensional pose information, and the current cross-modal geometric constraints are used to ensure the consistency of the current image data, current infrared thermal image data, and current text data in the spatial coordinate system.
[0115] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0116] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0117] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0118] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0119] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0120] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0121] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0122] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0123] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0124] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method of fault detection for a power transmission line, characterized by, The method comprises: acquiring current image data, current infrared thermal image data, and current text data of a power transmission line, wherein the current infrared thermal image data is used to describe a current temperature distribution of the power transmission line; based on the current image data, the current infrared thermal image data, and the current text data, determining a current cross-modal fusion feature of the power transmission line, wherein the current cross-modal fusion feature is used to fuse information of the current image data, the current infrared thermal image data, and the current text data; based on the current cross-modal fusion feature, using a fault type identification module included in a target fault detection model to obtain a first fault detection result of the power transmission line, wherein the first fault detection result at least includes a fault type and an initial fault position, the initial fault position is used to describe two-dimensional position information of the fault of the power transmission line, and the target fault detection model is used to determine the fault type and the fault position of the power transmission line; based on the current cross-modal geometric constraint of the power transmission line, the current cross-modal fusion feature, and the initial fault position included in the first fault detection result, using a fault position identification module included in the target fault detection model to obtain a second fault detection result of the power transmission line, wherein the second fault detection result is used to describe six-dimensional position information of the fault of the power transmission line, the six-dimensional position information includes three-dimensional position information and three-dimensional pose information, and the current cross-modal geometric constraint is used to ensure consistency of the current image data, the current infrared thermal image data, and the current text data in a spatial coordinate system.
2. The method of claim 1, wherein, The method comprises: based on the current environment data of the power transmission line, using a collection strategy optimization model to determine a collection strategy of an image collection device, wherein the collection strategy at least includes a collection path and a collection pose of the image collection device, and the collection strategy optimization model pre-learns an association relationship between the environment data of the power transmission line and the collection strategy; according to the collection strategy, collecting the current image data.
3. The method of claim 1, wherein, Before the method based on the current cross-modal fusion feature, using a fault type identification module included in a target fault detection model to obtain a first fault detection result of the power transmission line, the method further comprises: obtaining training image data, training infrared thermal image data, and training text data of the power transmission line; based on the training image data, the training infrared thermal image data, and the training text data, obtaining training cross-modal fusion features of the power transmission line; using the training cross-modal fusion features and training cross-modal geometric constraints of the power transmission line to train an initial fault detection model to obtain the target fault detection model.
4. The method of claim 3, wherein, The method based on the training image data, the training infrared thermal image data, and the training text data to obtain the training cross-modal fusion features of the power transmission line comprises: performing feature extraction on the training image data to obtain training target image features of the power transmission line; feature extraction is performed on the training infrared thermal image data to obtain training target temperature features of the power transmission line; feature extraction is performed on the training text data to obtain training target text features of the power transmission line; based on the training target image features, the training target temperature features, and the training target text features, the training cross-modal fusion features are determined.
5. The method of claim 4, wherein, The feature extraction on the training image data to obtain the training target image features of the power transmission line includes: feature extraction is performed on the training image data to obtain training initial image features of the power transmission line; Fourier transform is performed on the training initial image features to obtain frequency domain image features of the power transmission line; based on a preset frequency band segmentation threshold, the frequency domain image features are segmented to obtain frequency domain image sub-features corresponding to the plurality of frequency bands respectively; based on the frequency domain image sub-features corresponding to the plurality of frequency bands respectively, inverse Fourier transform is performed to obtain time domain image sub-features corresponding to the plurality of frequency bands respectively; based on the time domain image sub-features corresponding to the plurality of frequency bands respectively, the training target image features of the power transmission line are determined.
6. The method of claim 3, wherein, The training of the initial fault detection model based on the training cross-modal fusion features and the training cross-modal geometric constraints of the power transmission line to obtain the target fault detection model includes: based on the detection loss and the pose loss of the power transmission line, a loss function of the initial fault detection model is determined, wherein the detection loss is used to quantify the error of the first detection result, and the pose loss is used to quantify the error of the second detection result; based on the function value corresponding to the loss function and a preset loss function threshold, the initial parameter value of the target parameter of the initial fault detection model is adjusted, wherein the target parameter is a parameter of an adapter included in the initial fault detection model, and the adapter is used to fuse image data, infrared thermal image data and text data; when the function value corresponding to the loss function is less than or equal to the preset loss function threshold, the parameter value of the target parameter is determined as a target parameter value; based on the target parameter value, the target fault detection model is determined.
7. The method of claim 6, wherein, Before the initial parameter value of the target parameter of the initial fault detection model is adjusted based on the function value corresponding to the loss function and the preset loss function threshold, the method further includes: the training cross-modal fusion features are input into the initial fault detection model to obtain a third fault detection result of the power transmission line; the training cross-modal fusion features are respectively input into a teacher model to obtain a fourth fault detection result of the power transmission line, wherein the teacher model is used to guide the learning process of the initial fault detection model; based on the third fault detection result and the fourth fault detection result, detection errors of the initial fault detection model and the teacher model are determined; based on the detection error and a preset error threshold, the initial parameter value is determined.
8. The method according to any one of claims 1 to 7, characterized in that, The current cross-modal geometric constraint of the power transmission line, the current cross-modal fusion feature, and the initial fault location included in the first fault detection result are used to obtain a second fault detection result of the power transmission line by using a fault location identification module included in the target fault detection model, including: Obtain the current depth image and the current point cloud data of the power transmission line, wherein the current depth image is used to determine the three-dimensional structure and spatial layout of the power transmission line; Determine the current cross-modal geometric constraint based on the current depth image, the current point cloud data, and the geometric model of the power transmission line; Based on the current cross-modal geometric constraint, the current cross-modal fusion feature, and the initial fault location, a fault location identification module included in the target fault detection model is used to obtain a second fault detection result.
9. A fault detection device for a power transmission line, characterized in that Comprising: The data acquisition module is used to acquire the current image data, the current infrared thermal image data, and the current text data of the power transmission line, wherein the current infrared thermal image data is used to describe the current temperature distribution of the power transmission line; The current cross-modal fusion feature determination module is used to determine the current cross-modal fusion feature of the power transmission line based on the current image data, the current infrared thermal image data, and the current text data, wherein the current cross-modal fusion feature is used to fuse the information of the current image data, the current infrared thermal image data, and the current text data; The first fault detection result determination module is used to obtain a first fault detection result of the power transmission line by using a fault type identification module included in the target fault detection model based on the current cross-modal fusion feature, wherein the first fault detection result at least includes: fault type and initial fault location, the initial fault location is used to describe the two-dimensional position information of the power transmission line fault, and the target fault detection model is used to determine the fault type and fault location of the power transmission line; The second fault detection result determination module is used to obtain a second fault detection result of the power transmission line by using a fault location identification module included in the target fault detection model based on the current cross-modal geometric constraint of the power transmission line, the current cross-modal fusion feature, and the initial fault location included in the first fault detection result, wherein the second fault detection result is used to describe the six-dimensional position information of the power transmission line fault, the six-dimensional position information includes three-dimensional position information and three-dimensional pose information, and the current cross-modal geometric constraint is used to ensure the consistency of the current image data, the current infrared thermal image data, and the current text data in the spatial coordinate system.
10. A non-volatile storage medium, comprising: The non-volatile storage medium stores a plurality of instructions, the instructions are suitable for being loaded and executed by the processor to perform the fault detection method of the power transmission line in any one of claims 1 to 8.
Citation Information
Cited By
Power transmission line defect detection method and device and electronic equipment
CN121582258A