An intelligent vision and state perception collaborative hidden danger monitoring method and system for faults in dense-channel power transmission lines

By building a state-aware collaborative hidden danger monitoring architecture in dense channel transmission lines, combining deep learning and Internet of Things technology, using drone monocular cameras for image acquisition and hierarchical detection, the problems of untimely identification of faults and inaccurate positioning of dense channel transmission lines are solved, and efficient and accurate fault monitoring and early warning are achieved.

CN120259929BActive Publication Date: 2025-08-05SICHUAN YAAN ELECTRIC POWER (GRP) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510740416.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-08-05
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The fault identification of dense channel transmission lines is not timely and the positioning is inaccurate. The existing visual detection models are low in detection rate, high error detection rate, and poor generalization performance in complex open environments, making it difficult to meet the requirements of power grid safety and reliability.

Method used

Build a state-aware collaborative potential hazard monitoring architecture, combine deep learning algorithms and Internet of Things technology, and build multi-dimensional perceptual potential hazard coordination models and self-learning mechanisms at the edge and cloud respectively. Use drone monocular cameras to collect images, conduct first-level detection at the edge, and hand over suspected images to the cloud for secondary detection, generating hierarchical early warning signals.

Benefits of technology

It improves the accuracy of identification and positioning of faults in dense channel transmission lines, simplifies operations, is suitable for promotion and application in more scenarios, and improves data transmission efficiency and detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259929B_ABST
    Figure CN120259929B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of power transmission line inspection and fault identification. It provides a method and system for collaborative hidden danger monitoring of dense channel transmission line faults using intelligent vision and state perception. It constructs a state perception collaborative hidden danger monitoring architecture, constructs a multi-dimensional hidden danger perception collaborative model and a self-learning hidden danger perception collaborative model on the edge and cloud, respectively. It uses a monocular camera mounted on a drone to capture images of dense channel transmission lines. It performs primary detection at the edge, and when a suspicious image is found, performs secondary detection in the cloud, outputs the target detection results, and generates graded warning signals. This can improve the efficiency of data transmission and enhance the precision and accuracy of intelligent vision and state perception collaborative hidden danger monitoring of dense channel transmission line faults.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of transmission line inspection and fault identification, and in particular to a method and system for collaborative hidden danger monitoring of dense channel transmission line faults using intelligent vision and state perception. Background Art

[0002] Dense transmission lines are the primary model for power grid construction both domestically and internationally. By 2024, my country's dense transmission lines will exceed 1 million kilometers. This model offers advantages such as high transmission capacity and efficient resource utilization, but it also faces challenges such as delayed fault identification and inaccurate location, which can easily lead to regional blackouts.

[0003] Multiple research reports indicate that several major power outages in China and abroad were caused by inaccurate fault identification and location in densely populated transmission lines. The challenges lie in: ① The complex and highly noisy operating environment, coupled with diverse manifestations of hidden dangers, make it difficult for a single monitoring technology to effectively detect them. ② The multitude of fault triggers, coupled with a small number of valid samples, makes feature extraction challenging, and limited edge computing power makes it difficult to identify small defects. ③ The complex topology and fault mechanisms, coupled with deep feature coupling, make accurate location difficult with traditional methods. These characteristics make visual inspection extremely challenging. Existing visual inspection models encounter technical bottlenecks in complex, open environments, including low detection rates for densely populated small objects, high false positive rates, and poor model generalization. This not only requires grid operators to reanalyze and review results, but also significantly limits the further application of AI technology in production sites. With the expansion of power grids and the increasing density of transmission lines, traditional manual inspections and single-sensor monitoring methods are no longer able to meet the safety and reliability requirements of modern power systems.

[0004] Therefore, designing an intelligent vision and state perception collaborative hidden danger monitoring method for dense channel transmission line faults and conducting timely and accurate detection and defect troubleshooting of transmission lines are important means and urgent tasks to ensure the safe and stable operation of the power grid. Summary of the Invention

[0005] In view of this, it is necessary to provide an intelligent vision and state perception collaborative hidden danger monitoring method for dense channel transmission line faults. By combining deep learning algorithms with Internet of Things technology, real-time monitoring, fault warning and precise positioning of hidden dangers of dense channel transmission lines can be achieved, thereby improving the accuracy of fault identification and positioning of dense channel transmission lines. It is easy to promote and apply, and the operation is relatively simple, making it suitable for promotion and application in more scenarios.

[0006] In a first aspect, an embodiment of the present application provides a method for monitoring hidden dangers of dense channel transmission line faults using intelligent vision and state perception collaboration, the method comprising:

[0007] S1: Build a state-aware collaborative hidden danger monitoring architecture, which includes edge devices, transmission channels, intermediate nodes, and the cloud.

[0008] S2: Constructing a multi-dimensional hidden danger perception collaborative model based on the fusion of residual network and dynamic convolutional network model on the edge, and constructing a multi-dimensional hidden danger perception collaborative model with a self-learning mechanism on the cloud;

[0009] S3: Use a monocular camera on a drone to collect images of dense channel transmission lines;

[0010] S4: The edge end detects the image to be detected based on the multi-dimensional hidden danger perception collaborative model. If the detection result is a suspected image, the drone is commanded to return to the target according to the set flight strategy and re-collect image data according to the predetermined hovering strategy. At the same time, the detection result of the suspected image and the re-collected image data are sent to the cloud, and the detection task is handed over to the cloud, and the edge end stops detection.

[0011] S5: Re-detecting the suspected image in the cloud based on the multi-dimensional hidden danger perception collaborative model of the self-learning mechanism, and outputting a target detection result;

[0012] S6: Generate graded warning signals based on the dynamic risk assessment model to achieve accurate warning of hidden danger location errors.

[0013] Optionally, in an implementation of the first aspect of the present invention, the S1: constructing a state-aware collaborative hidden danger monitoring architecture, the state-aware collaborative hidden danger monitoring architecture including an edge terminal, a transmission channel, an intermediate node, and a cloud, includes:

[0014] The state-aware collaborative hidden danger monitoring architecture consists of five layers: terminal layer, first communication layer, edge layer, second communication layer, and cloud layer.

[0015] Wherein, the terminal layer is used to collect data by receiving instructions;

[0016] The first communication layer is used for data transmission between the terminal layer and the edge layer;

[0017] The edge layer is used to receive transmission data, perform data processing and detection, and send control instructions based on the detection results;

[0018] The second communication layer is used for data transmission between the edge layer and the cloud layer;

[0019] The cloud layer is used to receive and transmit data, perform deeper data processing and detection, and perform data management, data storage, and data visualization.

[0020] Optionally, in an implementation of the first aspect of the present invention, S2: constructing a multi-dimensional hidden danger perception collaborative model based on the fusion of a residual network and a dynamic convolutional network model on the edge terminal, including:

[0021] The multi-dimensional perception collaborative model includes an encoder and a decoder;

[0022] The encoder is composed of a ResNet-CNN architecture, including a residual network ResNet module, a multi-scale feature fusion module, and a first image feature self-attention module.

[0023] The residual network ResNet module is used as the backbone network and consists of two modules. The first module includes an input layer, a convolution layer, a batch normalization layer, a ReLU activation function layer, and a maximum pooling layer. The input layer has a size of 3 channels. The RGB image is taken as input; the second module contains 7 sequential networks with the same structure, each of which consists of multiple bottleneck residual blocks, each of which contains a convolutional layer, a batch normalization layer and a ReLU activation layer, and the output feature size is ;

[0024] The multi-scale feature fusion module contains five different processing channels, each with a different number of layers. The input feature size is 2048 and contains two convolutional layers to generate feature maps for different channels. The feature maps are resized by bilinear interpolation to match the spatial dimensions of the backbone features. The resized feature maps are concatenated with the backbone features to capture both low-level and high-level features at different spatial resolutions. The total number of output channels of the concatenation module is 2544×H×W, which are then fed into the self-attention module for processing.

[0025] The first image feature self-attention module consists of parallel query convolution, key convolution, and value convolution. The input is the features of the connection module, which are converted into query, key, and value tensors through the convolution operations of query convolution, key convolution, and value convolution respectively. The attention weight is calculated by the similarity between the query and key tensors, and then weighted focusing is performed with the value tensor to selectively focus on relevant image areas and enhance the representation of important features.

[0026] The decoder includes an embedding layer, a GRU recurrent neural network, and a linear output layer, wherein the embedding layer is used to receive input text descriptions, convert the text into an embedded vector representation, and splice and fuse it with the image feature tensor output by the encoder; the GRU recurrent neural network is used to process sequence data using a gating mechanism, apply Dropout regularization before inputting the linear layer, capture temporal context information through hidden states, and ultimately output a predicted tag distribution; the linear output layer is used to map the GRU hidden states to the vocabulary space to generate a probability distribution of the description text; this architecture achieves end-to-end image description generation by jointly optimizing the visual-linguistic feature representation.

[0027] Optionally, in an implementation of the first aspect of the present invention, the multi-dimensional hidden danger perception collaborative model with a self-learning mechanism is constructed on the cloud, including:

[0028] Build a self-learning mechanism based on the parallel three-dimensional ResNet, Faster-RCNN, and Vgg+Net networks;

[0029] By continuously collecting new image data and detection results at the edge, mixed samples are obtained;

[0030] Performing data augmentation on the mixed samples and screening training samples through an adversarial network;

[0031] Use three three-dimensional ResNet, Faster-RCNN, and Vgg+Net networks to predict the training samples and output their respective prediction results;

[0032] The samples of the respective prediction results are divided into obvious samples and not obvious samples according to the accuracy of the prediction, and self-learning is performed to calculate the loss weight corresponding to each sample. The training loss is optimized by weighted optimization to enable the network to better learn obvious samples and not obvious samples. The multi-dimensional perception collaborative model of hidden dangers is continuously updated and optimized through the self-learning mechanism to improve the detection performance and adaptability of the model.

[0033] Optionally, in an implementation of the first aspect of the present invention, after S3: using a monocular camera mounted on a drone to capture images of dense channel transmission lines, the method further includes:

[0034] The input image is preprocessed and features are extracted at key point positions. The coordinates of the features are ;

[0035] According to the depth value , convert the coordinates of the features into three-dimensional coordinates ,in,

[0036] ;

[0037] is the horizontal set distance of the monocular camera, is the baseline between the structured light projector and the monocular infrared camera, .

[0038] Optionally, in an implementation of the first aspect of the present invention, commanding the drone to return to the target according to the set flight strategy and to collect image data again according to a predetermined hovering strategy includes:

[0039] Divide the entire flight area into multiple sub-areas;

[0040] Selecting an optimal hovering point in each sub-area, wherein the optimal hovering point is the one that maximizes the quality of information collection;

[0041] Combining the optimal hovering points into a whole flight path;

[0042] Combining particle swarm optimization and Bayesian optimization algorithms, it generates trajectories that balance obstacle avoidance and efficiency in real time;

[0043] Re-collect image data from multiple angles through drone control commands;

[0044] The method combines particle swarm optimization and Bayesian optimization algorithms to generate trajectories that balance obstacle avoidance and efficiency in real time. This includes: using Bayesian optimization to establish a global probability model to handle environmental uncertainty and obstacle avoidance constraints, and updating the model in real time to reflect the current obstacle status; PSO performs local path optimization on this basis, and quickly adjusts the particle swarm to adapt to dynamic changes;

[0045] Obstacle positions, speeds, and predicted trajectories are acquired in real time through sensors, and the Bayesian optimization constraint model is updated. PSO is used to generate candidate paths, evaluate fitness, and perform iterative optimization based on the fitness. The fitness function is:

[0046] ;

[0047] in, is the path length, is the obstacle avoidance penalty calculated by the distance field, is the smoothness of the path calculated by curvature or acceleration, is the dynamic constraint, are the weight factors of the corresponding parameters respectively.

[0048] Optionally, in an implementation of the first aspect of the present invention, generating a graded warning signal based on a dynamic risk assessment model to achieve accurate warning of hidden danger location errors includes:

[0049] Integrate multi-dimensional data to build a risk assessment model, including target detection results and electrical, mechanical, and environmental data collected by multiple terminals;

[0050] Perform risk assessment using the risk assessment model to obtain a risk value;

[0051] Generate graded warning signals based on risk values to achieve accurate warning of hidden danger location errors;

[0052] The risk value of the risk assessment model is updated through incremental data, and automatically adjusted according to new data to maintain the timeliness of the assessment.

[0053] In a second aspect, an embodiment of the present application provides an intelligent visual and state perception collaborative hidden danger monitoring system for dense channel transmission line faults, which is applied to the intelligent visual and state perception collaborative hidden danger monitoring method for dense channel transmission line faults as described in the first aspect, and is characterized by comprising:

[0054] Monitoring architecture building module: Build a state-aware collaborative hidden danger monitoring architecture, which includes edge devices, transmission channels, intermediate nodes, and the cloud;

[0055] Model construction module: constructing a multi-dimensional hidden danger perception collaborative model based on the fusion of residual network and dynamic convolutional network model on the edge end, and constructing a multi-dimensional hidden danger perception collaborative model with a self-learning mechanism on the cloud end;

[0056] Image acquisition module: uses the monocular camera on the drone to collect images of dense channel transmission lines;

[0057] The first detection module detects the image to be detected at the edge based on the multi-dimensional hidden danger perception collaborative model. If the detection result is a suspected image, the drone is commanded to return to the target according to the set flight strategy and re-collect image data according to the predetermined hovering strategy. At the same time, the detection result of the suspected image and the re-collected image data are sent to the cloud, and the detection task is handed over to the cloud, and the edge stops detection.

[0058] The second detection module re-detects the suspected image in the cloud based on the multi-dimensional hidden danger perception collaborative model of the self-learning mechanism and outputs the target detection result;

[0059] Grading warning module: Generates graded warning signals based on the dynamic risk assessment model to achieve accurate warning of hidden danger location errors.

[0060] In a third aspect, an embodiment of the present application provides an electronic device, characterized by including:

[0061] processor;

[0062] a memory for storing processor-executable instructions;

[0063] Wherein, the processor is configured to implement the intelligent vision and state perception collaborative hidden danger monitoring method for dense channel transmission line faults as described in the first aspect when executing the instructions.

[0064] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a program, and the program instructs a device to execute the intelligent vision and state perception collaborative hidden danger monitoring method for dense channel transmission line faults as described in the first aspect.

[0065] This application provides a method and system for collaborative hidden danger monitoring of dense channel transmission line faults using intelligent vision and state perception. The system constructs a state perception collaborative hidden danger monitoring architecture, building a multi-dimensional hidden danger perception collaborative model and a self-learning hidden danger perception collaborative model on the edge and cloud, respectively. The system uses a monocular camera mounted on an unmanned aerial vehicle (UAV) to capture images of dense channel transmission lines. Primary detection is performed at the edge, and secondary detection is performed in the cloud when suspicious images are detected, outputting target detection results. A graded warning signal is generated. This improves data transmission efficiency and enhances the precision and accuracy of collaborative hidden danger monitoring using intelligent vision and state perception for dense channel transmission line faults.

[0066] Beneficial effects:

[0067] (1) Using a monocular camera to capture images of dense channel transmission lines can improve the efficiency of data transmission.

[0068] (2) By using a multi-dimensional hidden danger perception collaborative model that integrates an improved residual network and a dynamic convolutional network model, the evaluation results are made more accurate.

[0069] (3) By performing the first-level detection at the edge and performing the second-level detection in the cloud when there is a suspicious image, the target detection results are more accurate;

[0070] (4) It is easy to promote and apply, and the operation is relatively simple, making it suitable for promotion and application in more application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 A flowchart of a method for collaborative hidden danger monitoring of dense channel transmission line faults using intelligent vision and state perception is provided in one embodiment of the present application.

[0072] Figure 2 This is an overall structural diagram of the multi-dimensional perception collaboration model provided in one embodiment of the present application.

[0073] Figure 3 This is a diagram of the encoder-decoder network architecture provided in one embodiment of the present application.

[0074] Figure 4 This is a diagram of the encoder architecture provided in one embodiment of the present application.

[0075] Figure 5 This is a diagram of the residual network ResNet module architecture provided in one embodiment of the present application.

[0076] Figure 6 This is a diagram of the multi-scale feature fusion module architecture provided in one embodiment of the present application.

[0077] Figure 7 Schematic diagram of the first image feature self-attention module provided in one embodiment of the present application.

[0078] Figure 8 A schematic diagram of a decoder provided in one embodiment of the present application.

[0079] Figure 9 Schematic diagram of a multi-dimensional collaborative model for hidden danger perception of a self-learning mechanism provided in one embodiment of the present application.

[0080] Figure 10 A schematic diagram of the module of the intelligent vision and state perception collaborative hidden danger monitoring system for dense channel transmission line faults provided in one embodiment of the present application.

[0081] Figure 11 A schematic diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0082] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0083] It should be noted that, in the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art in the art to which this application relates. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0084] It should be noted that, in the embodiments of the present application, words such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order. Features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way.

[0085] Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0086] Example 1

[0087] This application provides a method and system for collaborative hidden danger monitoring of dense channel transmission line faults using intelligent vision and state perception. The system constructs a state perception collaborative hidden danger monitoring architecture, building a multi-dimensional hidden danger perception collaborative model and a self-learning hidden danger perception collaborative model on the edge and cloud, respectively. The system uses a monocular camera mounted on an unmanned aerial vehicle (UAV) to capture images of dense channel transmission lines. Primary detection is performed at the edge, and secondary detection is performed in the cloud when suspicious images are detected, outputting target detection results. A graded warning signal is generated. This improves data transmission efficiency and enhances the precision and accuracy of collaborative hidden danger monitoring using intelligent vision and state perception for dense channel transmission line faults.

[0088] Figure 1 A flowchart of a method for collaborative hidden danger monitoring of dense channel transmission line faults using intelligent vision and state perception is provided in one embodiment of the present application.

[0089] like Figure 1 As shown, a method for monitoring hidden dangers of dense channel transmission line faults by intelligent vision and state perception collaboration includes:

[0090] S1: Build a state-aware collaborative hidden danger monitoring architecture, which includes edge terminals, transmission channels, intermediate nodes and the cloud.

[0091] It is understood that, in this embodiment, the S1: constructing a state-aware collaborative hidden danger monitoring architecture includes an edge terminal, a transmission channel, an intermediate node, and a cloud, including:

[0092] The state-aware collaborative hidden danger monitoring architecture consists of five layers: terminal layer, first communication layer, edge layer, second communication layer, and cloud layer.

[0093] Wherein, the terminal layer is used to collect data by receiving instructions;

[0094] The first communication layer is used for data transmission between the terminal layer and the edge layer;

[0095] The edge layer is used to receive transmission data, perform data processing and detection, and send control instructions based on the detection results;

[0096] The second communication layer is used for data transmission between the edge layer and the cloud layer;

[0097] The cloud layer is used to receive and transmit data, perform deeper data processing and detection, and perform data management, data storage, and data visualization.

[0098] Specifically, the edge computing terminal and the cloud computing terminal are the core of the architecture. The edge computing terminal is responsible for image acquisition, hidden danger target detection and data transmission, while the cloud is mainly responsible for collecting and organizing the environmental images and hidden danger detection results corresponding to each edge layer node server; the edge layer server cluster and communication channel serve as storage and transmission modules connecting the edge and cloud platforms.

[0099] The edge is responsible for data processing and forwarding, providing services such as intelligent perception and data analysis. The edge layer includes devices such as edge gateways, controllers, the cloud, and sensors. The edge device layer is responsible for data collection, the edge node layer processes data, and the cloud layer performs global analysis. The edge computing end processes data and provides feedback, and the cloud trains models and deploys them to the edge, supporting real-time services. After preliminary processing, data is uploaded to an edge server cluster or the cloud via communication channels (such as edge gateways). The edge also collects data and uploads results to the cloud. The cloud relies on high-performance server clusters for large-scale data storage, in-depth analysis, and complex model training.

[0100] S2: Construct a multi-dimensional hidden danger perception collaborative model based on the fusion of residual network and dynamic convolutional network model on the edge end, and construct a multi-dimensional hidden danger perception collaborative model with a self-learning mechanism on the cloud end.

[0101] Figure 2 This is the overall structure diagram of the multi-dimensional perception collaboration model provided by one embodiment of the present application. Figure 2 As shown, it can be understood that in this embodiment, the S2: constructs a hidden danger multi-dimensional perception collaborative model based on the fusion of residual network and dynamic convolutional network model on the edge end, including: the multi-dimensional perception collaborative model includes an encoder and a decoder.

[0102] Figure 3 This is a diagram of the encoder-decoder network architecture provided in one embodiment of the present application. Figure 4 This is a diagram of the encoder architecture provided in one embodiment of the present application. Figure 3 、 Figure 4 As shown in the figure, the encoder is composed of a ResNet fusion CNN architecture, including: a residual network ResNet module, a multi-scale feature fusion module, and a first image feature self-attention module.

[0103] The residual network ResNet module is used as the backbone network and consists of two modules. The first module includes an input layer, a convolution layer, a batch normalization layer, a ReLU activation function layer, and a maximum pooling layer. The input layer has a size of 3 channels. The RGB image is taken as input; the second module contains 7 sequential networks with the same structure, each of which consists of multiple bottleneck residual blocks, each of which contains a convolutional layer, a batch normalization layer and a ReLU activation layer, and the output feature size is .

[0104] The multi-scale feature fusion module contains five different processing channels, each with a different number of layers, an input feature size of 2048, and two convolutional layers for generating feature maps of different channels. The feature maps are resized by bilinear interpolation to match the spatial dimensions of the backbone features. The adjusted feature maps are connected to the backbone features to capture both low-level and high-level features at different spatial resolutions. The total number of output channels of the connection module is 2544×H×W, which are then input into the self-attention module for processing.

[0105] The first image feature self-attention module consists of parallel query convolution, key convolution, and value convolution. The input is the feature of the connection module, which is converted into query, key, and value tensors through the convolution operations of query convolution, key convolution, and value convolution respectively; by calculating the attention weight based on the similarity between the query and key tensors, and then performing weighted focusing with the value tensor, it selectively focuses on the relevant image areas and enhances the representation of important features.

[0106] Specifically, the decoder includes an embedding layer, a GRU recurrent neural network, and a linear output layer, wherein the embedding layer is used to receive input text descriptions, convert the text into an embedded vector representation, and perform splicing and fusion with the image feature tensor output by the encoder; the GRU recurrent neural network is used to process sequence data using a gating mechanism, apply Dropout regularization before inputting the linear layer, capture temporal context information through hidden states, and finally output a predicted tag distribution; the linear output layer is used to map the GRU hidden states to the vocabulary space to generate a probability distribution of the description text; this architecture achieves end-to-end image description generation by jointly optimizing the visual-linguistic feature representation.

[0107] Figure 5 This is a diagram of the residual network ResNet module architecture provided in one embodiment of the present application. Specifically, Figure 5 As shown in the figure, the residual network ResNet module is used as the backbone module to process images through multiple convolutional layers and pooling layers to extract hierarchical features of different scales. It includes 8 network modules. Each network module is fine-tuned by incorporating the parameters of the layer into the gradient calculation during training. The first network module includes a convolutional layer, a batch normalization layer, a ReLU activation function layer, and a maximum pooling layer. The network module is based on a 3-channel size. The RGB image is taken as input and processed by the ResNet backbone network. The backbone network contains the 1st to 7th sequential network modules. Each sequential network module consists of multiple bottleneck residual blocks. Each bottleneck residual block contains a convolution layer, a batch normalization layer and a ReLU activation layer. The output of the model is the feature representation of the input image. After 32 times downsampling, the output feature size is ; Finally, these features will be fused with the multi-scale features extracted by the fusion model according to different alignment scales.

[0108] Figure 6 This is a diagram of the architecture of a multi-scale feature fusion module provided in one embodiment of the present application. Specifically, Figure 6 As shown, the multi-scale feature fusion module includes multiple additional branches that perform convolution operations on the backbone network features and extract features from them. Each additional branch has a different number of layers, which is used to extract different levels of information from the backbone network features, making the encoder features more robust. Each branch has an input feature size of 2048 dimensions and contains two convolutional layers, each generating feature maps for different channels. The feature maps generated by these additional branches are resized to the same spatial size as the backbone feature map through bilinear interpolation and then concatenated with the original backbone features, simultaneously capturing low-level and high-level features at different spatial resolutions. The feature fusion module enhances the model's representation learning ability, enabling it to extract more comprehensive and discriminative features from the input image, ultimately improving model performance. The total number of channels output by the concatenation module is 2544×H×W, and these features are fed into the self-attention module for processing.

[0109] Figure 7 Schematic diagram of the first image feature self-attention module provided in one embodiment of the present application. Specifically, Figure 7As shown, the first image feature self-attention module is used to optimize features, selectively focusing on relevant image regions and improving the overall representation. After passing through the attention module, the output undergoes a series of operations, including average pooling, fully connected layers, ReLU activation, and dropout layers, to generate an image feature vector or a compressed representation of the image embedding. The self-attention mechanism in the encoder is a key component for modeling long-range image dependencies and capturing contextual information. Through convolution operations and attention weight calculation, this mechanism enables the model to adaptively focus on relevant image regions and strengthen the representation of important features. The image data self-attention model architecture consists of three convolutional layers with a kernel size of 1: query convolution, key convolution, and key-value convolution, which are used to reduce the dimensionality of the input features. This module takes the concatenated image features as input and generates query, key, and key-value tensors through convolution operations. Attention weights are then calculated based on the query-key similarity, and the value tensor is weighted and focused accordingly. Finally, the attended features are fused with the original features using a learnable scaling factor α.

[0110] The self-attention mechanism in the encoder is a key component for modeling long-range dependencies and capturing contextual information about the image. By applying convolution operations and calculating attention weights, the model can selectively focus on relevant image regions and enhance the representation of important features. The architecture of the self-attention model on image data consists of three convolutional layers with a kernel size of 1 to reduce the dimensionality of the input features. It takes the concatenated image features as input and applies convolution operations to convert them into query, key, and key-value tensors. It then calculates attention weights based on the similarity between the query and key and uses these weights to focus on different parts of the value tensor. The focused features are combined with the original features using a scaling parameter; the features output by the encoder module are passed as the initial input to the decoder module for the subsequent image description generation process; the features output by the encoder module are passed as the initial input to the decoder module for the image description generation task.

[0111] Figure 8 This is a schematic diagram of a decoder provided in one embodiment of the present application. Specifically, Figure 8As shown in the figure, the decoder module is mainly composed of the following components: Embedding layer: Receives input text description, converts the text into an embedded vector representation, and concatenates and fuses it with the image feature tensor output by the encoder; GRU recurrent neural network: Uses a gating mechanism to process sequence data, applies Dropout regularization (default p=0.5) before inputting the linear layer, captures temporal context information through hidden states, and ultimately outputs a predicted tag distribution; Linear output layer: Maps the GRU hidden state to the vocabulary space to generate a probability distribution of the description text. This architecture achieves end-to-end image description generation by jointly optimizing visual-linguistic feature representations. The hidden state dimension of the GRU is set to 512, and the embedding layer dimension is aligned with the image feature dimension (usually 1024 dimensions) to ensure effective feature fusion.

[0112] It can be understood that in this embodiment, the multi-dimensional hidden danger perception collaborative model with a self-learning mechanism is constructed on the cloud. Figure 9 This is a schematic diagram of a multi-dimensional collaborative model for hidden danger perception of a self-learning mechanism provided in one embodiment of the present application. Figure 9 As shown, including:

[0113] Build a self-learning mechanism based on the parallel three-dimensional ResNet, Faster-RCNN, and Vgg+Net networks

[0114] By continuously collecting new image data and detection results at the edge, mixed samples are obtained;

[0115] Performing data augmentation on the mixed samples and screening training samples through an adversarial network;

[0116] Three three-dimensional ResNet, Faster-RCNN, and Vgg+Net networks are used to predict the training samples and output their respective prediction results.

[0117] The samples of the respective prediction results are divided into obvious samples and not obvious samples according to the accuracy of the prediction, and self-learning is performed to calculate the loss weight corresponding to each sample. The training loss is optimized by weighted optimization to enable the network to better learn obvious samples and not obvious samples. The multi-dimensional perception collaborative model of hidden dangers is continuously updated and optimized through the self-learning mechanism to improve the detection performance and adaptability of the model.

[0118] Specifically, data augmentation and data expansion effectively alleviate data insufficiency and overfitting by generating diverse training samples (such as scaling, cropping, and rotation), thereby improving the model's generalization capabilities. For example, data augmentation is the first step in model training. Using strategies like mixup to generate multiple augmented versions allows the model to learn features under different transformations. This augmented data increases diversity without changing its essence, making CNNs more adaptable to the complexity of real-world scenarios.

[0119] Specifically, the ResNet residual learning framework can train deeper networks, addressing the difficulty of training deep networks. Furthermore, by integrating residual networks, it has achieved excellent results on ImageNet. This demonstrates that using models of varying depth (such as ResNet) can improve accuracy and that model integration is effective. This supports the idea of training models from different frameworks and improving results through integration, potentially corresponding to the overcoming deficiencies of different models and fusing results discussed in the previous section. VGG emphasizes increasing network depth to 16-19 layers and using small convolution kernels, which improved performance. This suggests that different network structures (such as VGG's depth and small convolution kernels) may have different feature extraction capabilities, and combining multiple such models can complement each other.

[0120] After image data is expanded, it is trained separately using three different convolutional neural network models. This overcomes the limitations of each network model and further improves recognition accuracy. The recognition results of different network models can also be combined to enhance the effectiveness of subsequent self-learning mechanisms. CNN models of different frameworks overcome specific limitations through their respective structural designs, thus complementing each other in feature extraction and classification tasks: ResNet: This introduces a residual learning framework to address the vanishing gradient problem in deep networks, allowing for the training of ultra-deep networks with more than 152 layers. It also uses skip connections to preserve low-level features, improving classification accuracy. VGG: This utilizes stacked 3×3 convolutional kernels and a 16-19-layer deep network to enhance local feature extraction capabilities, resulting in excellent performance in the ImageNet challenge. Improved CNNs: These utilize features such as a dual-channel design (using different convolutional kernels to extract local / global features) and an adaptive feature fusion module to enhance representational capabilities through multi-scale feature fusion.

[0121] S3: Use the monocular camera on the drone to collect images of dense channel transmission lines.

[0122] Understandably, monocular cameras, due to their simple structure (requiring only a single lens and sensor), are significantly less expensive than binocular or multi-camera systems. Their compact size and low power consumption make them suitable for large-scale deployment. Monocular cameras, centered around high-resolution two-dimensional imaging, combine dynamic range extension, low-light optimization, and intelligent processing technologies, making them widely used in security, autonomous driving, and other fields. However, their depth perception relies on algorithms and requires calibration and correction to address optical distortion.

[0123] Since monocular cameras can only provide two-dimensional images, they can increase the transmission speed of image data, but lack depth information, which is also a key feature. Depth information can be provided through triangulation or deep learning.

[0124] Specifically, in this embodiment, after the step S3: using a monocular camera mounted on a drone to capture images of dense channel power transmission lines, the step further includes:

[0125] The input image is preprocessed and features are extracted at key point positions. The coordinates of the features are ;

[0126] According to the depth value , convert the coordinates of the features into three-dimensional coordinates ,in,

[0127] ;

[0128] is the horizontal set distance of the monocular camera, is the baseline between the structured light projector and the monocular infrared camera, .

[0129] S4: The image to be detected is detected at the edge based on the multi-dimensional perception collaborative model of hidden dangers. When the detection result is a suspected image, the drone is commanded to return to the target according to the set flight strategy and collect image data again according to the predetermined hovering strategy; at the same time, the detection result of the suspected image and the re-collected image data are sent to the cloud, the detection task is handed over to the cloud, and the edge stops detection.

[0130] It is understood that, in this embodiment, the commanding the drone to return to the target according to the set flight strategy and to collect image data again according to the predetermined hovering strategy includes:

[0131] Divide the entire flight area into multiple sub-areas;

[0132] Selecting an optimal hovering point in each sub-area, wherein the optimal hovering point is the one that maximizes the quality of information collection;

[0133] Combining the optimal hovering points into a whole flight path;

[0134] Combining particle swarm optimization and Bayesian optimization algorithms, it generates trajectories that balance obstacle avoidance and efficiency in real time;

[0135] Image data is collected again at multiple angles through drone control instructions.

[0136] Among them, the combination of particle swarm optimization and Bayesian optimization algorithm generates a trajectory that takes into account both obstacle avoidance and efficiency in real time, including: using Bayesian optimization to establish a global probability model, handling environmental uncertainty and obstacle avoidance constraints, and updating the model in real time to reflect the current obstacle status; PSO performs local path optimization on this basis and quickly adjusts the particle swarm to adapt to dynamic changes.

[0137] Obstacle positions, speeds, and predicted trajectories are acquired in real time through sensors, and the Bayesian optimization constraint model is updated. PSO is used to generate candidate paths, evaluate fitness, and perform iterative optimization based on the fitness. The fitness function is:

[0138] ;

[0139] in, is the path length, is the obstacle avoidance penalty calculated by the distance field, is the smoothness of the path calculated by curvature or acceleration, is the dynamic constraint, are the weight factors of the corresponding parameters respectively.

[0140] Specifically, PSO performs localized fine-grained optimization within the global search direction provided by BO. Each particle represents a candidate trajectory (such as a Bezier curve control point or path node), and the fitness function comprehensively considers path length, smoothness, and obstacle avoidance penalties (e.g., a penalty function). Improved PSO strategies can dynamically adjust inertia weights and learning factors to accelerate convergence (e.g., the exponentially decreasing strategy in [1]). Hybrid mechanisms (such as those combined with simulated annealing or genetic algorithms) can also be introduced to avoid becoming trapped in local optimal solutions.

[0141] S5: Re-detecting the suspected image in the cloud based on the multi-dimensional hidden danger perception collaborative model of the self-learning mechanism, and outputting the target detection result.

[0142] Understandably, the cloud, as the core of computing terminals, relies on high-performance server clusters for large-scale data storage, in-depth analysis, and complex model training. It also provides massive storage space for historical data and global analysis results, while supporting remote monitoring and command issuance.

[0143] Specifically, in this embodiment, the cloud re-detects the suspected image based on the multi-dimensional hidden danger perception collaborative model of the self-learning mechanism and outputs the target detection result. The detection result not only comprehensively considers the multi-dimensional factor data, but also continuously updates the detection model through the self-learning mechanism, and the obtained detection result is more accurate.

[0144] S6: Generate graded warning signals based on the dynamic risk assessment model to achieve accurate warning of hidden danger location errors.

[0145] It is understood that, in this embodiment, the generation of graded warning signals based on the dynamic risk assessment model to achieve accurate warning of hidden danger location errors includes:

[0146] Integrate multi-dimensional data to build a risk assessment model, including target detection results and electrical, mechanical, and environmental data collected by multiple terminals;

[0147] Perform risk assessment using the risk assessment model to obtain a risk value;

[0148] Generate graded warning signals based on risk values to achieve accurate warning of hidden danger location errors;

[0149] The risk value of the risk assessment model is updated through incremental data, and automatically adjusted according to new data to maintain the timeliness of the assessment.

[0150] Example 2

[0151] like Figure 10 As shown, the present application provides an intelligent vision and state perception collaborative hidden danger monitoring system for dense channel transmission line faults, which is applied to the intelligent vision and state perception collaborative hidden danger monitoring method for dense channel transmission line faults as described in Example 1, including: a monitoring architecture construction module 11, a model construction module 12, an image acquisition module 13, a first detection module 14, a second detection module 15, and a graded warning module 16.

[0152] It can be understood that, in this embodiment, the monitoring architecture construction module 11 constructs a state-aware collaborative hidden danger monitoring architecture, and the state-aware collaborative hidden danger monitoring architecture includes an edge end, a transmission channel, an intermediate node and a cloud.

[0153] It can be understood that in this embodiment, the model construction module 12: constructs a hidden danger multi-dimensional perception collaborative model based on the fusion of residual network and dynamic convolutional network model on the edge end, and constructs a hidden danger multi-dimensional perception collaborative model with a self-learning mechanism on the cloud.

[0154] It can be understood that, in this embodiment, the image acquisition module 13 utilizes a monocular camera carried by the drone to acquire images of the dense channel power transmission line.

[0155] It can be understood that in this embodiment, the first detection module 14: detects the image to be detected based on the hidden danger multi-dimensional perception collaborative model at the edge end, and when the detection result is a suspected image, commands the drone to return to the target according to the set flight strategy, and collects image data again according to the predetermined hovering strategy; at the same time, the detection result of the suspected image and the re-collected image data are sent to the cloud, the detection task is handed over to the cloud, and the edge end stops detection.

[0156] It can be understood that, in this embodiment, the second detection module 15 re-detects the suspected image based on the multi-dimensional hidden danger perception collaborative model of the self-learning mechanism in the cloud, and outputs the target detection result.

[0157] It can be understood that, in this embodiment, the graded warning module 16 generates graded warning signals based on the dynamic risk assessment model to achieve accurate warning of hidden danger location errors.

[0158] This application provides a method and system for collaborative hidden danger monitoring of dense channel transmission line faults using intelligent vision and state perception. The system constructs a state perception collaborative hidden danger monitoring architecture, building a multi-dimensional hidden danger perception collaborative model and a self-learning hidden danger perception collaborative model on the edge and cloud, respectively. The system uses a monocular camera mounted on an unmanned aerial vehicle (UAV) to capture images of dense channel transmission lines. Primary detection is performed at the edge, and secondary detection is performed in the cloud when suspicious images are detected, outputting target detection results. A graded warning signal is generated. This improves data transmission efficiency and enhances the precision and accuracy of collaborative hidden danger monitoring using intelligent vision and state perception for dense channel transmission line faults.

[0159] Using a monocular camera to capture images of densely packed transmission lines improves data transmission efficiency. A multi-dimensional hidden danger perception collaborative model, integrating an improved residual network with a dynamic convolutional network model, enhances evaluation results. By performing primary detection at the edge and secondary detection in the cloud when suspicious images are detected, target detection results are more accurate. This approach is easy to promote and apply, and its operation is relatively simple, making it suitable for widespread application in a wider range of scenarios.

[0160] Figure 11 This is an electronic device provided by an embodiment of the present application. Figure 11 As shown, the electronic device includes at least the following parts: a processor 101 and a memory 100 , a communication interface 103 , and a bus 102 .

[0161] In the embodiment of the present application, the memory 100 is used to store instructions executable by the processor 101. The processor 101 is configured to execute the instructions to implement the following Figure 10 The figure shows an intelligent vision and state perception collaborative hidden danger monitoring system for dense channel transmission line faults.

[0162] In an embodiment of the present application, a computer-readable storage medium includes instructions, and the instructions instruct a device to execute the method of the first aspect. For example, the instructions instruct the device to execute Figure 1 The method is shown in the process steps.

[0163] The program running in the electronic device involved in one embodiment of the present application may be a program that controls a central processing unit (CPU) and the like to implement the functions of the above-mentioned embodiment involved in one embodiment of the present invention (a program that causes a computer to function). The information processed by these devices is temporarily stored in random access memory (RAM) while being processed, and is then stored in various ROMs such as read-only memory (Flash ROM) and a hard disk drive (HDD), where it is read, modified, and written as needed by the CPU.

[0164] It should be noted that a portion of the electronic device of the above embodiment may also be implemented by a computer. In this case, a program for implementing the control function may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read into a computer and executed.

[0165] It should be noted that the "computer" mentioned here refers to a computer built into an electronic device, employing hardware including an operating system (OS) and peripheral devices. Furthermore, "computer-readable recording medium" refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computers.

[0166] Furthermore, "computer-readable recording media" may include: media that dynamically store programs for a short period of time, such as communication lines when transmitting programs via networks such as the Internet or communication lines such as telephone lines; and media that store programs for a fixed period of time, such as volatile memory within computers acting as servers or clients in this context. Furthermore, the aforementioned program may be a program for implementing a portion of the aforementioned functions, or a program that can achieve the aforementioned functions by combining with a program already stored in a computer.

[0167] Furthermore, the electronic device in the above-described embodiments can also be implemented as a collection of multiple devices (a device group). Each device comprising the device group may include some or all of the functions or functional blocks of the electronic device in the above-described embodiments. A device group only needs to include all of the functions or functional blocks of the electronic device.

[0168] Those skilled in the art should recognize that the above embodiments are merely intended to illustrate the present application and are not intended to limit the present application. As long as they are within the spirit of the present application, appropriate changes and modifications to the above embodiments are within the scope of protection claimed in the present application.

Claims

1. A method for monitoring hidden dangers of dense channel transmission line faults by intelligent vision and state perception collaboration, characterized in that: The method comprises: S1: Build a state-aware collaborative hidden danger monitoring architecture, which includes edge devices, transmission channels, intermediate nodes, and the cloud. S2: Constructing a multi-dimensional hidden danger perception collaborative model based on the fusion of residual network and dynamic convolutional network model on the edge, and constructing a multi-dimensional hidden danger perception collaborative model with a self-learning mechanism on the cloud; S3: Use a monocular camera on a drone to collect images of dense channel transmission lines; S4: The edge end detects the image to be detected based on the multi-dimensional hidden danger perception collaborative model. If the detection result is a suspected image, the drone is commanded to return to the target according to the set flight strategy and re-collect image data according to the predetermined hovering strategy. At the same time, the detection result of the suspected image and the re-collected image data are sent to the cloud, and the detection task is handed over to the cloud, and the edge end stops detection. S5: Re-detecting the suspected image in the cloud based on the multi-dimensional hidden danger perception collaborative model of the self-learning mechanism, and outputting a target detection result; S6: Generate graded warning signals based on the dynamic risk assessment model to achieve accurate warning of hidden danger location errors; S2: Constructing a multi-dimensional hidden danger perception collaborative model based on the fusion of residual network and dynamic convolutional network model on the edge end, including: The multi-dimensional perception collaborative model includes an encoder and a decoder; The encoder is composed of a ResNet-CNN architecture, including a residual network ResNet module, a multi-scale feature fusion module, and a first image feature self-attention module. The residual network ResNet module is used as the backbone network and consists of two modules. The first module includes an input layer, a convolution layer, a batch normalization layer, a ReLU activation function layer, and a maximum pooling layer. The input layer has a size of 3 channels. The RGB image is taken as input; the second module contains 7 sequential networks with the same structure, each of which consists of multiple bottleneck residual blocks, each of which contains a convolutional layer, a batch normalization layer and a ReLU activation layer, and the output feature size is ; The multi-scale feature fusion module contains five different processing channels, each with a different number of layers. The input feature size is 2048 and contains two convolutional layers to generate feature maps for different channels. The feature maps are resized by bilinear interpolation to match the spatial dimensions of the backbone features. The resized feature maps are concatenated with the backbone features to capture both low-level and high-level features at different spatial resolutions. The total number of output channels of the concatenation module is 2544×H×W, which are then fed into the self-attention module for processing. The first image feature self-attention module consists of parallel query convolution, key convolution, and value convolution. The input is the features of the connection module, which are converted into query, key, and value tensors through the convolution operations of query convolution, key convolution, and value convolution respectively. The attention weight is calculated by the similarity between the query and key tensors, and then weighted focusing is performed with the value tensor to selectively focus on relevant image areas and enhance the representation of important features. The decoder includes an embedding layer, a GRU recurrent neural network, and a linear output layer, wherein the embedding layer is used to receive input text descriptions, convert the text into an embedded vector representation, and splice and fuse it with the image feature tensor output by the encoder; the GRU recurrent neural network is used to process sequence data using a gating mechanism, apply Dropout regularization before inputting the linear layer, capture temporal context information through hidden states, and ultimately output a predicted tag distribution; the linear output layer is used to map the GRU hidden states to the vocabulary space to generate a probability distribution of the description text; this architecture achieves end-to-end image description generation by jointly optimizing the visual-linguistic feature representation.

2. The method for monitoring hidden dangers of dense channel power transmission line faults by intelligent vision and state perception collaboration according to claim 1 is characterized in that: S1: Constructing a state-aware collaborative hidden danger monitoring architecture, which includes edge devices, transmission channels, intermediate nodes, and the cloud, including: The state-aware collaborative hidden danger monitoring architecture consists of five layers: terminal layer, first communication layer, edge layer, second communication layer, and cloud layer. Wherein, the terminal layer is used to collect data by receiving instructions; The first communication layer is used for data transmission between the terminal layer and the edge layer; The edge layer is used to receive transmission data, perform data processing and detection, and send control instructions based on the detection results; The second communication layer is used for data transmission between the edge layer and the cloud layer; The cloud layer is used to receive and transmit data, perform deeper data processing and detection, and perform data management, data storage, and data visualization.

3. The method for monitoring hidden dangers of dense channel power transmission line faults by intelligent vision and state perception collaboration according to claim 2 is characterized in that: The multi-dimensional hidden danger perception collaborative model with a self-learning mechanism is constructed on the cloud, including: Build a self-learning mechanism based on the parallel three-dimensional ResNet, Faster-RCNN, and Vgg+Net networks; By continuously collecting new image data and detection results at the edge, mixed samples are obtained; Performing data augmentation on the mixed samples and screening training samples through an adversarial network; Use three three-dimensional ResNet, Faster-RCNN, and Vgg+Net networks to predict the training samples and output their respective prediction results; The samples of the respective prediction results are divided into obvious samples and not obvious samples according to the accuracy of the prediction, and self-learning is performed to calculate the loss weight corresponding to each sample. The training loss is optimized by weighted optimization to enable the network to better learn obvious samples and not obvious samples. The multi-dimensional perception collaborative model of hidden dangers is continuously updated and optimized through the self-learning mechanism to improve the detection performance and adaptability of the model.

4. The method for monitoring hidden dangers of dense channel power transmission line faults by intelligent vision and state perception collaboration according to claim 3 is characterized in that: After the S3: using the monocular camera carried by the UAV to collect images of the dense channel transmission line, the method further includes: The input image is preprocessed and features are extracted at key point positions. The coordinates of the features are ; According to the depth value , convert the coordinates of the features into three-dimensional coordinates ,in, , is the horizontal set distance of the monocular camera, is the baseline between the structured light projector and the monocular infrared camera, .

5. The method for monitoring hidden dangers of dense channel power transmission line faults by intelligent vision and state perception collaboration according to claim 4 is characterized in that: The commanding the drone to return to the target according to the set flight strategy and to collect image data again according to the predetermined hovering strategy includes: Divide the entire flight area into multiple sub-areas; Selecting an optimal hovering point in each sub-area, wherein the optimal hovering point is the one that maximizes the quality of information collection; Combining the optimal hovering points into a whole flight path; Combining particle swarm optimization and Bayesian optimization algorithms, it generates trajectories that balance obstacle avoidance and efficiency in real time; Re-collect image data from multiple angles through drone control commands; The method combines particle swarm optimization and Bayesian optimization algorithms to generate trajectories that balance obstacle avoidance and efficiency in real time. This includes: using Bayesian optimization to establish a global probability model to handle environmental uncertainty and obstacle avoidance constraints, and updating the model in real time to reflect the current obstacle status; PSO performs local path optimization on this basis, and quickly adjusts the particle swarm to adapt to dynamic changes; Obstacle positions, speeds, and predicted trajectories are acquired in real time through sensors, and the Bayesian optimization constraint model is updated. PSO is used to generate candidate paths, evaluate fitness, and perform iterative optimization based on the fitness. The fitness function is: , in, is the path length, is the obstacle avoidance penalty calculated by the distance field, is the smoothness of the path calculated by curvature or acceleration, is the dynamic constraint, are the weight factors of the corresponding parameters respectively.

6. The method for monitoring hidden dangers of dense channel power transmission line faults by intelligent vision and state perception collaboration according to claim 5 is characterized in that: The generation of graded warning signals based on the dynamic risk assessment model to achieve accurate warning of hidden danger location errors includes: Integrate multi-dimensional data to build a risk assessment model, including target detection results data and electrical, mechanical, and environmental data collected by multiple terminals; Perform risk assessment using the risk assessment model to obtain a risk value; Generate graded warning signals based on risk values to achieve accurate warning of hidden danger location errors; The risk value of the risk assessment model is updated through incremental data, and automatically adjusted according to new data to maintain the timeliness of the assessment.

7. A system for monitoring hidden dangers of dense channel transmission line faults by intelligent vision and state perception collaboration, applied to the method for monitoring hidden dangers of dense channel transmission line faults by intelligent vision and state perception collaboration as claimed in any one of claims 1 to 6, characterized in that: include: Monitoring architecture building module: Build a state-aware collaborative hidden danger monitoring architecture, which includes edge devices, transmission channels, intermediate nodes, and the cloud; Model construction module: constructing a multi-dimensional hidden danger perception collaborative model based on the fusion of residual network and dynamic convolutional network model on the edge end, and constructing a multi-dimensional hidden danger perception collaborative model with a self-learning mechanism on the cloud end; Image acquisition module: uses the monocular camera on the drone to collect images of dense channel transmission lines; The first detection module detects the image to be detected at the edge based on the multi-dimensional hidden danger perception collaborative model. If the detection result is a suspected image, the drone is commanded to return to the target according to the set flight strategy and re-collect image data according to the predetermined hovering strategy. At the same time, the detection result of the suspected image and the re-collected image data are sent to the cloud, and the detection task is handed over to the cloud, and the edge stops detection. The second detection module re-detects the suspected image in the cloud based on the multi-dimensional hidden danger perception collaborative model of the self-learning mechanism and outputs the target detection result; Grading warning module: Generates graded warning signals based on the dynamic risk assessment model to achieve accurate warning of hidden danger location errors.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the intelligent vision and state perception collaborative hidden danger monitoring method for dense channel transmission line faults as described in any one of claims 1 to 6 when executing the instructions.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, and the program instructs the device to execute the intelligent vision and state perception collaborative hidden danger monitoring method for dense channel transmission line faults as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Power transmission line refined sensing method based on cloud side cooperation

    CN112491982A

  • Self-learning identification system and method for external hidden dangers based on power transmission line

    CN114863118A