Power transmission line hidden danger detection method
Through multimodal perception model and contrast learning technology, combined with the transmission line images and text data obtained by the drone, efficient hidden danger detection in harsh environments is achieved, the problems of unstable image quality and low recognition accuracy are solved, and the accuracy and intelligence of detection are improved.
Patent Information
- Application Number
- CN202510171599.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-17
AI Technical Summary
Under severe weather conditions, the hidden danger detection of transmission lines faces the problems of unstable image quality and low recognition accuracy, and traditional methods are difficult to conduct accurate detection in complex environments.
The multimodal perception model is adopted to obtain transmission line image data through drones, combine the corresponding text data, and use visual encoder and text encoder for feature extraction. Through comparative learning and language model optimization, the semantic alignment of image and text features is achieved, and the hidden danger description and recognition types are finally generated.
It improves the accuracy of hidden danger detection in harsh environments such as wind disasters, enhances the intelligence of the system, realizes real-time hidden danger detection and early warning of potential problems, and reduces computing resource consumption and data transmission delay.
Smart Images

Figure CN120107829A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of power transmission line detection, and in particular relates to a method for detecting hidden dangers of power transmission lines. Background Art
[0002] With the development of modern power grids, transmission lines, as an important part of the power system, bear the key task of transmitting electricity. In order to ensure the safe and stable operation of the power system, the hidden danger detection and maintenance of transmission lines have become an important part of ensuring power supply. However, in the process of inspection of transmission lines, there are often many challenges, especially in severe weather conditions such as wind disasters, the identification and location of hidden dangers become more difficult.
[0003] Common hidden dangers of power transmission lines include broken conductors, foreign objects on the conductors (such as floating objects), and loose hardware. Broken conductors may cause conductor breakage, which may cause large-scale power outages in severe cases; floating objects on the conductors, especially in windy weather, may cause changes in the electrical properties of the conductors or damage power equipment; loose hardware may cause loose line connections, affecting power transmission safety. Traditional manual inspection methods rely on visual inspections by staff or the use of portable detection equipment, which is inefficient and costly. In addition, the difficulty and risk of manual inspections are greatly increased in bad weather.
[0004] Under severe weather conditions such as wind disasters, the detection of hidden dangers in power transmission lines faces more complex challenges. Increased wind speed and changes in precipitation will seriously affect the stability of image quality during inspections. Strong winds can cause the wires to vibrate violently, affecting the clarity of the image; at the same time, rain and snowy weather can cause image blur and even obscure the appearance of some wires or hardware, increasing the difficulty of identifying hidden dangers. In addition, traditional image processing methods are difficult to effectively respond to these dynamic changes, resulting in reduced recognition accuracy and difficulty in accurate detection in complex environments.
[0005] At present, the detection technologies for hidden dangers in power transmission lines mainly include manual inspections, drone inspections, and automated detection methods based on image processing and sensor data. Manual inspections have the problems of long cycles, low efficiency, high costs, and difficulty in operation in bad weather; drone inspections can improve inspection efficiency, but are limited by image transmission, weather influences, and the limitations of recognition algorithms, making it difficult to achieve efficient recognition in complex environments. Although automated methods based on image processing and sensor data have made some progress in some scenarios, they still face problems such as low image quality and cumbersome data processing in complex environments such as wind disasters, rain and snow. Summary of the invention
[0006] The purpose of the present invention is to provide a method for detecting hidden dangers in a power transmission line in order to overcome the defects of the prior art.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] The present invention provides a method for detecting hidden dangers of a power transmission line, comprising:
[0009] Acquire hidden danger images of the transmission line, wherein the hidden danger images include broken strand images, wire foreign matter images, and loose pin images;
[0010] Acquire text data corresponding to the hidden danger image, wherein the text data includes a text description of a broken strand image, a text description of a wire foreign body image, and a text description of a loose pin image from a pin sensor;
[0011] A transmission line hidden danger detection model is constructed based on hidden danger images and text data, wherein the transmission line hidden danger detection model includes a visual encoder, a text encoder, a classification head, and a language model; and a transmission line hidden danger detection model is constructed based on hidden danger images and text data;
[0012] Acquire image data of power transmission lines taken by drones;
[0013] Input the acquired image data into the visual encoder of the power transmission line hidden danger detection model, input the visual features output by the visual encoder into the language model, and output the text data corresponding to the acquired image data through the language model;
[0014] Based on the output text data, determine whether there are hidden dangers in the transmission line and the corresponding hidden danger types and descriptions.
[0015] Furthermore, the construction of a transmission line hidden danger detection model based on hidden danger images and text data specifically includes:
[0016] Preprocessing the hidden danger image, wherein the preprocessing includes image normalization, data enhancement, denoising and deblurring operations;
[0017] Encoding the preprocessed hidden danger images through a visual encoder to extract the first visual features of each hidden danger image;
[0018] Encoding the text data by a text encoder to extract the first text feature of each text description;
[0019] The image encoder and the text encoder are trained according to the first visual features of each hidden danger image and the first text features of each text description by contrastive learning, and the parameters of the image encoder and the text encoder are updated by a first loss function;
[0020] Input the extracted first visual feature and the first text feature into the classification head, output the probability of each feature matching, and update the parameters of the image encoder, the text encoder and the classification head through the second loss function according to the output probability;
[0021] After the image encoder and the text encoder are trained, the preprocessed hidden danger images are encoded by the trained visual encoder to extract the second visual features of each hidden danger image;
[0022] Encode the text data through the trained text encoder to extract the second text features of each text description;
[0023] The language model receives the second visual feature extracted by the visual encoder, takes the second text feature as a target, updates the parameters of the language model based on the third loss function, and obtains a trained language model.
[0024] Furthermore, encoding the preprocessed hidden danger images by a visual encoder to extract the first visual features of each hidden danger image specifically includes:
[0025] Divide the hidden danger image into a number of image patches of fixed sizes;
[0026] Expand each image patch into a one-dimensional vector and encode its position;
[0027] The one-dimensional vector of each image patch and its position encoding are input into a visual encoder, wherein the visual encoder adopts a network structure based on a self-attention mechanism and uses a multi-layer self-attention mechanism to capture the long-range dependencies between different regions of the image;
[0028] The self-attention mechanism of the visual encoder is used to calculate the similarity between each patch and other patches, and the representation of the one-dimensional vector of each image patch is adjusted according to the similarity;
[0029] The adjusted one-dimensional vector of image patches is further processed by a feed-forward neural network including a non-linear activation function and layer normalization;
[0030] The one-dimensional vector features of the processed image patches are fused to obtain the first visual features of each hidden danger image.
[0031] Furthermore, encoding the text data by a text encoder to extract the first text feature of each text description specifically includes:
[0032] Perform word segmentation on the input text data and divide the text data into several sub-word units;
[0033] Map each subword unit to the corresponding word vector representation;
[0034] Input the word vector representation into a text encoder, wherein the text encoder is a language model BERT based on a bidirectional attention mechanism;
[0035] The language model BERT is used to calculate the contextual dependencies of each subword unit, and the word vector representation of each subword is updated based on the contextual information.
[0036] The updated sub-word representations are merged into the overall text feature representation of the text through a pooling operation or an aggregation strategy to obtain the first text feature of each text description.
[0037] Furthermore, the image encoder and the text encoder are trained according to the first visual features of each hidden danger image and the first text features of each text description by contrastive learning, specifically including:
[0038] Pairing the first visual feature of the hidden danger image with the first text feature of the text data to form positive samples and negative samples, wherein the positive samples include pairs where the image and its corresponding text description are semantically matched, and the negative samples include pairs where there is no actual semantic match between the image and its corresponding text description;
[0039] Calculate the image features v of positive and negative samples z With the text feature v T The image encoder and text encoder are optimized by the similarity and the first loss function.
[0040] Furthermore, the first loss function is:
[0041]
[0042] in, is the first loss function, N B is the data size of a training batch, τ is the learnable coefficient, <·,·> represents the dot product, The image features of the i-th hidden danger image in the training data, is the text feature corresponding to the i-th hidden danger image in the training data.
[0043] Furthermore, the step of inputting the extracted first visual feature and the first text feature into a classification head, outputting the probability of each feature matching, and updating the parameters of the image encoder, the text encoder, and the classification head through a second loss function according to the output probability specifically includes:
[0044] The first visual feature of the hidden danger image extracted by the visual encoder and the first text description text features extracted by the text encoder Input a classification head, which is a binary classification network module, and is used to output the probability p of whether the first visual feature matches the first text feature. v ;
[0045] The classification head processes the first visual feature v of the input z and the first text feature v T , output a probability p v , represents the matching degree between image and text in semantic space. If the image and text match semantically, the output probability is close to 1, and if they do not match, the output probability is close to 0;
[0046] During the training process, the positive sample is a pair of images and their corresponding text descriptions that match, and its label is y=1, and the negative sample is a pair of images and text descriptions that do not match, and its label is y=0;
[0047] According to the output matching probability p v The difference between the image and the actual label y is used to update the parameters of the image encoder, text encoder, and classification head through the second loss function.
[0048] Furthermore, the second loss function is:
[0049]
[0050] in, is the second loss function, is the mathematical expectation.
[0051] Furthermore, the language model is a language model LLaMa based on an autoregressive mechanism;
[0052] The language model receives the second visual feature extracted by the visual encoder, takes the second text feature as a target, and updates the parameters of the language model based on the third loss function, specifically including:
[0053] The second visual feature is input into the language model. Given the visual feature v z As input, the language model outputs a conditional probability distribution The conditional probability distribution Describes the probability of a word at each position in a text sequence, given the previously generated text and visual features v z In the case of, generate the word at position m in the text The conditional probability of A third loss function is calculated, and the language model is updated according to the calculated third loss function.
[0054] Furthermore, the third loss function is:
[0055]
[0056] in, is the third loss function, is the mathematical expectation, represents all text features before the predicted position m, Represents the text generated given the previous and image features v z In the case of, generate the word at position m in the text The conditional probability of .
[0057] Compared with the prior art, the present invention has the following advantages:
[0058] (1) The present invention first uses contrastive learning to ensure that the visual encoder and the text encoder can find a corresponding relationship in the semantic space by training the feature matching of the visual encoder and the text encoder. Specifically, each image feature is semantically matched with its corresponding text description. The goal is to make the visual features extracted by the visual encoder and the text features extracted by the text encoder have a higher similarity in the high-dimensional space through training. Through feature matching training, it is ensured that there is a high degree of consistency between the visual features output from the visual encoder and the text features output from the text encoder. Specifically, the visual features and the text features will be aligned according to their respective characteristics. This alignment is not only achieved at the surface level (such as label matching between images and texts), but also finds a contrast relationship in deep semantics, so that the model can understand the essential connection between images and texts. The trained visual encoder and text encoder output high-quality visual features and text features respectively. The visual features are used as the input of the language model, and the text features are used as the target (i.e., label) of the language model. In this way, the language model can learn to infer relevant text descriptions from visual features, further optimize its generation ability, and enable it to understand and generate natural language descriptions that are highly related to visual information. This process enables the language model to recognize the visual features in image data and generate hidden danger description text that is consistent with the actual scenario based on these features.
[0059] (2) The present invention adopts a language model based on an autoregressive mechanism (such as LLaMa), combines image and text features, and then generates and infers text through autoregression. This technical means enables the system to not only identify hidden dangers in images, but also generate text descriptions related to hidden dangers and achieve abnormal prediction. For example, after receiving image data of loose pins, the model not only identifies the loose pin problem, but also generates corresponding prediction text (such as "pin loose") through reasoning. This abnormal prediction capability enhances the intelligence of the system, so that hidden danger detection is not limited to the recognition of the current state, but also can provide early warning of potential problems.
[0060] (3) The present invention realizes real-time hidden danger detection by designing a lightweight multimodal perception model and deploying it on the UAV terminal side. Compared with the traditional method that relies on cloud computing, the terminal-side reasoning solution of the present invention can not only greatly reduce the consumption of computing resources, but also avoid data transmission delays, thereby improving the real-time performance and inspection efficiency of hidden danger detection. Especially during the inspection process, the model can identify and feedback hidden danger information in real time, providing more timely decision-making support for power operation and maintenance personnel.
[0061] (4) The present invention introduces multimodal contrastive learning technology, which enables the features of images and texts to be aligned in a shared space through contrastive learning methods. Specifically, the image encoder and the text encoder can ensure that the image and text are semantically consistent through joint training. During the training process, the semantic matching degree of the image and text is optimized through the comparison of positive and negative samples. Multimodal contrastive learning not only improves the representation ability of image and text features, but also enhances the robustness of the model, enabling it to maintain a high recognition accuracy in complex environments (such as wind disasters, heavy rain, etc.).
[0062] (5) The visual encoder of the present invention adopts a network structure based on ViT, which captures the long-term dependencies between different areas of the image through the self-attention mechanism, and can effectively extract the hidden danger features in the image. Compared with the traditional convolutional neural network, ViT divides the image into several patches of fixed size and enhances the global information processing capability of the image through the self-attention mechanism, so that hidden dangers can still be accurately identified in complex situations such as wire vibration and image blur. Through the introduction of this self-attention mechanism, the image recognition capability of the present invention is significantly improved, which is particularly suitable for complex scenes in power line inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a flow chart of the method of the present invention;
[0064] Figure 2 This is a transmission line hidden danger detection model diagram of the present invention. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0066] Embodiment 1:
[0067] The present invention provides a method for detecting hidden dangers in a power transmission line. Figure 1 As shown, including:
[0068] Obtaining hidden danger images of transmission lines, including broken strands, foreign objects in conductors, and loose pins;
[0069] Acquire text data corresponding to the hidden danger image, the text data including a text description of a broken strand image, a text description of a wire foreign body image, and a text description of a loose pin image from a pin sensor;
[0070] A transmission line hidden danger detection model is constructed based on hidden danger images and text data, such as Figure 2 As shown, the transmission line hidden danger detection model includes a visual encoder, a text encoder, a classification head, and a language model;
[0071] Acquire image data of power transmission lines taken by drones;
[0072] The acquired image data is input into the visual encoder of the power transmission line hidden danger detection model, and the text data corresponding to the acquired image data is output through the language model;
[0073] Based on the output text data, determine whether there are hidden dangers in the transmission line and the corresponding hidden danger types and descriptions.
[0074] A transmission line hidden danger detection model is constructed based on hidden danger images and text data, including:
[0075] Preprocess the hidden danger images, including image normalization, data enhancement, denoising and deblurring operations;
[0076] Encode the preprocessed hidden danger images through a visual encoder to extract the visual features of each hidden danger image;
[0077] Encode the text data through a text encoder to extract text features of each text description;
[0078] The image encoder and the text encoder are trained according to the visual features of each hidden danger image and the text features of each text description by contrastive learning, and the parameters of the image encoder and the text encoder are updated by a first loss function;
[0079] The extracted visual features and text features are input into the classification head, the probability of each feature matching is output, and the parameters of the image encoder, text encoder and classification head are updated through the second loss function according to the output probability;
[0080] The extracted visual features and text features are input into the language model, and the parameters of the language model are updated through the third loss function.
[0081] The preprocessed hidden danger images are encoded by the visual encoder to extract the visual features of each hidden danger image, including:
[0082] Divide the hidden danger image into a number of image patches of fixed sizes;
[0083] Expand each image patch into a one-dimensional vector and encode its position;
[0084] The one-dimensional vector of each image patch and its position encoding are input into the visual encoder. The visual encoder adopts a network structure based on the self-attention mechanism and uses a multi-layer self-attention mechanism to capture the long-range dependencies between different regions of the image.
[0085] The self-attention mechanism of the visual encoder is used to calculate the similarity between each patch and other patches, and the representation of the one-dimensional vector of each image patch is adjusted according to the similarity;
[0086] The adjusted one-dimensional vector of image patches is further processed through a feed-forward neural network, which includes non-linear activation functions and layer normalization;
[0087] The one-dimensional vector features of the processed image patches are fused to obtain the visual features of each hidden danger image.
[0088] The text data is encoded through a text encoder to extract the text features of each text description, including:
[0089] Perform word segmentation on the input text data and divide the text data into several sub-word units;
[0090] Map each subword unit to the corresponding word vector representation;
[0091] The word vector representation is input into the text encoder, which is a language model BERT based on the bidirectional attention mechanism;
[0092] The language model BERT is used to calculate the contextual dependencies of each subword unit, and the word vector representation of each subword is updated based on the contextual information.
[0093] The updated sub-word representations are merged into the overall text feature representation of the text through a pooling operation or aggregation strategy to obtain the text features of each text description.
[0094] The image encoder and text encoder are trained by contrastive learning based on the visual features of each hidden danger image and the text features of each text description, including:
[0095] Pairing the visual features of the hidden danger image with the text features of the text data to form positive samples and negative samples. The positive samples include pairs where the image and its corresponding text description are semantically matched, and the negative samples include pairs where there is no actual semantic match between the image and its corresponding text description.
[0096] Calculate the image features v of positive and negative samples z With the text feature v T The image encoder and text encoder are optimized by the similarity and the first loss function.
[0097] The first loss function is:
[0098]
[0099] in, is the first loss function, N B is the data size of a training batch, τ is the learnable coefficient, <·,·> represents the dot product, The image features of the i-th hidden danger image in the training data, is the text feature corresponding to the i-th hidden danger image in the training data.
[0100] The extracted visual features and text features are input into the classification head, and the probability of each feature matching is output. According to the output probability, the parameters of the image encoder, text encoder and classification head are updated through the second loss function, which specifically includes:
[0101] The visual features of the hidden danger image extracted by the visual encoder and the text description text features extracted by the text encoder Input classification head, which is a binary classification network module, is used to output the probability p of whether the visual features match the text features. v ;
[0102] The classification head processes the input visual features v z and text feature v T , output a probability p v , represents the matching degree between image and text in semantic space. If the image and text match semantically, the output probability is close to 1, and if they do not match, the output probability is close to 0;
[0103] During the training process, the positive sample is a pair of images and their corresponding text descriptions that match, and its label is y=1, and the negative sample is a pair of images and text descriptions that do not match, and its label is y=0;
[0104] According to the output matching probability p vThe difference between the image and the actual label y is used to update the parameters of the image encoder, text encoder, and classification head through the second loss function.
[0105] The second loss function is:
[0106]
[0107] in, is the second loss function, is the mathematical expectation.
[0108] The language model is the language model LLaMa based on the autoregressive mechanism.
[0109] The third loss function is:
[0110]
[0111] in, is the third loss function, is the mathematical expectation, represents all text features before the predicted position m, Represents the text generated given the previous and image features v z In the case of, generate the word at position m in the text The conditional probability of .
[0112] Embodiment 2:
[0113] The parts not mentioned in this embodiment are the same as those in Embodiment 1.
[0114] This embodiment demonstrates a method for detecting hidden dangers in power transmission lines based on multimodal data fusion and deep learning technology through specific implementation. This method combines image enhancement technology in a wind disaster environment, multimodal fusion of image and text data, and language model reasoning based on an autoregressive mechanism, ultimately achieving efficient and accurate detection of hidden dangers in power transmission lines.
[0115] In this embodiment, first, the transmission line images taken by the drone and the text data obtained from the line hardware sensor are used as input data. The input images include broken strands, wire foreign matter, and loose pins, while the text data includes the corresponding text descriptions of these images and the displacement information from the pin sensor. These input data are pre-processed and then processed in different modules of the model.
[0116] The image data is first preprocessed, including image normalization, data enhancement, denoising and deblurring. The preprocessed image data is fed into the image encoder, which processes the image data through the self-attention mechanism based on the ViT (Vision Transformer) structure. ViT divides the image into several patches of fixed size, each of which is expanded into a one-dimensional vector and input into the encoder in combination with the position encoding. Through the self-attention mechanism, each patch in the image can exchange information with other patches to capture long-range dependencies, thereby enhancing the global understanding of the image. This structure is particularly suitable for processing complex vibrations, blur or occlusion of wires in images, and can effectively improve the ability to extract image features, thereby providing more accurate feature representation for subsequent hidden danger identification.
[0117] The text data input together with the image data is processed by the text encoder. The text data is first processed by word segmentation to divide the text into sub-word units (tokens). Each sub-word unit is mapped to the corresponding vector space, and then feature extraction is performed through a language model based on the BERT bidirectional attention mechanism. The BERT model extracts the semantic information of the text and generates a feature representation of the text by learning the contextual relationship between each word in the text. In this embodiment, the text data mainly includes the text description corresponding to the image and the displacement description of the sensor data. These text data provide the model with additional information that the image data cannot directly obtain. By combining the text information of the image and sensor data, the model can perform hidden danger analysis from multiple dimensions, improving the comprehensiveness and accuracy of hidden danger detection.
[0118] The output features of the image encoder and text encoder are sent to the classification head for feature matching. The classification head uses a binary classification network module to output the probability of whether the image and text features match. The parameters of the classification head, image encoder, and text encoder are updated based on the output probability and the second loss function, where the second loss function is:
[0119]
[0120] in, is the second loss function, is the mathematical expectation.
[0121] The main purpose of setting the classification head and the second loss function in the present invention is to optimize the output features of the image encoder so that it can be more accurately aligned with the text description, thereby improving the expressiveness of the image features. The role of the classification head is to match the image features extracted by the image encoder with the text description to determine whether the two are semantically consistent. Through this matching process, the model can learn how to correspond the hidden danger features in the image with the corresponding text description during the training process, and calculate the matching degree between the image features and the text features through the second loss function, optimize the parameters of the image encoder, so that it can better extract key features related to hidden dangers from the image. This process effectively improves the recognition accuracy of the image encoder in practical applications, especially in harsh environments such as wind disasters, and can better handle image blur, noise or occlusion problems, improving the accuracy and robustness of hidden danger identification.
[0122] At the same time, the system further optimizes the semantic alignment between image and text features through multimodal contrastive learning methods. During the training process, positive samples are pairs of images and their corresponding text descriptions, and negative samples are pairs of images and irrelevant text. By calculating the contrast loss between positive and negative samples, the image encoder and text encoder can jointly learn more accurate feature representations, ensuring that the semantics between images and texts remain consistent, thereby improving the accuracy of hidden danger identification.
[0123] Multimodal contrastive learning aims to jointly learn the shared feature representation technology of different modalities (such as images, text, etc.) through contrastive learning methods. This method can learn effective semantic alignment between multiple modalities, so that information from different modalities can be compared and matched in the same representation space.
[0124] Positive samples and negative samples: In multimodal contrastive learning, positive samples refer to corresponding paired data between different modalities (such as an image and its correct description of the text), and negative samples refer to those mismatched pairs (such as an image and an irrelevant text description).
[0125] Contrastive Loss: The goal of contrastive learning is to optimize a loss function so that the representation of positive samples is as close as possible and the representation of negative samples is as far away as possible. The contrastive loss function is often InfoNCE Loss, which maximizes the distinction between positive and negative samples to ensure that related information of different modalities is more closely connected in the shared embedding space.
[0126] The corresponding features extracted by the image encoder and the text encoder are denoted as v z and v T , then the contrast loss includes the image-text contrast loss and the text-image contrast loss. The first loss function is as follows:
[0127]
[0128] in, is the first loss function, N B is the data size of a training batch, τ is the learnable coefficient, <·,·> represents the dot product, The image features of the i-th hidden danger image in the training data, is the text feature corresponding to the i-th hidden danger image in the training data.
[0129] After the image and text features are extracted, they are input into a language model based on an autoregressive mechanism (such as LLaMa) for inference. The text generation is trained through the autoregressive text prediction loss function. The third loss function is as follows:
[0130]
[0131] in, is the third loss function, is the mathematical expectation, represents all text features before the predicted position m, Represents the text generated given the previous and image features v z In the case of, generate the word at position m in the text The conditional probability of .
[0132] The autoregressive language model generates inference results based on visual features and text features. For example, by inputting an image of a pin, the model can predict the result ("Detection result: the pin is loose") through causal inference. This technical means can provide more intelligent decision-making support for operation and maintenance personnel, not only identifying current hidden dangers, but also reasoning and predicting potential abnormal situations.
[0133] In addition, the transmission line hidden danger detection model can also combine the images taken by drones and the data of pin sensors (such as sensor data of pin displacement, looseness, etc.) to further enrich the description information of hidden dangers. Through the fusion of multimodal data, the model can simultaneously use the visual features in the image and the text data provided by the pin sensor to obtain a more accurate and comprehensive description of hidden dangers. This can not only improve the accuracy of hidden danger identification, but also enhance the adaptability and robustness of the model, especially in the face of complex environments or changeable power lines, and can provide more comprehensive hidden danger detection and prediction support for operation and maintenance personnel.
[0134] Finally, after multimodal learning and reasoning, the system can accurately identify and classify the types of hidden dangers in the image, such as broken wires, hanging objects, loose pins, etc., and output text descriptions and location information of the hidden dangers in real time. All processing is completed on the drone side, reducing data transmission and calculation delays and improving the real-time performance and efficiency of the system.
[0135] Through the implementation of the above-mentioned technical features, the present invention can significantly improve the accuracy of hidden danger detection in harsh environments such as wind disasters, reduce the blind spots of traditional manual inspections, enhance the intelligence level of the system, and reduce computing costs and delays through end-side reasoning technology.
[0136] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.
[0137] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A method for detecting hidden dangers in a power transmission line, characterized in that: include: Acquire hidden danger images of the transmission line, wherein the hidden danger images include broken strand images, wire foreign matter images, and loose pin images; Acquire text data corresponding to the hidden danger image, wherein the text data includes a text description of a broken strand image, a text description of a wire foreign body image, and a text description of a loose pin image from a pin sensor; A transmission line hidden danger detection model is constructed based on hidden danger images and text data, wherein the transmission line hidden danger detection model includes a visual encoder, a text encoder, a classification head, and a language model; and a transmission line hidden danger detection model is constructed based on hidden danger images and text data; Acquire image data of power transmission lines taken by drones; Input the acquired image data into the visual encoder of the power transmission line hidden danger detection model, input the visual features output by the visual encoder into the language model, and output the text data corresponding to the acquired image data through the language model; Based on the output text data, determine whether there are hidden dangers in the transmission line and the corresponding hidden danger types and descriptions.
2. A method for detecting hidden dangers in power transmission lines according to claim 1, characterized in that: The construction of a transmission line hidden danger detection model based on hidden danger images and text data specifically includes: Preprocessing the hidden danger image, wherein the preprocessing includes image normalization, data enhancement, denoising and deblurring operations; Encoding the preprocessed hidden danger images through a visual encoder to extract the first visual features of each hidden danger image; Encoding the text data by a text encoder to extract the first text feature of each text description; The image encoder and the text encoder are trained according to the first visual features of each hidden danger image and the first text features of each text description by contrastive learning, and the parameters of the image encoder and the text encoder are updated by a first loss function; Input the extracted first visual feature and the first text feature into the classification head, output the probability of each feature matching, and update the parameters of the image encoder, the text encoder and the classification head through the second loss function according to the output probability; After the image encoder and the text encoder are trained, the preprocessed hidden danger images are encoded by the trained visual encoder to extract the second visual features of each hidden danger image; Encode the text data through the trained text encoder to extract the second text features of each text description; The language model receives the second visual feature extracted by the visual encoder, takes the second text feature as a target, updates the parameters of the language model based on the third loss function, and obtains a trained language model.
3. A method for detecting hidden dangers in power transmission lines according to claim 2, characterized in that: The encoding of the preprocessed hidden danger images by a visual encoder to extract the first visual features of each hidden danger image specifically includes: Divide the hidden danger image into a number of image patches of fixed sizes; Expand each image patch into a one-dimensional vector and encode its position; The one-dimensional vector of each image patch and its position encoding are input into a visual encoder, wherein the visual encoder adopts a network structure based on a self-attention mechanism and uses a multi-layer self-attention mechanism to capture the long-range dependencies between different regions of the image; The self-attention mechanism of the visual encoder is used to calculate the similarity between each patch and other patches, and the representation of the one-dimensional vector of each image patch is adjusted according to the similarity; The adjusted one-dimensional vector of image patches is further processed by a feed-forward neural network including a non-linear activation function and layer normalization; The one-dimensional vector features of the processed image patches are fused to obtain the first visual features of each hidden danger image.
4. A method for detecting hidden dangers in power transmission lines according to claim 2, characterized in that: The step of encoding the text data by a text encoder to extract the first text feature of each text description specifically includes: Perform word segmentation on the input text data and divide the text data into several sub-word units; Map each subword unit to the corresponding word vector representation; Input the word vector representation into a text encoder, wherein the text encoder is a language model BERT based on a bidirectional attention mechanism; The language model BERT is used to calculate the contextual dependencies of each subword unit, and the word vector representation of each subword is updated based on the contextual information. The updated sub-word representations are merged into the overall text feature representation of the text through a pooling operation or an aggregation strategy to obtain the first text feature of each text description.
5. A method for detecting hidden dangers in power transmission lines according to claim 2, characterized in that: The image encoder and the text encoder are trained according to the first visual features of each hidden danger image and the first text features of each text description by contrastive learning, specifically including: Pairing the first visual feature of the hidden danger image with the first text feature of the text data to form positive samples and negative samples, wherein the positive samples include pairs where the image and its corresponding text description are semantically matched, and the negative samples include pairs where there is no actual semantic match between the image and its corresponding text description; Calculate the image features v of positive and negative samples z With the text feature v T The image encoder and text encoder are optimized by the similarity and the first loss function.
6. A method for detecting hidden dangers in power transmission lines according to claim 5, characterized in that: The first loss function is: in, is the first loss function, N B is the data size of a training batch, τ is the learnable coefficient, <·,·> represents the dot product, The image features of the i-th hidden danger image in the training data, is the text feature corresponding to the i-th hidden danger image in the training data.
7. A method for detecting hidden dangers in power transmission lines according to claim 2, characterized in that: The first visual feature and the first text feature extracted are input into the classification head, the probability of each feature matching is output, and the image encoder, the text encoder and the classification head are updated with the second loss function according to the output probability, specifically including: The first visual feature of the hidden danger image extracted by the visual encoder and the first text description text features extracted by the text encoder Input a classification head, which is a binary classification network module, and is used to output the probability p of whether the first visual feature matches the first text feature. v ; The classification head processes the first visual feature v of the input z and the first text feature v T , output a probability p v , represents the matching degree between image and text in semantic space. If the image and text match semantically, the output probability is close to 1, and if they do not match, the output probability is close to 0; During the training process, the positive sample is a pair of images and their corresponding text descriptions that match, and its label is y=1, and the negative sample is a pair of images and text descriptions that do not match, and its label is y=0; According to the output matching probability p v The difference between the image and the actual label y is used to update the parameters of the image encoder, text encoder, and classification head through the second loss function.
8. A method for detecting hidden dangers in power transmission lines according to claim 7, characterized in that: The second loss function is: in, is the second loss function, is the mathematical expectation.
9. A method for detecting hidden dangers in power transmission lines according to claim 1, characterized in that: The language model is a language model LLaMa based on an autoregressive mechanism; The language model receives the second visual feature extracted by the visual encoder, takes the second text feature as a target, and updates the parameters of the language model based on the third loss function, specifically including: The second visual feature is input into the language model. Given the visual feature v z As input, the language model outputs a conditional probability distribution The conditional probability distribution Describes the probability of a word at each position in a text sequence, given the previously generated text and visual features v z In the case of, generate the word at position m in the text The conditional probability of A third loss function is calculated, and the language model is updated according to the calculated third loss function.
10. A method for detecting hidden dangers in power transmission lines according to claim 9, characterized in that: The third loss function is: in, is the third loss function, is the mathematical expectation, represents all text features before the predicted position m, Represents the text generated given the previous and image features v z In the case of, generate the word at position m in the text The conditional probability of .
Citation Information
Patent Citations
Power transmission line specific fault recognition system based on combination of natural language model and target detection algorithm
CN112395954A
Power transmission and transformation line hidden danger automatic safety identification system based on deep residual network
CN113378723A
Prompt learning method for modal interaction enhancement of visual language model
CN116503683A
Power line post-disaster inspection scene image understanding method based on LSTM
CN117576592A
Visual encoder training method and device, visual encoder description method and device, equipment and medium
CN117764043A
Cited By
Flying and hanging object identification method and device of lightweight open-set detection model, and medium
CN120747744A
Method and system for identifying and describing potential safety hazards of rail transit engineering construction
CN120913118A