An outdoor power transmission channel inspection intelligent detection method, system and device
Patent Information
- Application Number
- CN202510593730.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2045-05-09
AI Technical Summary
[0003]输电巡检场景中的隐患类别复杂多变,输电塔杆、绝缘子、导线、避雷器等目标往往需要分别设计和训练语义、实例、全景分割模型,导致无人机巡检时切换模型频繁,增加部署和切换成本;且在跨区域巡检时,不同线路、杆塔形态变化大,单一任务模型难以泛化
[0018] This invention proposes an intelligent detection method, system, and device for outdoor power transmission channel inspection. The method includes the following steps: generating structured text prompts based on power transmission inspection task parameters and environmental information; generating task-related text features using a text encoder; extracting visual features from the inspection image, aligning the semantic space of the visual features with the text features by comparing a loss function and a text-guided cross-modal attention mechanism; fusing geographic information system coordinates and meteorological data into the text features to generate an environmentally perceptible segmentation mask, and outputting the defect area detection results. Based on this intelligent detection method for outdoor power transmission channel inspection, an intelligent detection system and device for outdoor power transmission channel inspection are also proposed. This invention achieves high efficiency, intelligence, and automation in outdoor power transmission channel inspection by integrating image segmentation and multimodal fusion technologies.
Smart Images

Figure CN120635692B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power transmission channel inspection technology, and specifically relates to an intelligent detection method, system and equipment for outdoor power transmission channel inspection. Background Technology
[0002] The stable operation of outdoor power transmission channels is crucial for power supply. Power equipment bears a heavy power supply burden, and its location in complex environments, including strong winds, rain, snow, lightning strikes, and bird droppings, all pose challenges to its operation. Timely and effective detection of defects and potential hazards, ensuring the safe and reliable operation of outdoor power supply equipment under various conditions, and providing residents and businesses with normal and stable power supply are the important significance of power transmission line inspection work.
[0003] The types of hazards in power transmission line inspection scenarios are complex and varied. Targets such as transmission towers, insulators, conductors, and surge arresters often require separate design and training of semantic, instance, and panoramic segmentation models. This leads to frequent model switching during drone inspections, increasing deployment and switching costs. Furthermore, in cross-regional inspections, the morphology of different lines and towers varies greatly, making it difficult for single-task models to generalize. Power transmission line inspections not only rely on images but also require the integration of multi-source data such as tower numbers, line corridor geographic information, historical maintenance records, and weather conditions. Existing systems struggle to align on-site QR codes, geographic coordinates, text inspection logs, and visual features, resulting in low accuracy in identifying novel or unknown defects such as microcracks and creepage spots. In outdoor high-altitude inspection images, the contrast between large background areas such as the sky, mountains, and vegetation and local details such as metal connectors and steel cables is low. Conventional grid subsampling easily overlooks small cracks in insulators and corrosion points on hardware, causing subsequent segmentation to lose key defect areas and affecting the completeness of safety assessments. Summary of the Invention
[0004] To address the aforementioned technical issues, this invention proposes an intelligent detection method, system, and equipment for outdoor power transmission channel inspection. By integrating image segmentation and multimodal fusion technologies, it achieves high efficiency, intelligence, and automation in the inspection and detection of outdoor power transmission channels.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] An intelligent detection method for inspecting outdoor power transmission channels includes the following steps:
[0007] Structured text prompts are generated based on power transmission inspection task parameters and environmental information; task-related text features are generated using a text encoder.
[0008] Visual features of the inspection images are extracted, and the semantic spaces of the visual features and text features are aligned by comparing the loss function and the text-guided cross-modal attention mechanism.
[0009] By integrating geographic information system coordinates and meteorological data into the text features, an environmentally aware segmentation mask is generated, and the defect area detection results are output.
[0010] This invention also proposes an intelligent detection system for outdoor power transmission channel inspection, characterized by including a text processing module, an image processing module, and a detection output module;
[0011] The text processing module is used to generate structured text prompts based on power transmission inspection task parameters and environmental information; and to generate task-related text features using a text encoder.
[0012] The image processing module is used to extract visual features from the inspection images and align the semantic space of the visual features with that of the text features by comparing the loss function and the text-guided cross-modal attention mechanism.
[0013] The detection output module is used to fuse geographic information system coordinates and meteorological data into the text features, generate an environmentally aware segmentation mask, and output the defect area detection results.
[0014] This invention also proposes an intelligent inspection device for outdoor power transmission channels, comprising:
[0015] Memory, used to store computer programs;
[0016] A processor for implementing the steps of the method when executing the computer program.
[0017] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects:
[0018] This invention proposes an intelligent detection method, system, and device for outdoor power transmission channel inspection. The method includes the following steps: generating structured text prompts based on power transmission inspection task parameters and environmental information; generating task-related text features using a text encoder; extracting visual features from the inspection image, aligning the semantic space of the visual features with the text features by comparing a loss function and a text-guided cross-modal attention mechanism; fusing geographic information system coordinates and meteorological data into the text features to generate an environmentally perceptible segmentation mask, and outputting the defect area detection results. Based on this intelligent detection method for outdoor power transmission channel inspection, an intelligent detection system and device for outdoor power transmission channel inspection are also proposed. This invention achieves high efficiency, intelligence, and automation in outdoor power transmission channel inspection by integrating image segmentation and multimodal fusion technologies.
[0019] This invention integrates semantic segmentation, instance segmentation, and panoramic segmentation through a unified mask classification framework. It can complete multiple segmentation tasks for towers, conductors, insulators, and accessories using a single model, simplifying the system architecture, reducing hardware resource requirements, and improving inspection efficiency.
[0020] This invention utilizes a pre-trained text encoder to generate tower numbers, line segments, and defect type prompts, and performs cross-modal attention fusion with visual features. It makes full use of on-site inspection text descriptions, historical maintenance standards, and environmental text information to improve the ability to identify open-domain defects such as microcracks and corrosion.
[0021] This invention targets the perspective of high-altitude inspection, employing a balanced clustering strategy to generate irregularly distributed feature tokens, and combining local importance scoring for dynamic sampling and feature merging, ensuring that key details such as tower hardware and insulators can still be captured even in a single background of sky or vegetation. Attached Figure Description
[0022] Figure 1 This is a flowchart of an intelligent detection method for outdoor power transmission channel inspection proposed in Embodiment 1 of the present invention;
[0023] Figure 2 This is a schematic diagram of the architecture of an intelligent detection method for outdoor power transmission channel inspection proposed in Embodiment 1 of the present invention;
[0024] Figure 3 This is a schematic diagram of the unified mask classification and segmentation framework proposed in Embodiment 1 of the present invention;
[0025] Figure 4 This is a schematic diagram of image-text dual-modal fusion proposed in Embodiment 1 of the present invention;
[0026] Figure 5 This is a schematic diagram of the adaptive downsampling module proposed in Embodiment 1 of the present invention;
[0027] Figure 6 This is a schematic diagram of an intelligent detection system for outdoor power transmission channel inspection proposed in Embodiment 2 of the present invention;
[0028] Figure 7 This is a schematic diagram of an intelligent detection device for outdoor power transmission channel inspection proposed in Embodiment 3 of the present invention. Detailed Implementation
[0029] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure of the invention, components and arrangements of specific examples are described below. Furthermore, reference numerals and / or letters may be repeated in different examples. This repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. Descriptions of well-known components, processing techniques, and processes are omitted in this invention to avoid unnecessarily limiting the invention.
[0030] Example 1
[0031] Embodiment 1 of this invention proposes an intelligent detection method for outdoor power transmission channel inspection, which is used to solve the technical problems existing in the inspection of outdoor power transmission channels in the prior art.
[0032] Figure 1 This is a flowchart of an intelligent detection method for outdoor power transmission channel inspection proposed in Embodiment 1 of the present invention;
[0033] In step S100, a structured text prompt is generated based on the power transmission inspection task parameters and environmental information; and a text encoder is used to generate text features related to the task.
[0034] Figure 2 This is a schematic diagram of the architecture of an intelligent detection method for outdoor power transmission channel inspection proposed in Embodiment 1 of the present invention;
[0035] Structured text prompts include:
[0036] Inspection task type, tower number, and line segment parameters;
[0037] High-frequency keywords from the power industry terminology dictionary (such as "insufficient creepage distance" and "aging of silicone rubber in composite insulators") are embedded to enhance the semantic representation of professional terms.
[0038] It integrates real-time weather conditions and geographic coordinate information.
[0039] Figure 3 This is a schematic diagram of the unified mask classification and segmentation framework proposed in Embodiment 1 of the present invention; the process of generating task-related text features using a text encoder includes: inputting structured text prompts into the CLIP text encoder to generate task-related text features F. t Among them, text features F t Represented as: F t = CLIP-TextEncoder(“Detect {task type}, tower number: {number}, environment: {weather}”);
[0040] The task types include at least one of microcrack detection, insulator damage identification, and mud erosion area segmentation.
[0041] The CLIP text encoder is responsible for converting text into vector representations. Based on the Transformer architecture, it can handle long-distance dependencies and generate text vectors corresponding to image vectors.
[0042] Ensure F through linear projection t With visual feature F v Dimension matching (usually d=512 or d=768).
[0043] Example: Input prompt: "Detect microcracks, tower number: 0257, line section: 110kV line A3 section, environment: cloudy, high humidity" → Output F t It includes task type, location, and environmental semantics.
[0044] In step S110, visual features of the inspection image are extracted, and the semantic space of the visual features and text features is aligned by comparing the loss function and the text-guided cross-modal attention mechanism.
[0045] The process of extracting visual features from inspection images involves inputting the inspection images into a downsampling encoder to extract visual features. Figure 5 This is a schematic diagram of the adaptive downsampling module proposed in Embodiment 1 of the present invention;
[0046] The specific process includes: first, image preprocessing, which normalizes pixel values to the range of [0,1] or [-1,1] to accelerate model convergence; then, rotation, flipping, brightness / contrast adjustment, etc., are applied to enhance the model's generalization ability; finally, if the inspected image has a lot of noise, filtering is performed.
[0047] When designing the downsampling editor, a lightweight model suitable for real-time inspection scenarios is adopted. If the amount of data is small, ResNet / VGG pre-trained on ImageNet is used as the encoder, and the top layer is adjusted to adapt to the inspection task. Downsampling strategies and multi-scale feature fusion are set.
[0048] Introduce SE (Squeeze-and-Excitation) or CBAM modules to enhance the response of defect-related channels / spatial regions, use Batch Normalization to accelerate training, and Dropout to prevent overfitting.
[0049] The output is a three-dimensional tensor (C×H×W), which can be adjusted according to the task; anomaly detection is performed by comparing the feature distribution of normal samples.
[0050] The specific process of aligning the semantic spaces of the visual features and text features by comparing the loss function and the text-guided cross-modal attention mechanism includes:
[0051] Where S = CosSimilarity(F v ,F t );
[0052]
[0053] L mask =Focal(M pred M gt )+Dice(M pred M gt );
[0054] L class =CrossEntropy(C pred C gt );
[0055] Among them, F v For image features; F t For text features; CosSimilarity() represents cosine similarity;
[0056] L contrastive For contrast loss; i represents image; j represents text; M represents the number of samples in the batch, i.e., image-text pairs; τ represents the temperature coefficient; S ii S represents the similarity between positive sample pairs; ij Represents the elements of the similarity matrix;
[0057] L mask Represents mask loss; Focal(M) pred M gt Dice(M) is used to represent the degree of overlap between the real mask and the real class; pred M gt M is used to adjust the weights between the real mask and the real class. pred For a real mask; M gt For the real category;
[0058] L class Represents cross-entropy loss; C pred C represents the classification confidence level of the true mask. gt A true category index representing each target region.
[0059] In power transmission inspection scenarios, this module can simultaneously identify insulator cracks, hardware corrosion, and conductor wear in a single inference, while also outputting segmentation masks and defect categories, thus simplifying the inspection process.
[0060] In step S120, geographic information system coordinates and meteorological data are fused into the text features to generate an environmentally perceptive segmentation mask and output the defect area detection results.
[0061] By employing contrastive loss and cross-modal attention, visual features are aligned with text embeddings, enhancing the ability to identify unseen defects such as novel mud erosion and creep channels.
[0062] By combining Geographic Information System (GIS) coordinates and weather logs, environmental text (such as "cloudy" and "high humidity") is integrated into F t This improves the model's detection stability under changes in lighting and rain / fog obstruction.
[0063] The process of image and text fusion in power transmission inspection can be adaptively adjusted according to different line and environmental conditions to ensure robust identification of unknown or rare defects. Figure 4 This is a schematic diagram of image-text dual-modal fusion proposed in Embodiment 1 of the present invention;
[0064] To address the issues of monotonous backgrounds and sparse details in outdoor high-altitude inspection images, this invention proposes an irregular token generation and dynamic sampling strategy.
[0065] The visual features are used as tokens, and the K-center clustering algorithm is used to cluster the visual features to generate an irregular token set, and a neighborhood matrix A is constructed for each token;
[0066] The local influence between tokens is calculated using the domain matrix A.
[0067] The downsampling center is selected based on the importance score of each token, and the features of neighboring tokens are merged to achieve efficient preservation of small-sized cracks and corrosion spots.
[0068] The process of calculating the local influence between tokens using the domain matrix A is as follows:
[0069]
[0070] Among them, X l X is the output of the l-th decoded block. l-1 This is the output of the (l-1)th decoded block; A l Q is the neighborhood matrix in the l-th decoded block; l K is the query matrix for the l-th decoded block; l The key value of the l-th decoded block is V. l Let be the Value matrix of the l-th decoded block; x and y are different visual features in the token set.
[0071] In actual inspections, this method can focus on vulnerable components such as the tower top, insulator strings, and vibration dampers, so that minor defects are not diluted by the large area of sky or vegetation background.
[0072] The output defect area detection results include: insulator cracks, hardware corrosion, and conductor wear; it also outputs the segmentation mask and defect category.
[0073] The present invention provides an intelligent detection method for outdoor power transmission channel inspection, which integrates image segmentation and multimodal fusion technology to achieve high efficiency, intelligence and automation of outdoor power transmission channel inspection.
[0074] The present invention, in embodiment 1, proposes an intelligent detection method for outdoor power transmission channel inspection. By using a unified mask classification framework, semantic segmentation, instance segmentation, and panoramic segmentation are integrated into one. A single model can complete multiple segmentation tasks for towers, conductors, insulators, and accessories, simplifying the system architecture, reducing hardware resource requirements, and improving inspection efficiency.
[0075] The present invention, in embodiment 1, proposes an intelligent detection method for outdoor power transmission channel inspection. This method utilizes a pre-trained text encoder to generate tower numbers, line segments, and defect type prompts, and performs cross-modal attention fusion with visual features. It makes full use of on-site inspection text descriptions, historical maintenance standards, and environmental text information to improve the ability to identify open-domain defects such as micro-cracks and corrosion.
[0076] The present invention provides an intelligent detection method for outdoor power transmission channel inspection. For high-altitude inspection, a balanced clustering strategy is used to generate irregularly distributed feature tokens. Combined with local importance scoring, dynamic sampling and feature merging are performed to ensure that key details such as tower hardware and insulators can still be captured even in a single background of sky or vegetation area.
[0077] Example 2
[0078] Based on the intelligent detection method for outdoor power transmission channel inspection proposed in Embodiment 1 of the present invention, Embodiment 2 of the present invention also proposes an intelligent detection system for outdoor power transmission channel inspection. Figure 6 This is a schematic diagram of an intelligent detection system for outdoor power transmission channel inspection proposed in Embodiment 2 of the present invention; the system includes a text processing module, an image processing module, and a detection output module;
[0079] The text processing module is used to generate structured text prompts based on power transmission inspection task parameters and environmental information; and to generate task-related text features using a text encoder.
[0080] The image processing module is used to extract visual features from the inspection images and align the semantic space of the visual features with that of the text features by comparing the loss function and the text-guided cross-modal attention mechanism.
[0081] The detection output module is used to fuse geographic information system coordinates and meteorological data into the text features, generate an environmentally aware segmentation mask, and output the defect area detection results.
[0082] During the execution of the text processing module,
[0083] Structured text prompts include:
[0084] Inspection task type, tower number, and line segment parameters;
[0085] High-frequency keywords embedded in the terminology dictionary of the power industry;
[0086] It integrates real-time weather conditions and geographic coordinate information.
[0087] The process of generating task-relevant text features using a text encoder includes: inputting structured text prompts into the CLIP text encoder to generate task-relevant text features; wherein, the text features are represented as:
[0088] F t = CLIP-TextEncoder(“Detect {task type}, tower number: {number}, environment: {weather}”);
[0089] The task types include at least one of microcrack detection, insulator damage identification, and mud erosion area segmentation.
[0090] In the image processing module, the process of extracting visual features from the inspection image includes: inputting the inspection image into a downsampling encoder to extract visual features.
[0091] The specific process of aligning the semantic spaces of the visual features and text features by comparing the loss function and the text-guided cross-modal attention mechanism includes:
[0092] Where S = CosSimilarity(F v ,F t );
[0093]
[0094] L mask =Focal(M pred M gt )+Dice(M pred M gt );
[0095] L class =CrossEntropy(C pred C gt );
[0096] Among them, F v For image features; F t For text features; CosSimilarity() represents cosine similarity;
[0097] L contrastive For contrast loss; i represents image; j represents text; M represents the number of samples in the batch, i.e., image-text pairs; τ represents the temperature coefficient; S ii S represents the similarity between positive sample pairs; ij Represents the elements of the similarity matrix;
[0098] L mask Represents mask loss; Focal(M) pred M gt Dice(M) is used to represent the degree of overlap between the real mask and the real class; pred M gt M is used to adjust the weights between the real mask and the real class. pred For a real mask; M gt For the real category;
[0099] L class Represents cross-entropy loss; C pred C represents the classification confidence level of the true mask. gt A true category index representing each target region.
[0100] In the detection output module, the visual features are used as tokens, and the K-center clustering algorithm is used to cluster the visual features to generate an irregular token set, and a neighborhood matrix for each token is constructed.
[0101] Calculate the local influence between tokens using the domain matrix;
[0102] The downsampling center is selected based on the importance score of each token, and the features of neighboring tokens are merged to achieve efficient preservation of small-sized cracks and corrosion spots.
[0103] The process of calculating the local influence between tokens using the aforementioned domain matrix is as follows:
[0104]
[0105] Among them, X l X is the output of the l-th decoded block. l-1 This is the output of the (l-1)th decoded block; A l Q is the neighborhood matrix in the l-th decoded block; l K is the query matrix for the l-th decoded block; lThe key value of the l-th decoded block is V. l Let be the Value matrix of the l-th decoded block; x and y are different visual features in the token set.
[0106] The output defect area detection results include: insulator cracks, hardware corrosion, and conductor wear; it also outputs the segmentation mask and defect category.
[0107] The intelligent detection system for outdoor power transmission channel inspection proposed in Embodiment 2 of this invention achieves high efficiency, intelligence and automation of outdoor power transmission channel inspection by integrating image segmentation and multimodal fusion technology.
[0108] The intelligent detection system for outdoor power transmission channel inspection proposed in Embodiment 2 of this invention integrates semantic segmentation, instance segmentation and panoramic segmentation through a unified mask classification framework. It can complete multiple segmentation tasks of towers, conductors, insulators and accessories using a single model, simplifying the system architecture, reducing hardware resource requirements and improving inspection efficiency.
[0109] The present invention, embodiment 2, proposes an intelligent detection system for outdoor power transmission channel inspection. It utilizes a pre-trained text encoder to generate tower numbers, line segments, and defect type prompts, and performs cross-modal attention fusion with visual features. It makes full use of on-site inspection text descriptions, historical maintenance standards, and environmental text information to improve the ability to identify open domain defects such as micro-cracks and corrosion.
[0110] The present invention, in embodiment 2, proposes an intelligent detection system for outdoor power transmission channel inspection. For high-altitude inspection perspective, it adopts a balanced clustering strategy to generate irregularly distributed feature tokens, and combines local importance scoring for dynamic sampling and feature merging to ensure that key details such as tower hardware and insulators can still be captured even in a single background of sky or vegetation area.
[0111] Example 3
[0112] The present invention also proposes a device, Figure 7 This is a schematic diagram of an intelligent detection device for outdoor power transmission channel inspection proposed in Embodiment 3 of the present invention, comprising:
[0113] Memory, used to store computer programs;
[0114] When a processor executes the computer program, the method steps are as follows:
[0115] In step S100, a structured text prompt is generated based on the power transmission inspection task parameters and environmental information; and a text encoder is used to generate text features related to the task.
[0116] In step S110, visual features of the inspection image are extracted, and the semantic space of the visual features and text features is aligned by comparing the loss function and the text-guided cross-modal attention mechanism.
[0117] In step S120, geographic information system coordinates and meteorological data are fused into the text features to generate an environmentally perceptive segmentation mask and output the defect area detection results.
[0118] It should be noted that the present invention also provides an electronic device, including: a communication interface capable of interacting with other devices such as network devices; and a processor connected to the communication interface to enable information interaction with other devices, used to execute a smart detection method for outdoor power transmission channel inspection provided by one or more of the above technical solutions when running a computer program, wherein the computer program is stored in a memory. In practical applications, the various components of the electronic device are coupled together through a bus system. It is understood that the bus system is used to realize the connection and communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus. The memory in the embodiments of this application is used to store various types of data to support the operation of the electronic device. Examples of this data include any computer program used to operate on the electronic device. It is understood that the memory can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache.By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), SyncLink Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM). The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memory. The methods disclosed in the embodiments of this application can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by integrated logic circuits in the processor hardware or by instructions in software. The processor can be a general-purpose processor, a DSP (Digital Signal Processing, i.e., a chip capable of implementing digital signal processing technology), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules can be located in a storage medium, which is located in memory. The processor reads the program from the memory and, in conjunction with its hardware, completes the steps of the aforementioned method. When the processor executes the program, it implements the corresponding processes in the various methods of the embodiments of this application; for simplicity, these will not be elaborated further here.
[0119] The description of the relevant parts of the intelligent detection device for outdoor power transmission channel inspection provided in Embodiment 3 of this application can be found in the detailed description of the corresponding parts of the intelligent detection method for outdoor power transmission channel inspection provided in Embodiment 1 of this application, and will not be repeated here.
[0120] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that the elements inherent in a process, method, article, or apparatus that includes a list of elements are included. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Additionally, portions of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.
[0121] While specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art can make other modifications or variations based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A smart detection method for inspecting outdoor power transmission channels, characterized in that, Includes the following steps: Structured text prompts are generated based on power transmission inspection task parameters and environmental information; task-related text features are generated using a text encoder. Visual features of the inspection images are extracted, and the semantic spaces of the visual features and text features are aligned by comparing the loss function and the text-guided cross-modal attention mechanism. By integrating geographic information system coordinates and meteorological data into the text features, an environmentally perceptual segmentation mask is generated, and the defect area detection results are output. The method further includes: using the visual features as tokens, clustering the visual features using the K-center clustering algorithm to generate an irregular token set, and constructing a neighborhood matrix for each token; using the neighborhood matrix to calculate the local influence between tokens; selecting a downsampling center based on the importance score of each token, merging the features of neighboring tokens, and achieving efficient preservation of small-sized cracks and corrosion spots; The process of calculating the local influence between tokens using the aforementioned domain matrix is as follows: in, For the first The output of each decoded block, For the first The output of each decoded block; For the first The domain matrix in each decoding block; For the first The query matrix of each decoded block; For the first The average key value of each decoded block; For the first Value matrix of each decoded block; and These are all different visual features within the token set.
2. The intelligent detection method for outdoor power transmission channel inspection according to claim 1, characterized in that, The structured text prompts include: Inspection task type, tower number, and line segment parameters; High-frequency keywords embedded in the terminology dictionary of the power industry; It integrates real-time weather conditions and geographic coordinate information.
3. The intelligent detection method for outdoor power transmission channel inspection according to claim 2, characterized in that, The process of generating task-relevant text features using a text encoder includes: inputting structured text prompts into the CLIP text encoder to generate task-relevant text features; wherein, the text features are represented as: =CLIP-TextEncoder("Detect {task type}, tower number: {number}, environment: {weather}"); The task types include at least one of microcrack detection, insulator damage identification, and mud erosion area segmentation.
4. The intelligent detection method for outdoor power transmission channel inspection according to claim 1, characterized in that, The process of extracting visual features from inspection images includes: inputting the inspection image into a downsampling encoder to extract visual features.
5. The intelligent detection method for outdoor power transmission channel inspection according to claim 1, characterized in that, The specific process of aligning the semantic spaces of the visual features and text features by comparing the loss function and the text-guided cross-modal attention mechanism includes: in, ; ; ; ; in, Image features; Text features; Indicates cosine similarity; To compare the losses; Representative image; Represents text; This represents the number of samples processed in a batch, i.e., image-text pairs; Represents the temperature coefficient; Represents the similarity between positive sample pairs; Represents the elements of the similarity matrix; Represents mask loss; Used to indicate the degree of overlap between the real mask and the real category; Used to adjust the weights between the real mask and the real category; For real masks; For the real category; Represents cross-entropy loss; The classification confidence level representing the true mask; A true category index representing each target region.
6. The intelligent detection method for outdoor power transmission channel inspection according to claim 1, characterized in that, The output defect area detection results include: insulator cracks, hardware corrosion, and conductor wear; it also outputs the segmentation mask and defect category.
7. An intelligent detection system for outdoor power transmission channel inspection, used to execute the intelligent detection method for outdoor power transmission channel inspection as described in any one of claims 1 to 6, characterized in that, It includes a text processing module, an image processing module, and a detection output module; The text processing module is used to generate structured text prompts based on power transmission inspection task parameters and environmental information; and to generate task-related text features using a text encoder. The image processing module is used to extract visual features from the inspection images and align the semantic space of the visual features with that of the text features by comparing the loss function and the text-guided cross-modal attention mechanism. The detection output module is used to fuse geographic information system coordinates and meteorological data into the text features, generate an environmentally aware segmentation mask, and output the defect area detection results.
8. An intelligent inspection device for outdoor power transmission channels, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Visual detection method for suspension state defects of high-speed rail overhead line system in open environment
CN119941640A