Outdoor power transmission channel inspection intelligent detection method, system and equipment
Through image segmentation and multimodal fusion technology, the text encoder and environmental data are used to align visual features and generate segmentation masks, which solves the problem of identifying various defects in outdoor power transmission channel inspections and achieves efficient and intelligent detection effects.
Patent Information
- Application Number
- CN202510593730.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-05-09
AI Technical Summary
Existing technologies make it difficult to efficiently and intelligently identify multiple defects during outdoor transmission channel inspections, especially in complex environments where the accuracy of identifying tiny cracks and corrosion spots is low. Existing systems also find it difficult to uniformly process multi-source data and background interference.
Image segmentation and multimodal fusion technology are used to generate task-related features through a text encoder. Combined with geographic information and meteorological data, the system aligns visual features with text features, generates environment-aware segmentation masks, and outputs defect area detection results.
It achieves efficient, intelligent and automated inspection of outdoor power transmission channels, simplifies the system architecture, improves the ability to identify minor defects, reduces hardware resource requirements, and can still accurately detect key details in complex backgrounds.
Smart Images

Figure CN120635692A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power transmission channel inspection, and in particular relates to an intelligent detection method, system and equipment for outdoor power transmission channel inspection. Background Art
[0002] The stable operation of outdoor power transmission channels is crucial to power supply. Power equipment carries a heavy load and is located in complex environments. High winds, rain, snow, lightning strikes, bird droppings, and other conditions can pose operational challenges. Transmission line inspections are crucial for timely and effective detection of defects and hidden dangers, ensuring the safe and reliable operation of outdoor power supply equipment in various environments and providing regular, stable electricity to residents and businesses.
[0003] The types of hidden dangers in power transmission inspection scenarios are complex and varied. Semantic, instance-based, and panoramic segmentation models are often designed and trained separately for transmission towers, insulators, conductors, and lightning arresters. This results in frequent model switching during drone inspections, increasing deployment and switching costs. Furthermore, during cross-regional inspections, the morphology of different lines and towers varies significantly, making single-task models difficult to generalize. Transmission line inspections rely not only on images but also on multiple sources of data, including tower numbers, geographic information about line corridors, historical maintenance records, and weather conditions. Existing systems struggle to align on-site QR codes, geographic coordinates, and text-based inspection logs with visual features, resulting in low accuracy in identifying new or unknown defects such as microcracks and creepage spots. In outdoor high-altitude inspection images, the contrast between large background areas like the sky, mountains, and vegetation and local details like metal connectors and cables is low. Conventional grid downsampling can easily overlook small cracks on insulators and corrosion points on hardware, leading to missed critical defect areas in subsequent segmentation and compromising the integrity of safety assessments. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes an intelligent detection method, system and equipment for outdoor power transmission channel inspection. By integrating image segmentation and multimodal fusion technology, the inspection and detection of outdoor power transmission channels are made efficient, intelligent and automated.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] An intelligent detection method for outdoor power transmission channel inspection includes the following steps:
[0007] Generate structured text prompts based on transmission inspection task parameters and environmental information; use text encoders to generate task-related text features;
[0008] Extract visual features from inspection images and align the semantic spaces of these visual features with those of text features using a contrastive loss function and a text-guided cross-modal attention mechanism.
[0009] The geographic information system coordinates and meteorological data are integrated with the text features to generate an environment-aware segmentation mask and output defect area detection results.
[0010] The present invention also proposes an outdoor power transmission channel inspection intelligent detection system, which is characterized by comprising a text processing module, an image processing module and a detection output module;
[0011] The text processing module is used to generate structured text prompts based on the power transmission inspection task parameters and environmental information; and use a text encoder to generate text features related to the task;
[0012] The image processing module is used to extract visual features of the inspection image and align the semantic space of the visual features with the text features through a contrast loss function and a text-guided cross-modal attention mechanism;
[0013] The detection output module is used to fuse geographic information system coordinates and meteorological data with the text features, generate an environment-aware segmentation mask, and output defect area detection results.
[0014] The present invention also proposes an outdoor power transmission channel inspection intelligent detection device, comprising:
[0015] memory for storing computer programs;
[0016] A processor is configured to implement the method steps described when executing the computer program.
[0017] The effects provided in the summary of the invention are only the effects of the embodiments, not all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects:
[0018] The present invention proposes an intelligent inspection method, system, and equipment for outdoor power transmission channel inspections. The method includes the following steps: generating structured text prompts based on power transmission inspection task parameters and environmental information; generating task-related text features using a text encoder; extracting visual features from inspection images, and aligning the semantic space of the visual features with the text features through a contrast loss function and a text-guided cross-modal attention mechanism; fusing geographic information system coordinates and meteorological data with the text features to generate an environmentally aware segmentation mask and output defect area detection results. Based on an intelligent inspection method for outdoor power transmission channel inspections, an intelligent inspection system and equipment for outdoor power transmission channel inspections are also proposed. The present invention achieves efficient, intelligent, and automated inspection and detection of outdoor power transmission channels by integrating image segmentation and multimodal fusion technologies.
[0019] The present invention integrates semantic segmentation, instance segmentation and panoramic segmentation through a unified mask classification framework. A single model can be used to complete multiple segmentation tasks of poles, conductors, insulators and accessories, simplifying the system architecture, reducing hardware resource requirements and improving inspection efficiency.
[0020] The present invention uses a pre-trained text encoder to generate tower numbers, line sections, and defect type prompts, and performs cross-modal attention fusion with visual features, making full use of on-site inspection text descriptions, historical maintenance standards, and environmental text information to improve the recognition ability of open domain defects such as tiny cracks and corrosion.
[0021] Aiming at the high-altitude inspection perspective, the present invention adopts a balanced clustering strategy to generate irregularly distributed feature tokens, and combines local importance scores for dynamic sampling and feature merging, ensuring that key details such as tower hardware and insulators can still be captured in areas with a single background of sky or vegetation. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a flow chart of an intelligent detection method for outdoor power transmission channel inspection proposed in Example 1 of the present invention;
[0023] Figure 2 This is a schematic diagram of the architecture of an intelligent detection method for outdoor power transmission channel inspection proposed in Example 1 of the present invention;
[0024] Figure 3 This is a schematic diagram of the unified mask classification and segmentation framework proposed in Example 1 of the present invention;
[0025] Figure 4 This is a schematic diagram of the image-text bimodal fusion proposed in Example 1 of the present invention;
[0026] Figure 5 This is a schematic diagram of the adaptive downsampling module proposed in Example 1 of the present invention;
[0027] Figure 6 This is a schematic diagram of an outdoor power transmission channel inspection intelligent detection system proposed in Example 2 of the present invention;
[0028] Figure 7 This is a schematic diagram of an intelligent detection device for inspecting outdoor power transmission channels proposed in Example 3 of the present invention. DETAILED DESCRIPTION
[0029] In order to clearly illustrate the technical features of this solution, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings. The disclosure below provides many different embodiments or examples for realizing different structures of the present invention. In order to simplify the disclosure of the present invention, the components and settings of specific examples are described below. In addition, the present invention may repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplicity and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed. It should be noted that the components illustrated in the accompanying drawings are not necessarily drawn to scale. The present invention omits descriptions of well-known components and processing technologies and processes to avoid unnecessary limitations on the present invention.
[0030] Example 1
[0031] Embodiment 1 of the present invention proposes an intelligent detection method for outdoor power transmission channel inspection, which is used to solve the technical problems existing in the outdoor power transmission channel inspection in the prior art.
[0032] Figure 1 This is a flow chart of an intelligent detection method for outdoor power transmission channel inspection proposed in Example 1 of the present invention;
[0033] In step S100, a structured text prompt is generated according to the power transmission inspection task parameters and environmental information; and a text encoder is used to generate text features related to the task.
[0034] Figure 2 This is a schematic diagram of the architecture of an intelligent detection method for outdoor power transmission channel inspection proposed in Example 1 of the present invention;
[0035] Structured text prompts include:
[0036] Inspection task type, tower number, and line section parameters;
[0037] Embed high-frequency keywords from the power field terminology dictionary (such as "insufficient creepage distance" and "aging of composite insulator silicone rubber") to enhance the semantic representation of professional terms;
[0038] Integrates real-time weather conditions and geographic coordinate information.
[0039] Figure 3 Schematic diagram of the unified mask classification and segmentation framework proposed in Example 1 of the present invention; the process of using the text encoder to generate task-related text features includes: inputting the structured text prompt into the CLIP text encoder to generate task-related text features F t ; Among them, the text feature F t Expressed as: F t = CLIP-TextEncoder("Detect {task type}, tower number: {number}, environment: {weather}");
[0040] Among them, the task types include at least one of microcrack detection, insulator damage identification, and mud erosion area segmentation.
[0041] The CLIP text encoder is responsible for converting text into vector representations. Based on the Transformer architecture, it can handle long-distance dependencies and generate text vectors corresponding to image vectors.
[0042] Ensure that F t With visual features F v Dimensionality matches (usually d=512 or d=768).
[0043] Example: Input prompt: "Detect microcracks, tower number: 0257, line section: 110kV line A3 section, environment: cloudy and high humidity" → output F t Contains task type, location, and environment semantics.
[0044] In step S110 , visual features of the inspection image are extracted, and the semantic spaces of the visual features and text features are aligned through a contrast loss function and a text-guided cross-modal attention mechanism.
[0045] The process of extracting visual features of the inspection image is to input the inspection image into a downsampling encoder to extract visual features. Figure 5 This is a schematic diagram of the adaptive downsampling module proposed in Example 1 of the present invention;
[0046] The specific process includes: first preprocessing the image, that is, normalizing the pixel values to the range of [0, 1] or [-1, 1] to accelerate model convergence; then applying rotation, flipping, brightness / contrast adjustment, etc. to enhance the generalization ability of the model; finally, if the inspection image has a lot of noise, filtering is performed.
[0047] When designing the downsampling editor, adopt a lightweight model suitable for real-time inspection scenarios. If the amount of data is small, use ResNet / VGG pre-trained on ImageNet as the encoder and adjust the top layer to suit the inspection task. Set the downsampling strategy and multi-scale feature fusion.
[0048] Introduce the SE (Squeeze-and-Excitation) or CBAM module to enhance the response of defect-related channels / spatial regions, use Batch Normalization to accelerate training, and use Dropout to prevent overfitting.
[0049] The output is a three-dimensional tensor (C×H×W), which can be adjusted according to the task; anomaly detection is performed by comparing the feature distribution of normal samples.
[0050] The specific process of aligning the semantic space of the visual features and text features by contrasting the loss function and the text-guided cross-modal attention mechanism includes:
[0051] Where S = CosSimilarity(F v ,F t );
[0052]
[0053] L mask =Focal(M pred ,M gt )+Dice(M pred ,M gt );
[0054] L class =CrossEntropy(C pred ,C gt );
[0055] Among them, F v is the image feature; F t is the text feature; CosSimilarity() represents cosine similarity;
[0056] L contrastive is the contrast loss; i represents image; j represents text; M represents the number of batch samples, i.e., image-text pairs; τ represents the temperature coefficient; S ii Represents the similarity of the positive sample pair; S ij Represents the similarity matrix element;
[0057] L mask Represents mask loss; Focal(M pred ,M gt ) is used to indicate the degree of overlap between the true mask and the true category; Dice (M pred ,M gt ) is used to adjust the weight between the true mask and the true category; M pred is the real mask; M gt is the real category;
[0058] L class represents the cross entropy loss; C pred Represents the classification confidence of the true mask; C gt Represents the true category index of each target region.
[0059] In power transmission inspection scenarios, this module can simultaneously identify insulator cracks, hardware corrosion, and conductor wear in a single inference, while outputting segmentation masks and defect categories to simplify the inspection process.
[0060] In step S120 , the geographic information system coordinates and meteorological data are integrated with the text features to generate an environment-aware segmentation mask and output a defect area detection result.
[0061] By using contrastive loss and cross-modal attention, visual features are aligned with text embeddings to enhance the recognition of unseen defects (such as new types of mud corrosion and creepage channels).
[0062] Combine Geographic Information System (GIS) coordinates and weather logs to incorporate environmental text (such as "cloudy", "high humidity") into F t , improving the model's detection stability under lighting changes and rain and fog.
[0063] The image and text fusion process can be adaptively adjusted according to different line and environmental conditions during transmission inspection to ensure robust recognition of unknown or rare defects. Figure 4 This is a schematic diagram of the image-text bimodal fusion proposed in Example 1 of the present invention;
[0064] To address the problems of single background and sparse details in outdoor high-altitude inspection images, the present invention designs an irregular token generation and dynamic sampling strategy.
[0065] The visual features are used as tokens, a K-center clustering algorithm is used to cluster the visual features to generate an irregular token set, and a domain matrix A of each token is constructed;
[0066] Calculate the local influence between tokens using the domain matrix A;
[0067] The downsampling center is selected according to the importance score of each token, and the neighborhood token features are merged to achieve efficient retention of small-sized cracks and corrosion spots.
[0068] The process of calculating the local influence between tokens using the domain matrix A is as follows:
[0069]
[0070] Among them, X l is the output of the lth decoding block, X l-1 is the output of the l-1th decoding block; A l is the neighborhood matrix in the lth decoding block; Q l is the query matrix of the lth decoding block; K l is the key value of the lth decoding block; V l is the Value matrix of the lth decoding block; x and y are different visual features in the token set.
[0071] In actual inspections, the process implemented by this method can focus on vulnerable parts such as tower tops, insulator strings, and shock-absorbing hammers, so that subtle defects are not diluted by the large sky or vegetation background.
[0072] The output defect area detection results include: insulator cracks, hardware corrosion and wire wear; segmentation mask and defect category are also output.
[0073] An intelligent inspection method for outdoor power transmission channels proposed in Example 1 of the present invention realizes efficient, intelligent and automated inspection of outdoor power transmission channels by integrating image segmentation and multimodal fusion technology.
[0074] An intelligent detection method for outdoor power transmission channel inspection proposed in Example 1 of the present invention integrates semantic segmentation, instance segmentation and panoramic segmentation through a unified mask classification framework. A single model can be used to complete multiple segmentation tasks of towers, conductors, insulators and accessories, simplifying the system architecture, reducing hardware resource requirements, and improving inspection efficiency.
[0075] Embodiment 1 of the present invention proposes an intelligent detection method for outdoor power transmission channel inspection, which uses a pre-trained text encoder to generate tower numbers, line sections, and defect type prompts, and performs cross-modal attention fusion with visual features, making full use of on-site inspection text descriptions, historical maintenance standards, and environmental text information to improve the recognition ability of open domain defects such as tiny cracks and corrosion.
[0076] Example 1 of the present invention proposes an intelligent detection method for outdoor power transmission channel inspection. For high-altitude inspection perspectives, a balanced clustering strategy is used to generate irregularly distributed feature tokens, and dynamic sampling and feature merging are combined with local importance scores to ensure that key details such as tower hardware and insulators can still be captured in areas with a single background of sky or vegetation.
[0077] Example 2
[0078] Based on the outdoor power transmission channel inspection intelligent detection method proposed in Example 1 of the present invention, Example 2 of the present invention further proposes an outdoor power transmission channel inspection intelligent detection system. Figure 6 This is a schematic diagram of an outdoor power transmission channel inspection intelligent detection system proposed in Example 2 of the present invention; the system includes a text processing module, an image processing module, and a detection output module;
[0079] The text processing module is used to generate structured text prompts based on the transmission inspection task parameters and environmental information; and use the text encoder to generate text features related to the task;
[0080] The image processing module is used to extract visual features of inspection images and align the semantic space of the visual features with the text features through a contrast loss function and a text-guided cross-modal attention mechanism.
[0081] The detection output module is used to fuse geographic information system coordinates and meteorological data into the text features, generate an environment-aware segmentation mask, and output defect area detection results.
[0082] During the execution of the text processing module,
[0083] Structured text prompts include:
[0084] Inspection task type, tower number, and line section parameters;
[0085] High-frequency keywords embedded in the terminology dictionary of the power sector;
[0086] Integrates real-time weather conditions and geographic coordinate information.
[0087] The process of using a text encoder to generate task-related text features includes: inputting a structured text prompt into the CLIP text encoder to generate task-related text features; wherein the text features are represented as:
[0088] F t = CLIP-TextEncoder("Detect {task type}, tower number: {number}, environment: {weather}");
[0089] Among them, the task types include at least one of microcrack detection, insulator damage identification, and mud erosion area segmentation.
[0090] In the image processing module, the process of extracting visual features of the inspection image includes: inputting the inspection image into a downsampling encoder to extract visual features.
[0091] The specific process of aligning the semantic space of the visual features and text features by contrasting the loss function and the text-guided cross-modal attention mechanism includes:
[0092] Where S = CosSimilarity(F v ,F t );
[0093]
[0094] L mask =Focal(M pred ,M gt )+Dice(M pred ,M gt );
[0095] L class =CrossEntropy(C pred ,C gt );
[0096] Among them, F v is the image feature; F t is the text feature; CosSimilarity() represents cosine similarity;
[0097] L contrastive is the contrast loss; i represents image; j represents text; M represents the number of batch samples, i.e., image-text pairs; τ represents the temperature coefficient; S ii Represents the similarity of the positive sample pair; S ij Represents the similarity matrix element;
[0098] L mask Represents mask loss; Focal(M pred ,M gt ) is used to indicate the degree of overlap between the true mask and the true category; Dice (M pred ,M gt ) is used to adjust the weight between the true mask and the true category; M pred is the real mask; M gt is the real category;
[0099] L class represents the cross entropy loss; C pred Represents the classification confidence of the true mask; C gt Represents the true category index of each target region.
[0100] In the detection output module, the visual features are used as tokens, and the K-center clustering algorithm is used to cluster the visual features to generate an irregular token set, and a domain matrix for each token is constructed;
[0101] Calculate the local influence between tokens using the domain matrix;
[0102] The downsampling center is selected according to the importance score of each token, and the neighborhood token features are merged to achieve efficient retention of small-sized cracks and corrosion spots.
[0103] The process of calculating the local influence between tokens using the domain matrix is as follows:
[0104]
[0105] Among them, X l is the output of the lth decoding block, X l-1 is the output of the l-1th decoding block; A l is the neighborhood matrix in the lth decoding block; Q l is the query matrix of the lth decoding block; K lis the key value of the lth decoding block; V l is the Value matrix of the lth decoding block; x and y are different visual features in the token set.
[0106] The output defect area detection results include: insulator cracks, hardware corrosion and wire wear; segmentation mask and defect category are also output.
[0107] An intelligent inspection and detection system for outdoor power transmission channels proposed in Example 2 of the present invention realizes efficient, intelligent and automated inspection and detection of outdoor power transmission channels by integrating image segmentation and multimodal fusion technology.
[0108] An intelligent detection system for outdoor power transmission channel inspection proposed in Example 2 of the present invention integrates semantic segmentation, instance segmentation and panoramic segmentation through a unified mask classification framework. A single model can be used to complete multiple segmentation tasks of poles, conductors, insulators and accessories, simplifying the system architecture, reducing hardware resource requirements and improving inspection efficiency.
[0109] An intelligent detection system for outdoor power transmission channel inspection proposed in Example 2 of the present invention uses a pre-trained text encoder to generate tower numbers, line sections, and defect type prompts, and performs cross-modal attention fusion with visual features, making full use of on-site inspection text descriptions, historical maintenance standards, and environmental text information to improve the ability to recognize open domain defects such as tiny cracks and corrosion.
[0110] Example 2 of the present invention proposes an intelligent detection system for outdoor power transmission channel inspection. It uses a balanced clustering strategy to generate irregularly distributed feature tokens for high-altitude inspection perspectives, and combines local importance scores for dynamic sampling and feature merging to ensure that key details such as tower hardware and insulators can still be captured in areas with a single background of sky or vegetation.
[0111] Example 3
[0112] The present invention also proposes a device, Figure 7 This is a schematic diagram of an outdoor power transmission channel inspection intelligent detection device proposed in Example 3 of the present invention, including:
[0113] memory for storing computer programs;
[0114] The processor is used to implement the following method steps when executing the computer program:
[0115] In step S100, a structured text prompt is generated according to the power transmission inspection task parameters and environmental information; and a text encoder is used to generate text features related to the task.
[0116] In step S110 , visual features of the inspection image are extracted, and the semantic spaces of the visual features and text features are aligned through a contrast loss function and a text-guided cross-modal attention mechanism.
[0117] In step S120 , the geographic information system coordinates and meteorological data are integrated with the text features to generate an environment-aware segmentation mask and output a defect area detection result.
[0118] It should be noted that the technical solution of the present invention also provides an electronic device, including: a communication interface capable of exchanging information with other devices such as network devices; a processor connected to the communication interface to realize information exchange with other devices, and used to execute an outdoor power transmission channel inspection intelligent detection method provided by one or more of the above technical solutions when running a computer program, and the computer program is stored on a memory. Of course, in actual application, the various components in the electronic device are coupled together through a bus system. It can be understood that the bus system is used to realize connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus and a status signal bus. The memory in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device. It can be understood that the memory can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disk, or compact disc read-only memory (CD-ROM); magnetic surface memory can be magnetic disk memory or magnetic tape memory. Volatile memory can be random access memory (RAM), which is used as an external cache.By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronized dynamic random access memory (SLDRAM), direct RAM bus random access memory (DRRAM). The memory described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memory. The method disclosed in the above embodiments of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or by instructions in the form of software. The above-mentioned processor can be a general-purpose processor, a DSP (Digital Signal Processing, i.e., a chip capable of implementing digital signal processing technology), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory. The processor reads the program in the memory and completes the steps of the above method in combination with its hardware. When the processor executes the program, the corresponding processes in the various methods of the embodiments of the present application are implemented. For the sake of brevity, they are not described here.
[0119] For the description of the relevant parts of the outdoor power transmission channel patrol intelligent detection equipment provided in Example 3 of the present application, please refer to the detailed description of the corresponding parts of the outdoor power transmission channel patrol intelligent detection method provided in Example 1 of the present application, and will not be repeated here.
[0120] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements are inherent to the elements. In the absence of further restrictions, the elements limited by the statement "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device comprising the elements. In addition, the above-mentioned technical solutions provided in the embodiments of the present application are not described in detail in accordance with the corresponding technical solutions in the prior art to achieve the same principle, so as to avoid excessive elaboration.
[0121] Although the above description is of specific embodiments of the present invention in conjunction with the accompanying drawings, it does not limit the scope of protection of the present invention. For those skilled in the art, other different forms of modifications or variations can be made based on the above description. It is not necessary and impossible to list all embodiments here. Based on the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without expending creative effort are still within the scope of protection of the present invention.
Claims
1. An intelligent detection method for outdoor power transmission channel inspection, characterized in that: The following steps are involved: Generate structured text prompts based on transmission inspection task parameters and environmental information; use text encoders to generate task-related text features; Extract visual features from inspection images and align the semantic spaces of these visual features with those of text features using a contrastive loss function and a text-guided cross-modal attention mechanism. The geographic information system coordinates and meteorological data are integrated with the text features to generate an environment-aware segmentation mask and output defect area detection results.
2. The outdoor power transmission channel inspection intelligent detection method according to claim 1 is characterized in that: The structured text prompts include: Inspection task type, tower number, and line section parameters; High-frequency keywords embedded in the terminology dictionary of the power sector; Integrates real-time weather conditions and geographic coordinate information.
3. The outdoor power transmission channel inspection intelligent detection method according to claim 2, characterized in that: The process of using a text encoder to generate task-related text features includes: inputting structured text prompts into the CLIP text encoder to generate task-related text features; wherein the text features are represented as: F t = CLIP-TextEncoder("Detect {task type}, tower number: {number}, environment: {weather}"); Among them, the task types include at least one of microcrack detection, insulator damage identification, and mud erosion area segmentation.
4. The outdoor power transmission channel inspection intelligent detection method according to claim 1, characterized in that: The process of extracting visual features of the inspection image includes: inputting the inspection image into a downsampling encoder to extract visual features.
5. The outdoor power transmission channel inspection intelligent detection method according to claim 1, characterized in that: The specific process of aligning the semantic space of the visual features and text features by contrasting the loss function and the text-guided cross-modal attention mechanism includes: Where S = CosSimilarity(F v ,F t ); L mask =Focal(M pred ,M gt )+Dice(M pred ,M gt ); L class =CrossEntropy(C pred ,C gt ); Among them, F v is the image feature; F t is the text feature; CosSimilarity() represents cosine similarity; L contrastive is the contrast loss; i represents image; j represents text; M represents the number of batch samples, i.e., image-text pairs; τ represents the temperature coefficient; S ii Represents the similarity of the positive sample pair; S ij Represents the similarity matrix element; L mask Represents mask loss; Focal(M pred ,M gt ) is used to indicate the degree of overlap between the true mask and the true category; Dice (M pred ,M gt ) is used to adjust the weight between the true mask and the true category; M pred is the real mask; M gt is the real category; L class represents the cross entropy loss; C pred Represents the classification confidence of the true mask; C gt Represents the true category index of each target region.
6. The outdoor power transmission channel inspection intelligent detection method according to claim 4, characterized in that: The method further comprises: The visual features are used as tokens, a K-center clustering algorithm is used to cluster the visual features to generate an irregular token set, and a domain matrix for each token is constructed; Calculate the local influence between tokens using the domain matrix; The downsampling center is selected according to the importance score of each token, and the neighborhood token features are merged to achieve efficient retention of small-sized cracks and corrosion spots.
7. The outdoor power transmission channel inspection intelligent detection method according to claim 1, characterized in that: The process of calculating the local influence between tokens using the domain matrix is as follows: Among them, X l is the output of the lth decoding block, X l-1 is the output of the l-1th decoding block; A l is the neighborhood matrix in the lth decoding block; Q l is the query matrix of the lth decoding block; K l is the key value of the lth decoding block; V l is the Value matrix of the lth decoding block; x and y are different visual features in the token set.
8. The outdoor power transmission channel inspection intelligent detection method according to claim 1, characterized in that: The output defect area detection results include: insulator cracks, hardware corrosion and wire wear; segmentation mask and defect category are also output.
9. An outdoor power transmission channel inspection intelligent detection system, characterized in that: Including text processing module, image processing module and detection output module; The text processing module is used to generate structured text prompts based on the power transmission inspection task parameters and environmental information; and use a text encoder to generate text features related to the task; The image processing module is used to extract visual features of the inspection image and align the semantic space of the visual features with the text features through a contrast loss function and a text-guided cross-modal attention mechanism; The detection output module is used to fuse geographic information system coordinates and meteorological data with the text features, generate an environment-aware segmentation mask, and output defect area detection results.
10. An outdoor power transmission channel inspection intelligent detection equipment, characterized in that: include: memory for storing computer programs; A processor, configured to implement the method steps according to any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Positioning method of power transmission line inspection image insulator
CN110490261A
Intelligent industrial defect detection method and system based on fusion model
CN117853492A
Power transmission line operation and maintenance method and device based on multi-modal feature fusion and computer readable storage medium
CN119339197A
Detection method of power transformer oil leakage detection system based on multi-mode prompt and multi-scale segmentation
CN119646668A
Visual detection method for suspension state defects of high-speed rail overhead line system in open environment
CN119941640A