A power transmission line facility inspection method and system
Patent Information
- Application Number
- CN202610972828.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明提供了一种输电线路设施巡检方法及系统,解决了当前主流的输电线路设施巡检识别方式整体巡检工作灵活性较差的技术问题
[0062]本发明的上述技术方案提供了一种输电线路设施巡检方法,获取无人机巡检原始图像和输电设施图像,并根据无人机巡检原始图像和输电设施图像,生成高维度深层视觉特征序列、结构化描述文本和适配新目标的任务特定元映射器参数;采用大规模预训练语言模型根据结构化描述文本、无人机巡检原始图像和适配新目标的任务特定元映射器参数,生成结构化关联列表;通过大规模预训练语言模型根据结构化描述文本、自然语言测量指令和图文范例任务支持集,执行非对称目标精确测量,得到非对称目标精确测量结构化数据;采用大规模预训练语言模型根据高维度深层视觉特征序列、结构化关联列表和适配新目标的任务特定元映射器参数,开展上下文感知的防震锤序列距离标定,输出防震锤连续距离链表;采用大规模预训练语言模型根据结构化描述文本、结构化关联列表、非对称目标精确测量结构化数据和防震锤连续距离链表,开展逻辑校验与数据净化,输出逻辑验证报告和洁净数据集;采用大规模预训练语言模型根据逻辑验证报告、洁净数据集、报告格式自然语言指令、标准报告范例任务支持集和适配新目标的任务特定元映射器参数进行巡检分析,得到综合巡检分析报告;基于上述方案,本发明依托输电设施图像生成适配新目标的任务特定元映射器参数,无需对大规模预训练语言模型整体重新训练,也无需准备海量标注样本,仅依靠专用元映射器参数就能完成新型输电设施的识别适配。在生成结构化关联列表、开展防震锤序列距离标定的环节中,方案复用适配新目标的任务特定元映射器参数,可灵活适配不同排布方式、不同组合形态的输电线路附属设施,不必针对差异化的线路布局单独设计配套算法。执行非对称目标精确测量时,依靠自然语言测量指令搭配图文范例任务支持集驱动模型完成测量作业,能够结合输电设施的不同结构、现场不同测量要求灵活定义测量任务。逻辑校验与数据净化环节会整合全流程产生的各类巡检数据,统一完成校验、筛选工作,依托现有模型与参数实现自动化数据处理,无需为不同类型的巡检数据单独定制校验规则,减少了场景切换过程中的规则调整工作量。最后的巡检分析与报告生成环节,结合报告格式自然语言指令和标准报告范例任务支持集开展作业,可根据运维工作的不同需求,灵活生成对应样式、对应分析侧重点的综合巡检分析报告,满足多元化的成果输出要求,整体有效提升了输电线路巡检方式的灵活性。
Smart Images

Figure CN122820177A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power transmission line operation and maintenance technology, and in particular to a method and system for inspecting power transmission line facilities. Background Technology
[0002] The power system is continuously upgrading and developing towards intelligence, and the status monitoring and operational status analysis of power grid equipment has become a core technological support for ensuring the safe and stable operation of the entire power grid. Relying on the continuous iteration and optimization of technologies such as sensors, drone inspections, and satellite remote sensing, the operational data of various power grid equipment can be collected and summarized in real time. The massive field data resources also provide a solid foundation for intelligent operation and maintenance of transmission lines, fault prediction, and operational risk assessment.
[0003] With the overall trend of intelligent power development, the operation and maintenance scenarios of transmission lines are becoming increasingly complex, and the difficulty and workload of inspection work are constantly increasing. Under these circumstances, the drawbacks of traditional transmission line facility inspection methods are becoming increasingly prominent. Whether in terms of work efficiency, detection accuracy, or overall intelligence level, there are obvious deficiencies, making it difficult to meet the actual requirements of current transmission line operation and maintenance work.
[0004] Current mainstream methods for identifying and locating power transmission line facilities primarily rely on conventional deep learning models to perform facility identification, localization, and ranging. This technology operates by training the model using a large number of manually labeled image samples, and then using the trained, fixed model parameters to complete on-site inspections. Because the model's recognition capabilities are entirely built upon existing labeled samples, if a new or rare type of power transmission facility appears during the inspection, the existing sample library cannot match the new target. Staff must then collect massive amounts of corresponding images, complete sample labeling, and retrain the entire model, making rapid functional adaptation impossible. Limited by the strong dependence on massive amounts of labeled data and the cumbersome process of adapting to new targets, existing methods for inspecting power transmission line facilities are poorly adaptable to scenarios such as equipment upgrades and the appearance of unusual facilities, resulting in poor overall flexibility in inspection work. Summary of the Invention
[0005] This invention provides a method and system for inspecting power transmission line facilities, which solves the technical problem of poor overall inspection flexibility in current mainstream power transmission line facility inspection and identification methods.
[0006] The first aspect of this invention provides a method for inspecting power transmission line facilities, comprising:
[0007] Acquire raw images of UAV inspections and images of power transmission facilities, and generate high-dimensional deep visual feature sequences, structured descriptive text, and task-specific meta-mapper parameters adapted to new targets based on the raw images of UAV inspections and the images of power transmission facilities.
[0008] A large-scale pre-trained language model is used to generate a structured association list based on the structured description text, the original images of the UAV inspection, and the task-specific meta-mapping parameters of the new target.
[0009] The large-scale pre-trained language model performs asymmetric target precision measurement based on the structured description text, natural language measurement instructions, and image-text example task support set, and obtains asymmetric target precision measurement structured data.
[0010] The large-scale pre-trained language model is used to perform context-aware vibration damper sequence distance calibration based on the high-dimensional deep visual feature sequence, the structured association list, and the task-specific meta-mapper parameters adapted to the new target, and outputs a continuous distance linked list of vibration dampers.
[0011] The large-scale pre-trained language model is used to perform logical verification and data purification based on the structured description text, the structured association list, the asymmetric target precise measurement structured data, and the vibration damper continuous distance linked list, and outputs a logical verification report and a clean dataset.
[0012] The large-scale pre-trained language model is used to perform inspection analysis based on the logical verification report, the clean dataset, the natural language instructions for the report format, the standard report example task support set, and the task-specific meta-mapping parameters for adapting to the new target, to obtain a comprehensive inspection analysis report.
[0013] Optionally, the step of generating a high-dimensional deep visual feature sequence, structured descriptive text, and task-specific meta-mapper parameters adapted to the new target based on the original images of the UAV inspection and the images of the power transmission facilities includes:
[0014] The original images of the UAV inspection are pixel-vectorized to obtain a high-dimensional deep visual feature sequence;
[0015] The high-dimensional deep visual feature sequence is concatenated with visual prefix parameters that can be learned online to obtain a concatenated feature sequence.
[0016] The spliced feature sequence is subjected to feature focusing and semantic extraction to obtain the updated guiding visual prefix;
[0017] The updated guiding visual prefix is subjected to autoregressive text generation processing to obtain structured descriptive text.
[0018] The images of the power transmission facilities and the corresponding text annotations are used to form a facility recognition task support set;
[0019] Based on the facility identification task support set, gradient update processing of the general meta-parameters of the meta-mapper is performed to obtain task-specific meta-mapper parameters adapted to the new target.
[0020] Optionally, the step of generating a structured association list using a large-scale pre-trained language model based on the structured descriptive text, the original UAV inspection image, and the task-specific meta-mapping parameters for adapting to the new target includes:
[0021] Extract facility information from the structured description text and the original images of the UAV inspection;
[0022] Based on the facility information and the original images from the UAV inspection, construct combined data of local facility information and global scene context;
[0023] From the combined data of local facility information and global scene context, typical auxiliary facilities of transmission lines, namely vibration dampers and line clamps, are identified, and the two types of facilities are used as associated objects to generate a candidate pairing set of vibration dampers and line clamps.
[0024] Using the task-specific meta-mapping parameters of the new target, visual feature encoding and rationality loss calculation are performed on the candidate pairs of vibration damper and wire clamp in the candidate pair set to obtain the exclusive guiding visual prefix of the candidate pairs of vibration damper and wire clamp.
[0025] The dedicated guiding visual prefix is input into the large-scale pre-trained language model to generate a structured association description text of the vibration damper and the clamp.
[0026] The structured association description text of the vibration damper-line clamp is standardized to obtain a structured association list.
[0027] Optionally, the step of performing asymmetric target precision measurement using the large-scale pre-trained language model based on the structured description text, natural language measurement instructions, and image-text example task support set to obtain asymmetric target precision measurement structured data includes:
[0028] The natural language measurement instructions and the image and text example task support set are combined to obtain combined measurement instruction and task support set data;
[0029] Based on the structured description text and the combined data of the measurement instructions and task support set, the general meta-parameters of the meta-mapper are updated by gradient to obtain task-specific meta-mapper parameters adapted to the measurement task.
[0030] Based on the task-specific meta-mapper parameters of the adapted measurement task and the structured description text, a measurement task-specific guiding visual prefix is determined;
[0031] The measurement task-specific guiding visual prefix is input into the large-scale pre-trained language model to generate structured text of key point coordinates;
[0032] Extract the pixel coordinates of key points from the structured text containing the key point coordinates;
[0033] The pixel distance value is obtained by calculating the Euclidean distance to the pixel coordinates of the key points.
[0034] By integrating the natural language measurement instructions, the structured text of the key point coordinates, and the pixel distance values, structured data for accurate measurement of asymmetric targets is obtained.
[0035] Optionally, the step of using the large-scale pre-trained language model to perform context-aware vibration damper sequence distance calibration based on the high-dimensional deep visual feature sequence, the structured association list, and the task-specific meta-mapper parameters adapted to the new target, and outputting a continuous distance linked list of vibration dampers, includes:
[0036] From the structured association list and the high-dimensional deep visual feature sequence, filter the vibration dampers attached to the same conductor or clamp system, sort them by spatial location and filter invalid targets through a preset height threshold to obtain the visual features of the vibration damper sequence.
[0037] The visual features of the vibration damper sequence are concatenated with visual prefix parameters that can be learned online to obtain the sequence concatenation features;
[0038] Using the task-specific meta-mapper parameters for adapting to the new target, the sequence splicing features are context-encoded to obtain an updated visual prefix rich in sequence context information;
[0039] The updated visual prefix rich in sequence context information is input into the large-scale pre-trained language model to perform iterative inference of the distance between adjacent vibration dampers, thereby obtaining the text of the distance measurement between each pair of adjacent vibration dampers.
[0040] The text describing the distance measurements of each pair of adjacent vibration dampers is organized into a linked list of continuous distances for the vibration dampers.
[0041] Optionally, the large-scale pre-trained language model performs logical verification and data cleansing based on the structured description text, the structured association list, the asymmetric target precise measurement structured data, and the vibration damper continuous distance linked list, outputting a logical verification report and a clean dataset, including:
[0042] Based on the structured description text, the structured association list, the structured data of precise measurement of the asymmetric target, and the continuous distance linked list of the vibration damper, construct a complete dataset to be verified;
[0043] Based on the complete dataset to be verified, gradient updates and feature encoding are performed on the general meta-parameters of the meta-mapper to obtain a guiding visual prefix specific to the verification task.
[0044] The verification task-specific guiding visual prefix is input into the large-scale pre-trained language model to generate logical verification text.
[0045] The logical verification text is sorted and summarized to obtain a logical verification report;
[0046] Based on the logical verification text, contradictory data in the complete dataset to be verified are filtered to obtain a clean dataset.
[0047] Optionally, the large-scale pre-trained language model is used to perform inspection analysis based on the logical verification report, the clean dataset, the report format natural language instructions, the standard report example task support set, and the task-specific meta-mapping parameters adapted to the new target, to obtain a comprehensive inspection analysis report, including:
[0048] The report format natural language instructions and the standard report example task support set are combined to obtain report generation task instruction and example combination data;
[0049] Global condensed data is generated by extracting data from the logical verification report and the clean dataset.
[0050] Using the task-specific meta-mapper parameters adapted to the new target, the global condensed data for report generation is encoded to obtain a report generation-specific guiding visual prefix.
[0051] The report generates a unique guiding visual prefix, which is then input into the large-scale pre-trained language model to generate a comprehensive inspection and analysis report.
[0052] A second aspect of the present invention provides a transmission line facility inspection system, comprising:
[0053] The acquisition module is used to acquire the original images of the UAV inspection and the images of the power transmission facilities, and generate a high-dimensional deep visual feature sequence, structured descriptive text, and task-specific meta-mapper parameters adapted to the new target based on the original images of the UAV inspection and the images of the power transmission facilities.
[0054] The generation module is used to generate a structured association list based on the structured description text, the original UAV inspection image, and the task-specific meta-mapping parameters of the new target using a large-scale pre-trained language model;
[0055] The measurement module is used to perform precise measurement of asymmetric targets using the large-scale pre-trained language model based on the structured description text, natural language measurement instructions, and image-text example task support set, to obtain structured data for precise measurement of asymmetric targets.
[0056] The calibration module is used to perform context-aware vibration damper sequence distance calibration using the large-scale pre-trained language model based on the high-dimensional deep visual feature sequence, the structured association list, and the task-specific meta-mapper parameters of the new target, and output a continuous distance linked list of vibration dampers.
[0057] The verification module is used to perform logical verification and data purification based on the structured description text, the structured association list, the asymmetric target precise measurement structured data, and the vibration damper continuous distance linked list using the large-scale pre-trained language model, and outputs a logical verification report and a clean dataset.
[0058] The analysis module is used to perform inspection analysis using the large-scale pre-trained language model based on the logical verification report, the clean dataset, the natural language instructions for the report format, the standard report example task support set, and the task-specific meta-mapping parameters for the new target, to obtain a comprehensive inspection analysis report.
[0059] A third aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the transmission line facility inspection method described above.
[0060] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the transmission line facility inspection method as described above.
[0061] As can be seen from the above technical solutions, the present invention has the following advantages:
[0062] The above-mentioned technical solution of the present invention provides a method for inspecting power transmission line facilities. This method acquires original images of UAV inspections and images of the power transmission facilities. Based on these images, it generates a high-dimensional deep visual feature sequence, structured descriptive text, and task-specific meta-mapper parameters adapted to new targets. A large-scale pre-trained language model is used to generate a structured association list based on the structured descriptive text, the original UAV inspection images, and the task-specific meta-mapper parameters adapted to new targets. The large-scale pre-trained language model then performs precise asymmetric target measurement based on the structured descriptive text, natural language measurement instructions, and a set of image-text example tasks, obtaining structured data for precise asymmetric target measurement. Finally, the large-scale pre-trained language model is used to perform upstream and downstream... The invention employs a text-aware method for calibrating the distance sequence of vibration dampers, outputting a continuous distance list of the dampers. A large-scale pre-trained language model is used to perform logical verification and data cleansing based on structured descriptive text, structured association lists, structured data from precise measurements of asymmetric targets, and the continuous distance list of vibration dampers, outputting a logical verification report and a clean dataset. Another large-scale pre-trained language model is used to perform inspection analysis based on the logical verification report, the clean dataset, natural language instructions for the report format, a standard report example task support set, and task-specific meta-mapping parameters adapted to the new target, resulting in a comprehensive inspection analysis report. Based on this approach, the invention generates task-specific meta-mapping parameters adapted to the new target from images of power transmission facilities. This eliminates the need for retraining the entire large-scale pre-trained language model or preparing massive amounts of labeled samples; the identification and adaptation of new power transmission facilities can be completed solely using dedicated meta-mapping parameters. In the stages of generating the structured association list and calibrating the distance sequence of vibration dampers, the solution reuses the task-specific meta-mapping parameters adapted to the new target, flexibly adapting to different layouts and combinations of power transmission line ancillary facilities, without requiring separate algorithms designed for differentiated line layouts. When performing precise measurements of asymmetric targets, the measurement operation is driven by a model that relies on natural language measurement commands combined with graphic and textual example task support sets. This allows for flexible definition of measurement tasks based on different structures of transmission facilities and varying on-site measurement requirements. The logic verification and data purification stages integrate various inspection data generated throughout the entire process, uniformly completing verification and filtering. Automated data processing is achieved using existing models and parameters, eliminating the need to customize verification rules for different types of inspection data and reducing the workload of rule adjustments during scenario switching. Finally, the inspection analysis and report generation stages utilize natural language commands for report formats and standard report example task support sets. This allows for the flexible generation of comprehensive inspection analysis reports with corresponding styles and analytical focuses, meeting diverse output requirements and effectively improving the flexibility of transmission line inspection methods. Attached Figure Description
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0064] Figure 1 This is a flowchart of the steps of a method for inspecting power transmission line facilities according to Embodiment 1 of the present invention;
[0065] Figure 2 This is a structural block diagram of a power transmission line facility inspection system provided in Embodiment 2 of the present invention. Detailed Implementation
[0066] This invention provides a method and system for inspecting power transmission line facilities, which solves the technical problem of poor overall inspection flexibility in current mainstream power transmission line facility inspection and identification methods.
[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be noted that in the optional embodiments of the present invention, the object information and other related data involved require the permission or consent of the object when the embodiments of the present invention are applied to specific products or technologies, and the collection, use, and processing of related data must comply with relevant laws, regulations, and standards. That is to say, if the embodiments of the present invention involve data related to the object, it needs to be obtained with the authorization and consent of the object, the authorization and consent of relevant departments, and in compliance with relevant laws, regulations, and standards. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments also need to be implemented with the authorization and consent of the object.
[0068] Please see Figure 1 , Figure 1 This is a flowchart illustrating the steps of a method for inspecting power transmission line facilities according to Embodiment 1 of the present invention.
[0069] The present invention provides a method for inspecting power transmission line facilities, comprising:
[0070] Step 101: Obtain the original images of the UAV inspection and the images of the power transmission facilities, and generate a high-dimensional deep visual feature sequence, structured descriptive text, and task-specific meta-mapper parameters adapted to the new target based on the original images of the UAV inspection and the images of the power transmission facilities.
[0071] The task-specific meta-mapping parameters adapted to the new target are special model parameters obtained by performing gradient update processing on the facility recognition task support set composed of power transmission facility images and their corresponding text annotations, using the general meta-mapping parameters as the initial basic parameters. These parameters are generated for power transmission facility recognition scenarios and can be reused in multiple inspection stages such as structured association list generation, vibration damper sequence distance calibration, inspection analysis and comprehensive report generation, to complete visual feature encoding, context encoding, inspection data encoding and other processing tasks.
[0072] The original images of drone inspections are visible light images of the entire scene collected by drones along the predetermined inspection route of the power transmission line. The images also include power transmission towers, conductors, ancillary facilities, and vegetation and terrain background around the line. They are original visual data covering the global spatial information of the inspection section.
[0073] The images of power transmission facilities are close-up local images of independent auxiliary facilities such as vibration dampers and clamps on power transmission lines, with redundant background interference removed. They are accompanied by manually labeled facility categories and condition text tags, and are only used for small-sample adaptation training of new facilities.
[0074] The high-dimensional deep visual feature sequence is a one-dimensional feature vector sequence generated by the frozen large-scale pre-trained visual encoder after it has completed deep semantic encoding of all pixels in the inspection image, according to the spatial arrangement order of the image. It carries implicit visual information such as facility edges, deformation, and relative position that cannot be directly identified by the naked eye.
[0075] The structured description text is a standardized machine-readable text generated in accordance with the unified semantic specification of power inspection. It contains three fixed structured fields: facility pixel coordinates, facility category, and appearance and operation status, without redundant natural language descriptions.
[0076] It should be noted that after the two types of images are acquired simultaneously, they are processed separately. The raw images from the UAV inspection are sent to the frozen visual encoder to complete pixel vectorization and semantic decoding, and output visual feature sequences and structured descriptive text. The images of power transmission facilities are paired with corresponding labeled text to form a small sample support set, and the meta-parameters are fine-tuned separately. Finally, the three types of pre-data are output simultaneously to provide basic input for subsequent full-process inspection reasoning.
[0077] Further, step 101 may include the following sub-steps:
[0078] S11. Perform pixel vectorization on the original images of the UAV inspection to obtain a high-dimensional deep visual feature sequence;
[0079] S12. Concatenate the high-dimensional deep visual feature sequence with the visual prefix parameters that can be learned online to obtain the concatenated feature sequence;
[0080] S13. Perform feature focusing and semantic extraction on the spliced feature sequence to obtain the updated guiding visual prefix;
[0081] S14. Perform autoregressive text generation processing on the updated guiding visual prefix to obtain structured descriptive text;
[0082] S15. Combine the images of power transmission facilities and the corresponding text annotations to form a support set for facility recognition tasks;
[0083] S16. Based on the facility identification task support set, perform gradient update processing on the general meta-parameters of the meta-mapper to obtain task-specific meta-mapper parameters adapted to the new target.
[0084] The facility identification task support set is a small-sample dedicated dataset built to adapt to the new power transmission facility identification task, consisting of 1 to 5 power transmission facility samples. Each sample contains two types of data: first, the image of the power transmission facility, which is a close-up image of the auxiliary facilities such as vibration dampers and clamps of the power transmission line, which removes redundant background interference and retains only the facility body and its key feature areas; second, the text annotation corresponding to the image, which is structured information that conforms to the semantic specifications of power inspection, including three core fields: facility category, model, appearance and operating status.
[0085] It should be noted that this invention employs a unified multimodal meta-learning framework, which consists of three core components: a large-scale pre-trained visual encoder (… A large-scale pre-trained language model ), and a lightweight, trainable meta-mapper that serves as a bridge between the two. To achieve efficient computation and rapid adaptation to new tasks, the visual encoder and language model maintain parameter freeze throughout the process, thus fully leveraging their inherently powerful knowledge capabilities. The system's learning and adaptation primarily focus on a lightweight meta-mapper, which is responsible for learning how to establish effective connections between visual features and linguistic information.
[0086] Furthermore, the raw image x captured by the drone inspection is input into the frozen visual encoder. The encoder is responsible for transforming the input image pixels into high-dimensional, machine-understandable deep visual features. The output of this process is a sequence of visual features, namely a high-dimensional deep visual feature sequence. This sequence is a vectorized representation of the image content in mathematical space.
[0087] Furthermore, the extracted visual feature sequences With a set of parameters that can be learned online, known as "visual prefixes". The sequences are then concatenated. The concatenated sequence is then fed into the metamapper. This mapper is built on a self-attention mechanism, and its core computation process follows the formula below:
[0088] ;
[0089] in, This is the nth feature vector in the visual feature sequence. The metamapper is the l-th learnable parameter vector in the sequence of learnable visual prefix parameters. The input is a concatenated feature sequence. By calculating the similarity between feature vectors within a sequence and performing a weighted summation, the meta-mapper can focus on and encode the most critical semantic information in an image. The output of this process is a set of updated guiding visual prefixes that have been integrated and refined with contextual information. It will serve as a customized instruction to guide the language model for subsequent processing.
[0090] Furthermore, the visual prefixes generated in the previous step, which contain the core semantics of the image, are then... Send to the frozen language model The language model uses this visual prefix as an initial condition and contextual guide to generate tokens describing the image content in an autoregressive manner. The generation of each subsequent token depends on the visual prefix and all previously generated tokens; this process can be expressed by the following formula:
[0091] ;
[0092] in, For the next generated token, it represents the output unit predicted and generated by the language model in step i+1; The frozen pre-trained language model is the core function that performs the generation task, and its subscript... This represents all network parameters contained within the model that remain unchanged throughout the process; To provide the language model with a complete input sequence, a guiding visual prefix containing the core semantics of the image is concatenated with the textual context previously generated by the model. Ultimately, the model outputs a complete, structured descriptive text containing the facility's location, category, and status.
[0093] Furthermore, the small-sample adaptability of this method is activated when new types of transmission facilities, unseen by the model, are encountered during inspections. Workers only need to provide a small number (e.g., 1 to 5) images of the new facility along with their textual annotations, forming a task support set. The system utilizes this support set to configure the general meta-parameters of the meta-mapper. Several gradient updates are performed to quickly adapt to the new task, resulting in a set of task-specific meta-mapper parameters adapted to the new objective. The mathematical expression for this inner-loop update process is:
[0094] ;
[0095] in: The task-specific meta-mapping parameters adapted to the new target are the output of this formula, representing the meta-mapping parameters for a specific task. The new version after minor adjustments and updates; These are the general meta-parameters of the meta-mapper, representing the initial parameters that the meta-mapper learns after meta-training and have good generalization ability. They are the starting point for this update. The learning rate for the internal loop is a hyperparameter that controls the step size of parameter updates. For the task loss function, the gradient with respect to the meta-parameters points to the loss that enables the task to perform a task-supported loss function. The direction of fastest growth is the opposite direction, which is the direction for parameter optimization. Through this mechanism, the model can quickly master the ability to accurately identify new targets without large-scale retraining.
[0096] In this embodiment, the process of pixel vectorizing the original images from UAV inspections to obtain a high-dimensional deep visual feature sequence is as follows: the original images captured by the UAV inspection are input into a frozen, large-scale pre-trained visual encoder. This encoder transforms the input image pixels into high-dimensional, machine-understandable deep visual features, and outputs a high-dimensional deep visual feature sequence. The process of concatenating high-dimensional deep visual feature sequences with online learnable visual prefix parameters to obtain concatenated feature sequences is as follows: The high-dimensional deep visual feature sequences... With a set of visual prefix parameter sequences that can be learned online By concatenating them sequentially, we obtain the concatenated feature sequence, which is the input sequence of the meta-mapper. The process of performing feature focusing and semantic extraction on the concatenated feature sequence to obtain the updated guiding visual prefix is as follows: the concatenated feature sequence is fed into a meta-mapper constructed based on a self-attention mechanism. According to the formula The calculation involves calculating the similarity between feature vectors within the sequence and then weighting and summing them. This process focuses on key semantic information in the image, encodes and extracts it, and yields the updated guiding visual prefix. The process of generating structured descriptive text by performing autoregressive text generation on the updated guiding visual prefixes is as follows: The updated guiding visual prefixes are... Feeding into a frozen large-scale pre-trained language model Words are generated one by one in an autoregressive manner, and the generation process satisfies the formula ,in For the lexical units generated in step i+1, the input sequence is formed by concatenating guiding visual prefixes with the generated text context, and the final output is a structured descriptive text containing facility location, category, and status information. The process of forming a facility recognition task support set by combining power transmission facility images and their corresponding text annotations is as follows: collect 1 to 5 power transmission facility images and their corresponding text annotations to form the facility recognition task support set; the annotations include information such as facility category and operating status. The process of performing gradient update processing on the general meta-parameters of the meta-mapper based on the facility recognition task support set to obtain task-specific meta-mapper parameters adapted to the new target is as follows: using the general meta-parameters of the meta-mapper... The initial parameters are based on the task loss function corresponding to the facility identification task support set. Through formula Perform several gradient updates to obtain task-specific meta-mapper parameters adapted to the new target. .
[0097] Step 102: Using a large-scale pre-trained language model, a structured association list is generated based on the structured description text, the original images of the UAV inspection, and the task-specific meta-mapper parameters adapted to the new target.
[0098] It should be noted that, taking the structured description text output from the previous steps, the original images of the UAV inspection, and the task-specific meta-mapping parameters adapted to the new target as input, the system first extracts the category and status information of the facilities from the structured description text, and extracts the spatial location information of the facilities from the original images of the UAV inspection. It then constructs a combination of local facility information and global scene context data, identifies typical transmission line auxiliary facilities such as vibration dampers and clamps from this combination data, and generates a candidate pairing set of vibration damper-clamps as associated objects. The system then uses the task-specific meta-mapping parameters adapted to the new target to perform visual feature encoding and rationality loss calculation on the candidate pairing set, obtaining a unique guiding visual prefix for the vibration damper-clamp candidate pairings. This unique guiding visual prefix is input into a large-scale pre-trained language model to generate a structured association description text for vibration damper-clamps. This description text is then standardized to obtain a structured association list.
[0099] The structured association list is an ordered list obtained by integrating the structured association description text of vibration dampers and clamps and standardizing it according to the unified data format and field specifications of power inspection. The list is organized in a machine-readable structured entry format. Each independent entry in the list corresponds to a pair of vibration dampers and clamps. Specifically, it includes a unique pairing number, vibration damper equipment identifier, clamp equipment identifier, corresponding association relationship labels of the two types of facilities, facility pixel spatial coordinates, location matching verification information, conductor segment number, line span information, and other complete structured content, thus completely preserving the pairing association and spatial distribution data of transmission line ancillary facilities.
[0100] The combined data of local facility information and global scene context is an integrated image and text data constructed by fusing local facility information extracted from structured descriptive text and global scene information obtained from the parsing of original images from UAV inspections. The data is bound and organized with a single power transmission facility as the basic unit. The local facility information specifically includes detailed content such as the facility's unique code, specific category, appearance and operating status, local pixel coordinates, and local appearance features. The global scene context information specifically includes the facility's global coordinates within the entire inspection image, environmental information such as surrounding terrain and vegetation, conductor layout, line section number, distribution of adjacent facilities, and overall line topology.
[0101] The candidate pairing set for vibration dampers and clamps is a set of potential pairing combinations of facilities generated based on spatial affiliation after accurately identifying typical auxiliary facilities of transmission lines, namely vibration dampers and clamps, from the combination of local facility information and global scene context data. The set uses a single pairing combination as the basic element and only includes vibration damper and clamp combinations within the same conductor segment, the same line span, and the same clamp system. It covers all pairing forms of the two types of facilities with physical matching possibilities under this inspection scenario. Each candidate pairing unit in the set synchronously stores basic attribute data such as the corresponding vibration damper and clamp device number, local pixel coordinates, and the line segment to which it belongs.
[0102] The structured association description text of the vibration damper and clamp is a standardized text generated by a large-scale pre-trained language model based on the exclusive guiding visual prefix corresponding to the candidate pairing set of vibration damper and clamp. The text strictly follows the preset semantic specifications for the association description of power transmission facilities, and corresponds to each candidate pairing in the set. Specifically, it includes structured fields such as pairing number, vibration damper equipment number and basic attributes, clamp equipment number and basic attributes, precise location correspondence between the two types of facilities, conductor segment number, relative pixel distance between facilities, and installation matching relationship description.
[0103] Furthermore, step 102 may include the following sub-steps:
[0104] S21. Extract facility information from structured description text and raw images from drone inspections;
[0105] S22. Based on facility information and original images from drone inspections, construct combined data of local facility information and global scene context.
[0106] S23. Identify typical auxiliary facilities of transmission lines, namely vibration dampers and line clamps, from the combination data of local facility information and global scene context, and generate a candidate pairing set of vibration damper-line clamp as the two types of facilities as associated objects.
[0107] S24. Using the task-specific meta-mapping parameters adapted to the new target, perform visual feature encoding and rationality loss calculation on the candidate pairs of vibration dampers and wire clamps in the candidate pairing set to obtain the exclusive guiding visual prefix of the candidate pairs of vibration dampers and wire clamps.
[0108] S25. Input the exclusive guiding visual prefix into a large-scale pre-trained language model to generate a structured relational description text of the vibration damper and the clamp.
[0109] S26. Standardize the structured association description text of the vibration damper and clamp to obtain a structured association list.
[0110] The vibration damper-clamp candidate pairing set is a set of potential facility pairing combinations generated based on spatial affiliation and physical compatibility constraints after accurately identifying typical transmission line auxiliary facilities such as vibration dampers and clamps from the combination of local facility information and global scene context data. The set is organized in the form of an ordered list, only including vibration damper and clamp combinations within the same conductor segment, the same line span, and the same clamp system. It covers all pairing forms with physical compatibility possibilities for the two types of facilities under this inspection scenario, and does not include invalid pairings that do not conform to the line layout logic, such as those involving spans or different conductor types. Each candidate pairing unit in the set synchronously stores the complete basic data of the corresponding vibration damper and clamp, specifically including the unique identifier of the vibration damper equipment, facility category, pixel spatial coordinates, appearance status information, the unique identifier of the clamp equipment, facility category, pixel spatial coordinates, installation location information, and line topology data such as the line segment number and span identifier to which the pairing belongs. Each candidate pairing unit represents a potential physical connection relationship.
[0111] It should be noted that the system acquires information about the identified facilities, including the location, category, and status of each facility. Simultaneously, it retains the original UAV images as global scene context. This allows the model to focus on both local targets and the overall layout of the power transmission line, providing sufficient visual information for subsequent establishment of physical relationships.
[0112] Furthermore, for each first-class target (such as a vibration damper), the model generates a set of candidate pairs with potentially associated second-class targets (such as wire clamps). Each candidate pair represents a potential physical connection, but has not yet been logically verified. This step ensures that all possible associations are considered, providing a foundation for subsequent reasoning.
[0113] Furthermore, for each candidate association, the model uses a meta-mapper to encode the relevant visual features into guiding signals, i.e., visual prefixes, that the language model can understand. The correctness of the association judgment is measured by a task loss function, which is calculated on the task support set during the meta-training phase, and its formula is:
[0114] ;
[0115] in, To support centralized imagery, facilities A and B are included. The tags that correctly associate and describe it, The association descriptions generated for the model. By calculating the loss, the model is able to evaluate the plausibility of each candidate association.
[0116] Furthermore, the generated visual prefixes are input into a frozen language model, which generates structured text describing the relationships between each pair of facilities in an autoregressive manner. The language model, combining visual guidance and contextual information, can output explicit logical relationships, such as "vibration damper A is connected to the conductor where clamp B is located," rather than simply judgments based on geometric distance.
[0117] Furthermore, during the meta-training phase, the model utilizes the task support set and task query set to adjust the meta-mapper parameters. Optimization is performed to enable it to quickly adapt to new scenes with a limited number of samples. Through outer loop optimization (meta-parameters), the model ultimately learns to accurately determine the spatial relationships between objects from the input visual features and global context, thereby generating low-loss association results. This optimization objective can be expressed as:
[0118] ;
[0119] in: To minimize the optimization operator, the overall objective of this formula is to find a set of optimal initial parameters for the meta-mapper. This minimizes its total loss across all tasks; For cross-task summation operators, it means that the optimization process is performed on a distribution containing multiple different tasks. In the process, for each sampled task The loss is summed, which ensures that the model learns general, scalable capabilities, rather than knowledge specific to a particular task; The loss on the query set for the meta-task measures the performance error of the meta-mapper on the query set of the task when using task-specific parameters (i.e., parameters that have been quickly adapted by the inner loop). Optimizing this loss means improving the final performance of the model after it has "learned a new task".
[0120] Furthermore, after inference is complete, the model outputs structured association information. Each element explicitly records the source target and the target's ID, category, location information, physical connection relationship, and confidence level. The output data is both human-understandable and usable for precise measurement and analysis in subsequent steps.
[0121] In this embodiment, the process of extracting facility information from structured description text and original UAV inspection images is as follows: The location, category, and status information of each facility are extracted from the structured description text, while the original UAV inspection images are retained as a global scene context, providing local target information and overall transmission line layout information for subsequent correlation analysis. The process of constructing combined data of local facility information and global scene context based on the facility information and original UAV inspection images is as follows: Using facility information extracted from the structured description text, containing each facility's unique identifier, category, appearance, operating status, and pixel coordinates in the UAV image, along with the complete original UAV inspection image, as input, for each identified facility, the local region of interest is located in the original image based on its pixel coordinates. The pixel information of this local region is then bound and associated with the facility's textual local information to obtain image-text bound data of the facility's local information. Simultaneously, the complete original UAV inspection image is retained as a global scene context, and global topology and scene information such as the transmission line direction, line span division, relative positions between facilities, and surrounding terrain environment are extracted from the image. Subsequently, the bound data of each facility is combined with the global scene context. The information is combined to form a combined data of local facility information and global scene context. Each facility entry in the data simultaneously contains its own local feature information, as well as the overall layout of the transmission line and the global scene information of the surrounding environment. Typical auxiliary facilities of the transmission line, namely vibration dampers and clamps, are identified from the combined data of local facility information and global scene context. The process of generating a candidate pairing set of vibration damper-clamps is as follows: For each vibration damper, the model generates a candidate pairing set with potentially associated clamps. Each candidate pairing represents a potential physical connection relationship. The steps consider all possible associations to provide a foundation for subsequent reasoning. Utilizing task-specific meta-mapper parameters adapted to the new target, the process of visual feature encoding and rationality loss calculation for candidate vibration damper-wire clamp pairs in the candidate pairing set, resulting in a unique guiding visual prefix for each candidate vibration damper-wire clamp pair, is as follows: For each candidate association, the model uses the meta-mapper to encode the relevant visual features into a guiding signal, i.e., a visual prefix, that the language model can understand. The correctness of the association judgment is measured by a task loss function, which is calculated on the task support set during the meta-training phase, using the following formula: The process of evaluating the rationality of each candidate association by calculating the loss and inputting the exclusive guiding visual prefix into a large-scale pre-trained language model to generate structured association description text for the vibration damper-line clamp is as follows: the generated visual prefix is input into the frozen language model, and the model generates structured text describing the association between each pair of facilities in an autoregressive manner. The language model combines visual guidance and contextual information to output clear logical relationships, such as "vibration damper A is connected to the conductor where line clamp B is located". The process of standardizing the structured association description text of the vibration damper-line clamp to obtain a structured association list is as follows: the structured association information output by the model is standardized, and each element clearly records the source target and the target's ID, category, location information, physical connection relationship and confidence level. The resulting structured association list data can be understood by humans and can also be used for precise measurement and analysis in subsequent steps.
[0122] Step 103: Using a large-scale pre-trained language model, perform precise asymmetric target measurement based on structured description text, natural language measurement instructions, and image-text example task support set to obtain structured data for precise asymmetric target measurement.
[0123] It should be noted that the natural language measurement instructions and the image and text example task support set are combined into a measurement instruction and task support set combined data. The pre-set meta-mapper general meta-parameters are updated by gradient with the structured description text to obtain task-specific meta-mapper parameters adapted to the measurement task. Based on these parameters and the structured description text, a measurement task-specific guiding visual prefix is generated. The prefix is then input into a large-scale pre-trained language model to generate key point coordinate structured text. The key point pixel coordinates are extracted and the Euclidean distance is calculated to obtain the pixel distance value. The natural language measurement instructions, key point coordinate structured text, and pixel distance value are integrated to obtain the structured data for accurate measurement of asymmetric targets.
[0124] Among them, the natural language measurement instructions are task instruction texts written in general natural language sentence patterns and specifically designed for the measurement scenario of asymmetric auxiliary facilities of transmission lines. The text contains a number of clear execution elements, including the specific type of asymmetric facility being measured, the name and number of key measurement points to be located, the measurement dimension with the distance between the two points as the core, the writing specifications of coordinate data and distance results, and the overall layout format of the output text. It can clearly define the execution boundary and output standard of this measurement task.
[0125] The image and text example task support set is a small-sample image and text pairing dataset built for the measurement scenario of asymmetric targets of power transmission facilities. The dataset is composed of multiple independent samples in an ordered manner. Each sample consists of two parts: a real-scene image and corresponding annotation text. The real-scene image is an on-site image of asymmetric power transmission facilities under different shooting angles, lighting conditions, and installation postures. The annotation text is a collection of reference content, including the standard pixel coordinates of preset measurement key points in the image, the standard distance value calculated based on the coordinates, and compliant measurement result description text.
[0126] The task-specific metamapper parameters adapted for measurement tasks are lightweight trainable parameters obtained by combining structured description text, measurement instructions and task support set combined data, and gradient updates based on measurement task loss. These parameters are optimized only for asymmetric target precision measurement tasks and will not change the main parameters of the large-scale pre-trained language model throughout the process.
[0127] The measurement task-specific guiding visual prefix is a feature sequence generated by relying on the task-specific meta-mapper parameters adapted to the measurement task and integrating information such as facility type and spatial location recorded in the structured description text. This feature sequence has fixed dimensions and standardized data format, and is fed into a large-scale pre-trained language model as a pre-guiding input.
[0128] The key point coordinate structured text is standardized text generated by a large-scale pre-trained language model based on the measurement task-specific guiding visual prefix. The text is arranged according to a preset field format, specifically including the number of the measured asymmetric facility, the unique identifier of each measurement key point, the pixel x-coordinate, pixel y-coordinate, coordinate recognition confidence, and other core fields.
[0129] The structured data for precise measurement of asymmetric targets is a complete structured dataset formed by integrating three types of data according to a unified standard: natural language measurement instructions, structured text of key point coordinates, and pixel distance values. The data is organized in a machine-readable standardized format and is divided into three main sections: task instructions, key point coordinates, and measurement results. The task instructions section fully preserves the original natural language measurement instructions, the key point coordinates section records the pixel coordinate information of all measured key points one by one, and the measurement results section records the pixel distance values obtained through coordinate calculations.
[0130] Furthermore, step 103 may include the following sub-steps:
[0131] S31. Combine the natural language measurement instructions and the image and text example task support set to obtain combined measurement instructions and task support set data;
[0132] S32. Based on the structured description text and the combined data of measurement instructions and task support set, perform gradient updates on the general meta-parameters of the meta-mapper to obtain task-specific meta-mapper parameters adapted to the measurement task.
[0133] S33. Based on the task-specific meta-mapper parameters and structured description text adapted to the measurement task, determine the measurement task-specific guiding visual prefix;
[0134] S34. Input the measurement task-specific guiding visual prefix into a large-scale pre-trained language model to generate structured text of key point coordinates;
[0135] S35. Extract the pixel coordinates of key points from the structured text containing key point coordinates;
[0136] S36. Calculate the Euclidean distance between the key point pixel coordinates to obtain the pixel distance value;
[0137] S37. Integrate natural language measurement instructions, key point coordinate structured text, and pixel distance values to obtain accurate measurement structured data for asymmetric targets.
[0138] The combined measurement instruction and task support set data is a structured dataset that binds natural language measurement instructions for this precise asymmetric target measurement task with a set of graphic and textual example task support sets adapted to this measurement scenario, according to a preset data format. The dataset is organized with each individual measurement task as an independent unit. This combined data contains two core categories: the first is the natural language measurement instructions, which specifically include the type of asymmetric facility being measured (e.g., tension clamps, vibration dampers), the specific names and number of key measurement points to be located, the measurement dimensions (e.g., center-to-center distance, edge spacing), and the field specifications for the output data (e.g., the format for writing key point pixel coordinates, the numerical format requirements for distance results), such as specifying concrete measurement requirements like "measure the distance from the center of the left fixing bolt of the tension clamp to the outer edge of the nearest vibration damper." The second category is the graphic and textual example task support set, consisting of a small number (k-shots, typically 1 to 5) of examples related to this task. The measurement task consists of image-text paired samples of the same type. Each sample contains a real-scene image and corresponding labeled text. The real-scene image is a measurement scene image of the power transmission facility under different shooting angles, lighting conditions, and installation postures, which fully presents the key parts of the measured facility. The labeled text records the standard pixel coordinates of the preset measurement key points in the scene, the standard distance value calculated based on the coordinates, and the structured measurement result example text that conforms to the output specifications. Natural language measurement instructions are bound one-to-one with the image-text example task support set. Each measurement instruction is matched with a set of image-text example samples of the same type, which together constitute a complete combined data unit.
[0139] It should be noted that this step first defines the precise measurement task for asymmetric structural targets. This task is issued through flexible natural language commands, rather than relying on a fixed algorithm. Operators can provide a specific instruction based on the specific shape of the target (such as different models of tension clamps) and measurement requirements, such as "measure the distance from the center of the left fixing bolt of the tension clamp to the outer edge of the nearest vibration damper." To enable the model to understand and execute this specific task, the system will prepare a measurement instruction combined with a task support set of data. This support set contains a small number of (k-shot) graphic examples of this type of measurement that have already been performed. These examples (images) Corresponding measurement result text These factors together form the basis for the model's few-shot learning and task adaptation.
[0140] Furthermore, upon receiving the measurement instructions and task support set, the core objective of the model is to quickly adapt to this entirely new measurement rule. This process is achieved through the meta-mapper. The parameters are updated via an inner-loop update. The model utilizes the support set. The task loss is calculated, and one or more gradient descent passes are performed on the meta-parameters to obtain a set of specific parameters optimized for the current measurement task, i.e., task-specific meta-mapper parameters adapted to the measurement task. The mathematical expression for this adaptation process is:
[0141] ;
[0142] in It is the measurement error loss of the model on the support set. This is the internal loop learning rate. Subsequently, this set of specific parameters was incorporated. The meta-mapper processes the query image to be measured, encoding and fusing its complex visual features, especially key region features relevant to the measurement instructions, into a set of updated visual prefixes—that is, measurement task-specific guiding visual prefixes. This set of visual prefixes not only contains information about the target's appearance, but more importantly, it contains guiding instructions on "how to perform this specific measurement".
[0143] Furthermore, the generated guiding visual prefixes It was then fed into a frozen large-scale language model. In this context, the initial conditions for logical reasoning and localization are used. The language model, in an autoregressive manner, generates text describing the measurement process step by step, guided by visual prefixes. The most crucial aspect is outputting the precise coordinates of the starting and ending keypoints required for the measurement. Its generation process follows the formula:
[0144] ;
[0145] The output word sequence here Instead of simple descriptive sentences, it uses structured coordinate information, such as: "Starting point: center of the bolt on the left side of the tension clamp, coordinates: [x1, y1]. Ending point: outer edge of the vibration damper, coordinates: [x2, y2]".
[0146] Furthermore, after the language model successfully locates and outputs the coordinates of all necessary key points, the system performs the final geometric calculation, typically calculating the Euclidean distance between two points to obtain the final precise pixel distance value. The final output of this step is a structured data object that comprehensively records all information from the measurement, including the input natural language commands, the names of the key points located by the model, their pixel coordinates, and the calculated final distance. This output method not only ensures the accuracy of the measurement results but also provides complete traceability and interpretability, offering high-quality data input for subsequent analysis.
[0147] In this embodiment, the process of combining natural language measurement instructions and image-text example task support sets to obtain combined measurement instruction and task support set data is as follows: Natural language measurement instructions that specify the type of asymmetric facility being measured, key measurement points, measurement dimensions, and output format, such as "measure the distance from the center of the left fixing bolt of the tension clamp to the outer edge of the nearest anti-vibration hammer," are combined with image-text example task support sets containing a small number of k-shots of completed measurements of the same type (each example includes an image of a power transmission facility measurement scene and corresponding standard measurement result text). This results in combined measurement instruction and task support set data, which provides the foundation for the model to perform few-shot learning and task adaptation. The process of updating the general meta-parameters of the meta-mapper based on the structured description text and the combined measurement instruction and task support set data to obtain task-specific meta-mapper parameters adapted to the measurement task is as follows: Using the general meta-parameter θ of the meta-mapper as the initial parameter, the combined measurement instruction and task support set data is used... Calculate measurement error loss According to the formula Metamapper The parameters are updated through one or more internal loop gradient descent iterations, where α is the internal loop learning rate, ultimately yielding task-specific meta-mapper parameters optimized for the current measurement task and adapted to the measurement task. The process of determining the task-specific guiding visual prefix based on the task-specific meta-mapper parameters and structured description text adapted to the measurement task is as follows: This involves incorporating the task-specific meta-mapper parameters adapted to the measurement task. The meta-mapper, combined with facility information recorded in the structured description text, encodes and fuses the visual features of the image to be measured, which contain key regional features related to the measurement instructions, to generate a set of updated visual prefixes, namely, measurement task-specific guiding visual prefixes. This prefix contains both target appearance information and execution guidance instructions for this specific measurement. The process of inputting the measurement task-specific guiding visual prefix into a large-scale pre-trained language model to generate structured text of keypoint coordinates is as follows: The measurement task-specific guiding visual prefix... Feeding into a frozen large-scale pre-trained language model The model uses an autoregressive approach, based on the formula The process involves progressively generating text, outputting structured text containing the names and pixel coordinates of the starting and ending key points required for measurement, such as "Starting point: Center of the left bolt of the tension clamp, coordinates: [x1, y1]. Ending point: Outer edge of the vibration damper, coordinates: [x2, y2]". The process of extracting the pixel coordinates of the key points from this structured text involves: parsing and extracting the x and y coordinates of the starting and ending key points in the pixel coordinate system from the generated structured text; and calculating the Euclidean distance between the pixel coordinates of the key points to obtain the pixel distance values. The process is as follows: Based on the Euclidean distance calculation formula between two points, the pixel coordinates of the two extracted key points are calculated to obtain the precise pixel distance value between the two points; The process of integrating natural language measurement instructions, structured text of key point coordinates, and pixel distance values to obtain structured data for precise measurement of asymmetric targets is as follows: Integrate the natural language instructions for this measurement, the structured text recording key point information, and the calculated pixel distance values according to a unified standard to form a complete structured data object. This object completely records the input instructions, key point information, and final distance results of this measurement.
[0148] Step 104: Using a large-scale pre-trained language model, based on the high-dimensional deep visual feature sequence, structured association list, and task-specific meta-mapper parameters adapted to the new target, conduct context-aware vibration damper sequence distance calibration and output a continuous distance linked list of vibration dampers.
[0149] The continuous distance linked list of vibration dampers is an ordered data set organized in the form of a structured linked list, output from the context-aware vibration damper sequence distance calibration. The entire list is organized with a single transmission line or a single clamp system as an independent unit, arranged sequentially according to the actual installation order of the vibration dampers along the conductor. Each node entry in the linked list contains complete association information, specifically including: the unique device identifier of the current vibration damper (corresponding one-to-one with the device code in the structured association list), the number of the conductor segment to which it belongs and the line span identifier, pairing association information with adjacent vibration dampers (such as the device identifier of the previous adjacent vibration damper), the pixel distance value between adjacent vibration dampers calculated based on high-dimensional deep visual feature sequences, and the confidence information of the distance calculation. The first and last nodes of the linked list correspond to the vibration dampers at both ends of the line segment, respectively. It only includes continuous pairings within the same conductor / clamp system, excluding invalid pairings across spans or systems, and completely and orderly records the spacing distribution data of all adjacent vibration dampers within the conductor / clamp system.
[0150] It should be noted that the vibration dampers under the same conductor or clamp system are selected from the structured association list. Combined with the spatial information of the high-dimensional deep visual feature sequence, the vibration damper sequence visual features are obtained by sorting them according to the installation order and filtering invalid targets. The vibration damper sequence visual features are then concatenated with visual prefix parameters that can be learned online. The context encoding is completed by adapting the parameters of the task-specific meta-mapping algorithm to the new target, generating an updated visual prefix containing sequence position information. After being input into a large-scale pre-trained language model, the model generates the distance information of adjacent vibration dampers by combining the sequence context. Finally, the distances are organized into a continuous distance list of vibration dampers according to the installation order.
[0151] Furthermore, step 104 may include the following sub-steps:
[0152] S41. From the structured association list and the high-dimensional deep visual feature sequence, filter the vibration dampers attached to the same conductor or clamp system, sort them by spatial location and filter invalid targets through a preset height threshold to obtain the visual features of the vibration damper sequence.
[0153] S42. Concatenate the visual features of the vibration damper sequence with the visual prefix parameters that can be learned online to obtain the sequence concatenation features;
[0154] S43. Using the task-specific meta-mapper parameters adapted to the new target, contextual encoding is performed on the sequence splicing features to obtain an updated visual prefix rich in sequence context information.
[0155] S44. Input the updated visual prefix rich in sequence context information into a large-scale pre-trained language model, perform iterative inference of the distance between adjacent vibration dampers, and obtain the text of the distance measurement between each pair of adjacent vibration dampers.
[0156] S45. Organize the distance measurement text of each pair of adjacent vibration dampers into a continuous distance linked list of vibration dampers.
[0157] The text for measuring the distance between adjacent vibration dampers is structured text generated by a large-scale pre-trained language model in an autoregressive iterative manner. During the generation process, the text output of each step serves as the context for subsequent generation steps, ensuring that the measurement proceeds continuously along the installation sequence of the vibration dampers, ultimately forming an ordered set of texts that correspond one-to-one with the vibration damper sequence. The text is organized in the form of "one line corresponding to a pair of adjacent vibration dampers". Each line of text contains clear structured fields, specifically covering: the sequence number of the adjacent vibration damper pair (corresponding one-to-one with the sequence order, such as "pair 1 (damper 1-damper 2)" "pair 2 (damper 2-damper 3)"), the unique device ID of each vibration damper (corresponding one-to-one with the device identifier in the structured association list and the high-dimensional deep visual feature sequence to ensure data traceability), the pixel position coordinates of each vibration damper in the image, the guide segment / clamp system number to which it belongs, the pixel distance value between adjacent vibration dampers obtained by model inference, and the confidence value of this distance measurement.
[0158] It should be noted that this step first constructs a sequence of vibration dampers belonging to the same conductor based on the output structured association list. The system will filter out all vibration dampers attached to the same clamp system or conductor path and sort them according to their spatial position in the image, forming an ordered sequence of facilities. Simultaneously, the system introduces a preset "height threshold L" as a constraint to filter out invalid targets with excessive vertical position deviations due to image distortion or occlusion. The output of this stage is a collection containing the visual features and global contextual information of all vibration dampers in the sequence, providing a complete and ordered data foundation for subsequent sequence analysis.
[0159] Furthermore, to enable the model to understand the inherent relationships within the entire vibration damper sequence, rather than viewing individual facilities in isolation, the system employs a meta-mapper. Contextual encoding is performed on the visual features of the entire sequence. The meta-mapper encodes the visual features of all vibration dampers in the sequence, i.e., the visual features of the vibration damper sequence. With learnable visual prefixes The data is fused and processed using a self-attention mechanism. The core calculation formula is:
[0160] ;
[0161] In this application, input This represents the set of visual information for the entire vibration damper sequence. Through self-attention computation, the feature representation of each vibration damper in the sequence is influenced and weighted by the features of all other vibration dampers in the sequence. This allows the model to capture the overall pattern of the sequence, such as the regularity or anomaly of spacing, thereby generating a set of updated visual prefixes rich in sequence context information. .
[0162] Furthermore, the generated serialized visual prefix It was then fed into a frozen large-scale language model. This guides the model to perform iterative distance reasoning. The instruction received by the model is not just to calculate distance, but to "calculate the distance between adjacent vibration dampers sequentially along this sequence." The language model generates the measurement results between adjacent vibration dampers pairwise in an autoregressive manner, based on the overall visual context of the sequence. The generation process follows the formula:
[0163] ;
[0164] Here, the generated text is the text of distance measurements for each pair of adjacent vibration dampers. It will include the measurement results of the previous pair of vibration dampers, providing context for the current measurement, such as... distance (damper 2, hammer 3) is 15.3 meters. This iterative reasoning method ensures that the measurements are performed sequentially along the traverse, rather than through simple, unordered pairwise calculations, thus guaranteeing the logical consistency of the results.
[0165] Furthermore, after completing iterative measurements of all adjacent vibration dampers in the sequence, the system summarizes and organizes all generated measurement results. Ultimately, the output of this step is a structured "continuous distance linked list." This list is an ordered list where each element explicitly records the ID of a pair of adjacent vibration dampers, their respective positions, and the precise distance value calculated through model inference. This linked list is not only a collection of values but also a data structure reflecting the distribution order of vibration dampers along the conductor in the physical world, providing a direct and reliable basis for subsequent compliance analysis and anomaly detection.
[0166] In this embodiment, the process of selecting vibration dampers attached to the same conductor or clamp system from a structured association list and a high-dimensional deep visual feature sequence, sorting them by spatial position, and filtering out invalid targets through a preset height threshold to obtain the visual features of the vibration damper sequence is as follows: Based on the structured association list, vibration dampers attached to the same clamp system or conductor path are selected, sorted according to their spatial position in the image, and an invalid target with excessive vertical position deviation due to image distortion or occlusion is filtered out by introducing a preset height threshold L. The high-dimensional deep visual features of these valid vibration dampers are extracted to form the visual features of the vibration damper sequence. Subsequently, the visual features of the vibration damper sequence are concatenated sequentially with online learnable visual prefix parameters to obtain the sequence concatenation features, which are the input sequences of the meta-mapper. Next, using the task-specific metamapper parameters adapted to the new target, the metamapper formula is applied. Contextual encoding is applied to the concatenated features of the sequence. A self-attention mechanism is used to calculate and weight the similarity between feature vectors within the sequence, ensuring that the feature representation of each vibration damper is influenced by the features of all other vibration dampers in the sequence. This captures the overall pattern of the sequence and generates an updated visual prefix rich in sequence context information. Then, the updated visual prefixes, rich in sequence context information, are fed into the frozen, large-scale pre-trained language model. The model uses an autoregressive approach, based on the formula Combining the overall visual context of the sequence with the generated preceding measurement results, the distance between adjacent vibration dampers is calculated iteratively, generating a pair-by-pair distance measurement text containing the ID, position, and spacing information of each pair of adjacent vibration dampers. Finally, all the measurement results of each pair of adjacent vibration dampers are summarized and organized into an ordered structured continuous distance list of vibration dampers according to the actual installation order of the vibration dampers along the conductor. Each element in the list records the ID, position, and precise spacing value obtained through model inference for a pair of adjacent vibration dampers.
[0167] Step 105: Using a large-scale pre-trained language model, logical verification and data cleansing are performed based on structured description text, structured association lists, asymmetric target precise measurement structured data, and vibration damper continuous distance linked lists. A logical verification report and a clean dataset are then output.
[0168] It should be noted that, taking structured description text, structured association list, structured data of precise measurement of asymmetric targets, and continuous distance chain list of vibration dampers as input, the rationality of the asymmetric target measurement data and the distribution pattern of vibration damper spacing are verified by cross-comparing facility categories, locations and associations. Data contradictions and outliers are identified, a logical verification report is generated, and abnormal data is removed or corrected to obtain a clean dataset that conforms to physical association and distribution patterns.
[0169] Furthermore, step 105 may include the following sub-steps:
[0170] S51. Construct a complete dataset to be verified based on the structured description text, structured association list, structured data of precise measurement of asymmetric targets, and continuous distance chain of vibration damper.
[0171] S52. Based on the complete dataset to be verified, gradient updates and feature encoding are performed on the general meta-parameters of the meta-mapper to obtain a guiding visual prefix specific to the verification task.
[0172] S53. Input the verification task-specific guiding visual prefix into a large-scale pre-trained language model to generate logical verification text;
[0173] S54. Organize and summarize the content of the logic verification text to obtain a logic verification report;
[0174] S55. Based on the logical verification text, filter out contradictory data in the complete dataset to be verified to obtain a clean dataset.
[0175] The complete dataset to be verified is a structured knowledge graph that integrates structured descriptive text, structured association lists, structured data from precise measurements of asymmetric targets, and a continuous distance list of vibration dampers generated in the preceding steps. It covers the entire process of the current inspection scenario. The dataset is organized into independent units based on a single inspection scenario and contains four core data categories: First, structured descriptive text, which fully records the category, location, and status of each transmission facility in the inspection images; second, structured association lists, which explicitly store the pairing relationships between vibration dampers and clamps, their respective conductor segments, and spatial location matching information; third, structured data from precise measurements of asymmetric targets, including natural language measurement instructions, key point pixel coordinates, and Euclidean distance calculation results; and fourth, a continuous distance list of vibration dampers, which records the ID, location, and spacing values of adjacent vibration dampers according to the conductor installation order. These four data categories are cross-indexed using the facility's unique ID as the association key, forming a complete data chain covering facility identification, association pairing, target measurement, and sequence calibration, providing a full data foundation with global context for subsequent logical verification.
[0176] The logical verification text is a set of structured texts generated by a large-scale pre-trained language model in an autoregressive manner, based on visual prefixes specific to each verification task. Each text is organized in an ordered format of "single task, single entry," and each entry strictly adheres to a pre-defined standardized expression. The content is divided into two types: First, verification-passed text, such as "Verification passed: Data uniqueness check, all vibration damper IDs appear only once in the distance chain; continuity check, the start / end coordinates of adjacent vibration dampers correspond one-to-one; integrity check, vibration dampers other than the first and last in the sequence contain pre- and post-sequence connections." Second, verification-failed text, which must clearly indicate abnormal details, such as "Verification failed: Logical contradiction found, vibration damper ID_B08 appears repeatedly in the distance chain, corresponding to entries 3 and 7; continuity check failed, the end coordinate of entry 5 in the distance chain does not match the start coordinate of entry 6, and the deviation exceeds a pre-defined threshold."
[0177] The logic verification report is a standardized, structured report formed by classifying, sorting, and summarizing logic verification texts. It adopts a three-tier structure of "overview-items-appendix". The overview section records the inspection scene identifier, verification time, and overall verification conclusion (e.g., "verification passed / verification found 3 logic anomalies"). The itemized verification section is categorized by verification task. Each inspection item includes "task name, logic rule description, verification result, and abnormal data entry (if any)". The logic rule description clearly lists the verification basis, such as "uniqueness rule: the same vibration damper ID appears only once in the distance chain; continuity rule: the deviation between the endpoint coordinates of adjacent vibration dampers and the starting coordinates of the next one must be within a preset range; integrity rule: the vibration damper in the middle of the sequence must have both preceding and following pairings". The appendix section includes the original entries and location indexes of all abnormal data.
[0178] The clean dataset is a structured dataset obtained by filtering and correcting contradictory and anomalous data in the complete dataset to be verified. Its data organization format is consistent with the complete dataset, and all data has been freed from duplication, contradictions, and discontinuities, conforming to the physical association rules and distribution patterns of transmission lines. The dataset contains four types of data: structured descriptive text, structured association lists, structured data of precise measurements of asymmetric targets, and continuous distance lists of vibration dampers. These are still cross-indexed using the facility's unique ID as the association key, and all association relationships, measurement results, and sequence calibration data meet logical consistency requirements, such as no duplicate vibration damper IDs in the distance list, continuous coordinates of adjacent entries, and no missing sequence information.
[0179] It should be noted that the input to this step is all the analysis results generated in the previous steps, with the core being the continuous distance linked list output in step 4. To achieve deep logical verification, the system also integrates all intermediate data generated in the preceding steps, including facility identification information, established physical associations, and precise measurement results of asymmetric targets. These data collectively form a complete knowledge graph about the current inspection scenario, providing a sufficient data foundation for subsequent semantic and logical consistency checks based on global context.
[0180] Furthermore, the core of this step is no longer simply data deduplication or cleaning, but rather intelligent logical verification of the analysis results. The system checks whether the data conforms to the basic laws of power transmission lines in reality, such as the same vibration damper appearing only once in a sequence, the endpoint of the previous element in the linked list corresponding to the starting point of the next element, and each vibration damper having both a preceding and following sequence connection. These rules are not implemented through fixed code, but rather through a series of verification tasks mastered by the model via meta-learning, thereby automatically judging the rationality and completeness of the data.
[0181] Furthermore, to enable the model to perform verification tasks, the system specifically trains its logical consistency reasoning ability during the meta-training phase. For a given verification task... (For example, "checking uniqueness"), which means the complete dataset to be validated. The model will use a task support set that contains positive examples (logically correct data) and negative examples (logically incorrect data). Metamapper The parameters are rapidly adapted. This adaptation process follows an internal loop update rule:
[0182] ;
[0183] Among them, mission losses It measures the accuracy of the model in determining whether the data logic is correct. This is achieved through outer loop optimization on a large number of different types of validation tasks.
[0184] ;
[0185] The final meta-parameters obtained by the model This enables it to have a deep understanding of the inherent logic of data structures. During inference, the meta-mapper encodes the structured features of the dataset to be validated into a guiding visual prefix, i.e., a validation task-specific guiding visual prefix. This prefix can guide the language model to perform corresponding logical checks.
[0186] Furthermore, the visual prefixes generated in the previous step, which contain data structure features and verification task instructions, are then... Input to the frozen language model Guided by this, the language model generates a detailed validation report using an autoregressive approach. The generation process follows this formula:
[0187] ;
[0188] The output text, i.e., the logic verification text. The verification results will be clearly stated, such as "Verification passed: The uniqueness, continuity, and integrity of the data all meet the requirements," or "Verification failed: A logical error was found; the vibration damper ID_B08 appears repeatedly in the distance chain." The system will automatically filter or mark data entries with logical contradictions based on the verification report. The final output of this step is a "clean dataset" that has undergone deep semantic and logical verification, ensuring the accuracy, uniqueness, and logical consistency of all data, providing a reliable and high-quality data source for subsequent professional analysis reports.
[0189] In this embodiment, the process of constructing the complete dataset to be verified based on structured description text, structured association list, structured data of precise measurement of asymmetric targets, and continuous distance linked list of vibration dampers is as follows: Integrating all intermediate data generated in the preceding steps, including facility identification information, established physical associations, precise measurement results of asymmetric targets, and continuous distance linked list of vibration dampers, forms a complete knowledge graph about the current inspection scenario, serving as the basic data for logical verification. Based on the complete dataset to be verified, the process of performing gradient updates and feature encoding on the general meta-parameters of the meta-mapper to obtain the verification task-specific guiding visual prefix is as follows: Utilizing a task support set containing logically correct (positive examples) and logically incorrect (negative examples) data, for a given verification task... (e.g., checking uniqueness, continuity, and completeness), the formula is updated through an inner loop. Gradient updates are performed on the general meta-parameters of the meta-mapper, while optimization is achieved through the outer loop. The model is then given the ability to reason logically about the data. The structured features of the dataset to be validated are then encoded into a validation task-specific guiding visual prefix that contains both data structure features and validation task instructions. The process of inputting the verification task-specific guiding visual prefixes into a large-scale pre-trained language model to generate logical verification text is as follows: the model, guided by the visual prefixes, follows a formula in an autoregressive manner. The process of generating logical verification text, which clearly indicates the verification result, is as follows: "Verification passed: The uniqueness, continuity, and integrity of the data all meet the requirements" or "Verification failed: A logical error was found; the vibration damper ID_B08 appears repeatedly in the distance chain." The process of organizing and summarizing the logical verification text to obtain a logical verification report involves: integrating all verification results, categorizing them by verification task, and forming a standardized report containing check items, verification conclusions, abnormal data entries, and explanations of logical rules. These logical rules include that the same vibration damper appears only once in the sequence, that the endpoint of the previous element in the distance chain corresponds to the starting point of the next element, and that each vibration damper has a preorder and postorder connection. Based on the logical verification text, the process of filtering contradictory data in the complete dataset to obtain a clean dataset involves: filtering or correcting logically contradictory data entries marked in the logical verification text, ultimately obtaining a clean dataset where all data meets the requirements of accuracy, uniqueness, and logical consistency.
[0190] Step 106: Using a large-scale pre-trained language model, perform inspection analysis based on the logical verification report, clean dataset, natural language instructions for the report format, standard report example task support set, and task-specific meta-mapping parameters adapted to the new target, to obtain a comprehensive inspection analysis report.
[0191] The report format natural language instructions are instruction texts written in natural language. They are specifically used to define the production specifications of comprehensive inspection and analysis reports, clearly specify the overall framework of the report, the order of chapters, the required fields for each section, the data display format, the standard for the use of professional terminology, and the text layout requirements. They provide clear guidance on format and content for generating reports using large-scale pre-trained language models.
[0192] The Standard Report Example Task Support Set is a small-sample text dataset built for the task of generating transmission line inspection reports. It consists of multiple examples of compliant comprehensive inspection analysis reports under different inspection scenarios. Each example has a complete content structure, standard text descriptions, and standardized data presentation formats, which are used to help the model learn the report writing logic and format requirements and adapt to report generation related tasks.
[0193] The comprehensive inspection analysis report is the final official document output from the intelligent inspection process of transmission lines. It is compiled according to the preset format and compilation requirements, with the logic verification report and clean dataset as the core data sources. It fully includes basic information of the inspection scenario, facility identification results, relationship of auxiliary facilities, asymmetric target measurement data, distribution of vibration damper spacing, data logic verification conclusions, and explanation of abnormal situations, comprehensively presenting the results of this inspection.
[0194] It should be noted that the report compilation specifications are determined by combining the natural language instructions for the report format and the standard report example task support set. Data feature encoding is completed by relying on the task-specific meta-mapping parameters adapted to the new target. A large-scale pre-trained language model reads all the inspection information in the logical verification report and the clean dataset. All contents are integrated and sorted out according to the established specifications, and finally a comprehensive inspection analysis report is generated.
[0195] Furthermore, step 106 may include the following sub-steps:
[0196] S61. Combine the natural language instructions for report format and the standard report example task support set to obtain the report generation task instruction and example combination data;
[0197] S62. Extract global condensed data from the logic verification report and clean dataset;
[0198] S63. Using the task-specific meta-mapping parameters adapted to the new target, the global condensed data for report generation is encoded to obtain a dedicated guiding visual prefix for report generation.
[0199] S64. Input the report into a large-scale pre-trained language model to generate a comprehensive inspection and analysis report.
[0200] The report generates a globally condensed set of information extracted and integrated from the logical verification report and the clean data set. It includes information on identified facilities, verified physical connections, precise measurement data, continuous distance lists, and logical verification conclusions. It is a structured data that comprehensively summarizes the core findings of this inspection, providing factual evidence and global context for report generation.
[0201] The report generation task instruction and example combination data is a dataset that combines natural language instructions for report formatting with a standard report example task support set. Each report instruction corresponds to a set of report example samples of the same type, which clarifies the report format requirements and compilation logic, and provides guidance and reference for large-scale pre-trained language models to learn report generation tasks.
[0202] It should be noted that the goal of this step is to transform the clean dataset output from step 5 into a comprehensive report containing in-depth analysis and decision support information. The system first receives a natural language instruction regarding the final report's format and content, such as "generate an A-level inspection report including an equipment overview, distance list, abnormal spacing highlight alerts, and status assessment." The specific format and analytical logic of this report can itself be defined as a meta-learning task. By providing the model with a small number of qualified report examples as a task support set during the meta-training phase. The model can learn to generate analysis reports with specific styles and depths, demonstrating the high flexibility of this method in task customization.
[0203] Furthermore, to generate a comprehensive and accurate report, the system integrates all the clean dataset output from step 5, including but not limited to detailed information on identified facilities, verified physical associations, precise measurement data, and continuous distance lists. This fully validated dataset will serve as the sole factual basis and global context for generating the report. This structured data is then processed by the system to construct a highly condensed final input that comprehensively summarizes all the core findings of this inspection, preparing for subsequent guided text generation.
[0204] Furthermore, this is the core step in transforming structured data into natural language reports. The global data context integrated in the previous step is processed by the meta-mapper. The final encoding is transformed into a highly summarized guiding visual prefix containing all the key information. Subsequently, this prefix was fed into a frozen large-scale language model. The final comprehensive analysis report is generated using an autoregressive approach. Each term in the report is generated based on preceding content and the initial global context, and its mathematical expression is as follows:
[0205] ;
[0206] In this process, the model not only restates the data, but also uses its reasoning ability acquired through meta-learning to perform data comparison, standard judgment and anomaly identification, thereby proactively generating high-value analytical conclusions in the report, such as "Anomaly alarm: The distance between the 3rd and 4th anti-vibration hammers on Line 2 is 18.1 meters, exceeding the standard threshold of 16.0 meters, and displacement is suspected."
[0207] Finally, after text generation, the system outputs two final deliverables. The first is a readable comprehensive analysis report (such as a PDF or Word document) for operations and maintenance personnel and managers. This report is rich in visuals, logically clear, and explicitly identifies all potential risk points, providing direct support for subsequent maintenance decisions. The second is a machine-readable structured data archive (such as JSON or XML format), which fully records the entire process from raw image recognition to the final analysis conclusions, facilitating data storage, long-term tracking, and integration with other digital systems. Thus, the intelligent analysis process for power transmission line facilities based on UAV imagery forms a closed loop, achieving end-to-end transformation from raw visual information to actionable intelligent intelligence.
[0208] In this embodiment, the process of combining report format natural language instructions and standard report example task support sets to obtain report generation task instructions and example combination data is as follows: Report format natural language instructions that specify the report format, chapter structure, and required content (such as equipment overview, distance list, abnormal distance highlighting alarm, and status assessment), such as "generate an A-level inspection report containing equipment overview, distance list, abnormal distance highlighting alarm, and status assessment," are bound one-to-one with a standard report example task support set containing a small number of similar compliant report examples. Each report instruction is matched with a set of report example samples of the same style and format, forming report generation task instruction and example combination data, providing a data foundation for the model to learn report writing logic and format requirements; from logical verification reports, clean datasets... The process of extracting and generating globally condensed data for the report specifically involves: integrating all data from the cleanroom dataset, including detailed information on identified facilities, verified physical associations, precise measurement data, continuous distance lists, and verification conclusions and anomaly explanations from the logical verification report; processing this structured data into a comprehensive set of information that summarizes the core findings of this inspection, serving as the sole factual basis and global context for report generation; and using task-specific meta-mapping parameters adapted to the new target to encode the globally condensed data for report generation, resulting in a unique guiding visual prefix for the report. This process involves a meta-mapping tool equipped with task-specific meta-mapping parameters adapted to the new target encoding the globally condensed data for report generation, transforming the structured data into a highly summarized guiding visual prefix containing all key information. The process of generating a comprehensive inspection and analysis report by inputting a unique guiding visual prefix from the report into a large-scale pre-trained language model is as follows: The unique guiding visual prefix generated from the report is then fed into the frozen large-scale pre-trained language model. The model is based on the formula Text is generated in an autoregressive manner. During the generation process, based on the preceding content and the initial global context, the reasoning ability obtained by meta-learning is used to perform data comparison, standard judgment and anomaly identification. The output text includes analysis conclusions such as anomaly alarms and status assessments, and finally forms a readable comprehensive analysis report for operation and maintenance personnel, as well as a machine-readable structured data archive.
[0209] For comparison of technical effects, existing technologies can be used as a reference. With the continuous development and intelligent upgrading of power systems, substations, as important nodes in power grid operation, require equipment status monitoring and operational trend analysis as key technologies to ensure the safety and stability of the power grid. In recent years, with the advancement of sensor technology, drone inspection, and satellite remote sensing technology, a large amount of equipment operation data has been collected in real time, providing a foundation for intelligent operation and maintenance, fault prediction, and risk assessment.
[0210] In existing technologies, the condition monitoring of critical transmission line accessories such as vibration dampers, tension clamps, and suspension clamps mainly relies on target detection and geometric ranging methods based on UAV images. These methods extract equipment location and category using convolutional neural networks or traditional target detection algorithms, and calculate the distance between the center point or edge points of a rectangular frame to form a basic spacing list. While these methods can achieve basic equipment positioning and identification, they have several shortcomings, including limited ability to identify new or rare equipment, insufficient spatial logic judgment in complex crossing lines leading to erroneous associations, and reliance on simple threshold rules for anomaly identification in ranging results, making it difficult to handle multi-dimensional features or dynamic changes. Recent research has attempted to combine visual information with textual descriptions or historical maintenance data, using deep learning models to correct image distortion, identify equipment offsets and abnormal states, and generate inspection reports. This has improved identification accuracy to some extent, but still relies on a large number of training samples, making it difficult to quickly adapt to new or rare equipment. Furthermore, it lacks systematic solutions for distance measurement and anomaly detection. Therefore, existing technologies still have significant shortcomings in terms of small sample adaptability, complex line spatial logic judgment, continuous vibration damper ranging, and automatic identification of abnormal states, making it difficult to meet the actual needs of intelligent inspection and trend analysis.
[0211] Furthermore, with the increasing complexity of transmission line operation and maintenance tasks, traditional inspection methods have significant shortcomings in terms of efficiency, accuracy, and intelligence. Most existing technologies rely on manual inspection or single-modality deep learning detection models. These methods typically suffer from the following problems: First, for new or rare transmission facilities, a large number of labeled samples are required for retraining, lacking rapid adaptability and resulting in low inspection efficiency. Second, in environments with multiple parallel or intersecting lines, traditional methods rely solely on geometric distance or center point matching for equipment association, which is prone to incorrect pairings and makes it difficult to guarantee the accuracy of spatial logical relationships. Third, for asymmetric or structurally complex facilities, existing measurement methods rely on fixed algorithms or manual programming, lacking flexibility and scalability. In addition, the verification and screening of analysis results often rely on simple data deduplication or manual checks, failing to deeply verify logical consistency and semantic integrity. Finally, traditional inspection systems lack the ability to automatically generate actionable reports and structured archives from the analysis results, making it difficult for the data to directly support operation and maintenance decisions and long-term management.
[0212] To address the aforementioned issues, this invention proposes a method for inspecting transmission line facilities, aiming to: rapidly identify new facilities under limited sample conditions; accurately establish physical relationships between facilities in complex line environments; flexibly and accurately measure asymmetric or complex structural targets; perform deep semantic and logical verification of analysis results; and achieve end-to-end intelligent report generation and structured data archiving, thereby significantly improving the automation, intelligence, and data availability of inspections.
[0213] Specifically, addressing the problem that existing technologies heavily rely on massive amounts of labeled data and cannot quickly adapt to new facility models, this invention introduces a meta-learning mechanism. By training a lightweight meta-mapper, bridging the frozen large-scale visual and language models, it achieves "learning how to learn." When new or rare facilities are encountered during inspections, there is no need for large-scale retraining of the entire system. Only a small number of samples are needed as a support set, and the meta-mapper parameters can be quickly fine-tuned through internal loop updates, thereby achieving efficient identification of new facilities. This capability significantly shortens the deployment cycle of new equipment and greatly improves the flexibility and engineering practicality of the method.
[0214] To address the problem that existing technologies rely solely on geometric calculations in complex scenarios, leading to erroneous associations, this invention upgrades association judgment from two-dimensional distance matching to visual-spatial logical reasoning based on global context through a multimodal large model. The meta-mapper encodes local features and the global scene in the image into guiding visual prefixes, guiding the language model to understand the physical topological relationships of facilities, such as determining whether a vibration damper is attached to the same conductor, rather than relying solely on two-dimensional distance. This context- and logic-driven reasoning approach effectively reduces recognition errors in multi-line parallel or intersecting environments, improving the accuracy of facility association.
[0215] To address the limitations of existing technologies in measuring asymmetric targets due to their rigidity and limited accuracy, this invention employs a natural language command-driven measurement method. Operators can precisely define the measurement start and end points using text commands, and the model can accurately map these high-level semantic commands to image pixels, enabling precise measurement of complex structures. This method overcomes the constraints of fixed algorithms, offering flexibility and scalability, adapting to targets of various structural forms, and significantly improving measurement accuracy and automation.
[0216] To address the issues of fragmented processes and complex inter-module interfaces in existing technologies, this invention provides a unified end-to-end intelligent analysis framework. From facility identification, logical association, and precise measurement to final report generation, the entire process is completed in a closed loop within a single multimodal meta-learning architecture. This integrated design avoids interface mismatches and error accumulation between modules, making the system simpler and more efficient, while also enhancing stability and maintainability.
[0217] To address the shortcomings of existing technologies, such as limited output and lack of operability, this invention utilizes a language model to generate comprehensive analysis reports. The system integrates deeply validated structured data into detailed, human-readable reports, automatically identifies anomalies, and provides analytical suggestions, such as displacement warnings or review prompts. This transforms the analysis results into actionable intelligence with decision-making support value, significantly enhancing the practicality and decision support capabilities of intelligent inspection.
[0218] In this embodiment of the invention, a method for inspecting power transmission line facilities is provided. This method acquires original images of the inspection by a drone and images of the power transmission facilities. Based on these images, a high-dimensional deep visual feature sequence, structured descriptive text, and task-specific meta-mapper parameters adapted to the new target are generated. A large-scale pre-trained language model is used to generate a structured association list based on the structured descriptive text, the original drone inspection images, and the task-specific meta-mapper parameters adapted to the new target. The large-scale pre-trained language model then performs precise asymmetric target measurement based on the structured descriptive text, natural language measurement instructions, and a set of image-text example tasks, obtaining structured data for precise asymmetric target measurement. Finally, the large-scale pre-trained language model is used to perform further measurements based on the high-dimensional deep visual feature sequence, the structured association list, and the task-specific meta-mapper parameters adapted to the new target. The following steps involve calibrating the distance of the vibration damper sequence and outputting a continuous distance list of the vibration dampers. A large-scale pre-trained language model is used to perform logical verification and data cleansing based on structured descriptive text, structured association lists, structured data from precise measurements of asymmetric targets, and the continuous distance list of the vibration dampers, outputting a logical verification report and a clean dataset. Another large-scale pre-trained language model is used to perform inspection analysis based on the logical verification report, the clean dataset, natural language instructions for the report format, a standard report example task support set, and task-specific meta-mapping parameters adapted to the new target, resulting in a comprehensive inspection analysis report. Based on the above scheme, this invention generates task-specific meta-mapping parameters adapted to the new target using images of power transmission facilities. This eliminates the need for retraining the entire large-scale pre-trained language model and preparing massive amounts of labeled samples; the identification and adaptation of new power transmission facilities can be completed solely using dedicated meta-mapping parameters. In the steps of generating the structured association list and calibrating the distance of the vibration damper sequence, the scheme reuses the task-specific meta-mapping parameters adapted to the new target, flexibly adapting to different layouts and combinations of power transmission line ancillary facilities, without needing to design separate matching algorithms for differentiated line layouts. When performing precise measurements of asymmetric targets, the measurement operation is driven by a model that relies on natural language measurement commands combined with graphic and textual example task support sets. This allows for flexible definition of measurement tasks based on different structures of transmission facilities and varying on-site measurement requirements. The logic verification and data purification stages integrate various inspection data generated throughout the entire process, uniformly completing verification and filtering. Automated data processing is achieved using existing models and parameters, eliminating the need to customize verification rules for different types of inspection data and reducing the workload of rule adjustments during scenario switching. Finally, the inspection analysis and report generation stages utilize natural language commands for report formats and standard report example task support sets. This allows for the flexible generation of comprehensive inspection analysis reports with corresponding styles and analytical focuses, meeting diverse output requirements and effectively improving the flexibility of transmission line inspection methods.
[0219] Please see Figure 2 , Figure 2 This is a structural block diagram of a power transmission line facility inspection system provided in Embodiment 2 of the present invention.
[0220] This invention provides a power transmission line facility inspection system, comprising:
[0221] The acquisition module 201 is used to acquire the original images of the UAV inspection and the images of the power transmission facilities, and generate a high-dimensional deep visual feature sequence, structured descriptive text, and task-specific meta-mapper parameters adapted to the new target based on the original images of the UAV inspection and the images of the power transmission facilities.
[0222] The generation module 202 is used to generate a structured association list based on the structured description text, the original images of the UAV inspection, and the task-specific meta-mapping parameters adapted to the new target using a large-scale pre-trained language model.
[0223] Measurement module 203 is used to perform precise measurement of asymmetric targets by using a large-scale pre-trained language model based on structured description text, natural language measurement instructions and image-text example task support set, and obtain structured data for precise measurement of asymmetric targets;
[0224] The calibration module 204 is used to perform context-aware vibration damper sequence distance calibration using a large-scale pre-trained language model based on high-dimensional deep visual feature sequences, structured association lists, and task-specific meta-mapper parameters adapted to the new target, and output a continuous distance linked list of vibration dampers.
[0225] The verification module 205 is used to perform logical verification and data purification based on structured description text, structured association list, asymmetric target accurate measurement structured data and vibration damper continuous distance chain list using a large-scale pre-trained language model, and outputs logical verification report and clean dataset.
[0226] Analysis module 206 is used to perform inspection analysis using a large-scale pre-trained language model based on the logical verification report, clean dataset, natural language instructions for the report format, standard report example task support set, and task-specific meta-mapping parameters adapted to the new target, to obtain a comprehensive inspection analysis report.
[0227] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system and modules described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0228] This invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs the steps of the transmission line facility inspection method as described in the above embodiments.
[0229] This invention also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps of the transmission line facility inspection method as described in the above embodiments.
[0230] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0231] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0232] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0233] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0234] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for inspecting power transmission line facilities, characterized in that, include: Acquire raw images of UAV inspections and images of power transmission facilities, and generate high-dimensional deep visual feature sequences, structured descriptive text, and task-specific meta-mapper parameters adapted to new targets based on the raw images of UAV inspections and the images of power transmission facilities. A large-scale pre-trained language model is used to generate a structured association list based on the structured description text, the original images of the UAV inspection, and the task-specific meta-mapping parameters of the new target. The large-scale pre-trained language model performs asymmetric target precision measurement based on the structured description text, natural language measurement instructions, and image-text example task support set, and obtains asymmetric target precision measurement structured data. The large-scale pre-trained language model is used to perform context-aware vibration damper sequence distance calibration based on the high-dimensional deep visual feature sequence, the structured association list, and the task-specific meta-mapper parameters adapted to the new target, and outputs a continuous distance linked list of vibration dampers. The large-scale pre-trained language model is used to perform logical verification and data purification based on the structured description text, the structured association list, the asymmetric target precise measurement structured data, and the vibration damper continuous distance linked list, and outputs a logical verification report and a clean dataset. The large-scale pre-trained language model is used to perform inspection analysis based on the logical verification report, the clean dataset, the natural language instructions for the report format, the standard report example task support set, and the task-specific meta-mapping parameters for adapting to the new target, to obtain a comprehensive inspection analysis report.
2. The method for inspecting transmission line facilities according to claim 1, characterized in that, The step of generating a high-dimensional deep visual feature sequence, structured descriptive text, and task-specific meta-mapper parameters adapted to the new target based on the original images of the UAV inspection and the images of the power transmission facilities includes: The original images of the UAV inspection are pixel-vectorized to obtain a high-dimensional deep visual feature sequence; The high-dimensional deep visual feature sequence is concatenated with visual prefix parameters that can be learned online to obtain a concatenated feature sequence. The spliced feature sequence is subjected to feature focusing and semantic extraction to obtain the updated guiding visual prefix; The updated guiding visual prefix is subjected to autoregressive text generation processing to obtain structured descriptive text. The images of the power transmission facilities and the corresponding text annotations are used to form a facility recognition task support set; Based on the facility identification task support set, gradient update processing of the general meta-parameters of the meta-mapper is performed to obtain task-specific meta-mapper parameters adapted to the new target.
3. The method for inspecting transmission line facilities according to claim 1, characterized in that, The method employs a large-scale pre-trained language model to generate a structured association list based on the structured descriptive text, the original UAV inspection images, and the task-specific meta-mapping parameters for adapting to the new target. This list includes: Extract facility information from the structured description text and the original images of the UAV inspection; Based on the facility information and the original images from the UAV inspection, construct combined data of local facility information and global scene context; From the combined data of local facility information and global scene context, typical auxiliary facilities of transmission lines, namely vibration dampers and line clamps, are identified, and the two types of facilities are used as associated objects to generate a candidate pairing set of vibration dampers and line clamps. Using the task-specific meta-mapping parameters of the new target, visual feature encoding and rationality loss calculation are performed on the candidate pairs of vibration damper and wire clamp in the candidate pair set to obtain the exclusive guiding visual prefix of the candidate pairs of vibration damper and wire clamp. The dedicated guiding visual prefix is input into the large-scale pre-trained language model to generate a structured association description text of the vibration damper and the clamp. The structured association description text of the vibration damper-line clamp is standardized to obtain a structured association list.
4. The method for inspecting transmission line facilities according to claim 1, characterized in that, The process involves using the large-scale pre-trained language model to perform precise asymmetric target measurement based on the structured descriptive text, natural language measurement instructions, and image-text example task support set, resulting in structured data for precise asymmetric target measurement, including: The natural language measurement instructions and the image and text example task support set are combined to obtain combined measurement instruction and task support set data; Based on the structured description text and the combined data of the measurement instructions and task support set, the general meta-parameters of the meta-mapper are updated by gradient to obtain task-specific meta-mapper parameters adapted to the measurement task. Based on the task-specific meta-mapper parameters of the adapted measurement task and the structured description text, a measurement task-specific guiding visual prefix is determined; The measurement task-specific guiding visual prefix is input into the large-scale pre-trained language model to generate structured text of key point coordinates; Extract the pixel coordinates of key points from the structured text containing the key point coordinates; The pixel distance value is obtained by calculating the Euclidean distance to the pixel coordinates of the key points. By integrating the natural language measurement instructions, the structured text of the key point coordinates, and the pixel distance values, structured data for accurate measurement of asymmetric targets is obtained.
5. The method for inspecting transmission line facilities according to claim 1, characterized in that, The process involves using the large-scale pre-trained language model to perform context-aware distance calibration of the vibration damper sequence based on the high-dimensional deep visual feature sequence, the structured association list, and the task-specific meta-mapper parameters adapted to the new target, outputting a continuous distance linked list for the vibration damper, including: From the structured association list and the high-dimensional deep visual feature sequence, filter the vibration dampers attached to the same conductor or clamp system, sort them by spatial location and filter invalid targets through a preset height threshold to obtain the visual features of the vibration damper sequence. The visual features of the vibration damper sequence are concatenated with visual prefix parameters that can be learned online to obtain the sequence concatenation features; Using the task-specific meta-mapper parameters for adapting to the new target, the sequence splicing features are context-encoded to obtain an updated visual prefix rich in sequence context information; The updated visual prefix rich in sequence context information is input into the large-scale pre-trained language model to perform iterative inference of the distance between adjacent vibration dampers, thereby obtaining the text of the distance measurement between each pair of adjacent vibration dampers. The text describing the distance measurements of each pair of adjacent vibration dampers is organized into a linked list of continuous distances for the vibration dampers.
6. The method for inspecting transmission line facilities according to claim 1, characterized in that, The large-scale pre-trained language model is used to perform logical verification and data cleansing based on the structured description text, the structured association list, the asymmetric target precise measurement structured data, and the vibration damper continuous distance linked list, outputting a logical verification report and a clean dataset, including: Based on the structured description text, the structured association list, the structured data of precise measurement of the asymmetric target, and the continuous distance linked list of the vibration damper, construct a complete dataset to be verified; Based on the complete dataset to be verified, gradient updates and feature encoding are performed on the general meta-parameters of the meta-mapper to obtain a guiding visual prefix specific to the verification task. The verification task-specific guiding visual prefix is input into the large-scale pre-trained language model to generate logical verification text. The logical verification text is sorted and summarized to obtain a logical verification report; Based on the logical verification text, contradictory data in the complete dataset to be verified are filtered to obtain a clean dataset.
7. The method for inspecting transmission line facilities according to claim 1, characterized in that, The process employs the large-scale pre-trained language model to perform inspection analysis based on the logical verification report, the clean dataset, the report format natural language instructions, the standard report example task support set, and the task-specific meta-mapping parameters adapted to the new target, resulting in a comprehensive inspection analysis report, including: The report format natural language instructions and the standard report example task support set are combined to obtain report generation task instruction and example combination data; Global condensed data is generated by extracting data from the logical verification report and the clean dataset. Using the task-specific meta-mapper parameters adapted to the new target, the global condensed data for report generation is encoded to obtain a report generation-specific guiding visual prefix. The report generates a unique guiding visual prefix, which is then input into the large-scale pre-trained language model to generate a comprehensive inspection and analysis report.
8. A transmission line facility inspection system, characterized in that, include: The acquisition module is used to acquire the original images of the UAV inspection and the images of the power transmission facilities, and generate a high-dimensional deep visual feature sequence, structured descriptive text, and task-specific meta-mapper parameters adapted to the new target based on the original images of the UAV inspection and the images of the power transmission facilities. The generation module is used to generate a structured association list based on the structured description text, the original UAV inspection image, and the task-specific meta-mapping parameters of the new target using a large-scale pre-trained language model; The measurement module is used to perform precise measurement of asymmetric targets using the large-scale pre-trained language model based on the structured description text, natural language measurement instructions, and image-text example task support set, to obtain structured data for precise measurement of asymmetric targets. The calibration module is used to perform context-aware vibration damper sequence distance calibration using the large-scale pre-trained language model based on the high-dimensional deep visual feature sequence, the structured association list, and the task-specific meta-mapper parameters of the new target, and output a continuous distance linked list of vibration dampers. The verification module is used to perform logical verification and data purification based on the structured description text, the structured association list, the asymmetric target precise measurement structured data, and the vibration damper continuous distance linked list using the large-scale pre-trained language model, and outputs a logical verification report and a clean dataset. The analysis module is used to perform inspection analysis using the large-scale pre-trained language model based on the logical verification report, the clean dataset, the natural language instructions for the report format, the standard report example task support set, and the task-specific meta-mapping parameters for the new target, to obtain a comprehensive inspection analysis report.
9. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the transmission line facility inspection method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the transmission line facility inspection method as described in any one of claims 1-7.