A power transmission line infrared-visible light multi-modal target detection and warning method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-08-11
AI Technical Summary
[0011]针对现有技术存在的上述不足,本发明的目的在于,提供一种输电线路红外-可见光多模态目标检测与告警方法,能够在复杂光照条件、低纹理红外图像、小尺度目标和新类别缺陷出现时保持检测稳定性,同时减少不符合运行规律的误报
[0025]1. 本方案通过同步采集可见光图像、红外图像以及气象信息、线路台账和负荷电流等多维度数据,实现了输电线路状态的全方位感知。相较于现有技术中仅依赖单一图像模态或简单叠加多源信息的方式,本方案为后续的智能分析提供了更加丰富和准确的数据基础,显著提升了系统对复杂环境的适应能力。
Smart Images

Figure CN122550902A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent inspection technology for power systems, and in particular to a method for infrared-visible multimodal target detection and alarm for transmission lines. Background Technology
[0002] Drones, fixed cameras, inspection robots, and vehicle-mounted equipment are now widely used in power transmission line inspections. Visible light images are suitable for identifying external damage, foreign objects, tree obstructions, and the condition of structural components, while infrared images are suitable for detecting thermal anomalies such as joints, clamps, insulators, and localized overheating of conductors. In field use, both types of images are often acquired simultaneously, but in many systems, they are still processed separately, with the results ultimately merged manually. This approach works in clear weather with clear targets, but it significantly increases the likelihood of missed detections and false alarms when encountering backlighting, fog, nighttime conditions, heat reflection, thin conductors, target obstruction, or complex backgrounds.
[0003] The existing target detection schemes mainly have the following problems.
[0004] First, the detection results of a single mode are unstable. Visible light is easily affected by illumination, shadows, and motion blur, while infrared images have low resolution and little texture, making it easy to identify cloud reflections, metal fittings exposed to sunlight, and background heat sources as defects.
[0005] Second, infrared and visible light were not reliably registered. Drone gimbal rotation, binocular mounting errors, focal length differences, and parallax can cause the same hardware to appear in different positions in two images. If the two detection frames are simply hard-matched based on timestamps, duplicate alarms for the same target, mismatch alarms, or hotspot position shifts will occur.
[0006] Third, the detection network only learns from the external characteristics and does not understand the operational constraints of the transmission line. For example, conductor sag is related to span, tension, temperature, icing, and wind speed; hot spots in clamps must also be judged in conjunction with load current, ambient temperature, and heat dissipation conditions. Drawing conclusions solely based on image confidence levels can easily lead to physically unreasonable results being used as alarms.
[0007] Fourth, the ability to detect new defects in open scenarios is insufficient. Construction machinery, kite strings, plastic film, bird nesting materials, temporary facilities, and other non-fixed-category targets may appear on the construction site. Traditional closed-set detectors require labeling before training, which cannot effectively handle targets that are "not seen in the training set but are indeed dangerous on-site."
[0008] Fifth, the system implementation chain is incomplete. Some solutions remain at offline identification, lacking mechanisms for edge inference, cloud verification, expert rules, work order linkage, data feedback, and model updates. In actual operation and maintenance, what is needed on-site is a closed-loop system that can operate stably during inspection tasks, explain the basis of alarms, and push high-risk points to personnel for verification.
[0009] In summary, existing power transmission line inspection technologies have significant shortcomings in multimodal image registration, physical constraint fusion, open scene adaptability, and closed-loop system operation, making it difficult to meet the requirements for stable and reliable target detection and alarm in complex environments.
[0010] Therefore, how to maintain detection stability under complex lighting conditions, low-texture infrared images, small-scale targets, and the emergence of new types of defects, while reducing false alarms that do not conform to operational patterns, and transforming alarm results into actionable maintenance recommendations, has become an urgent problem to be solved. Summary of the Invention
[0011] To address the aforementioned shortcomings of existing technologies, the present invention aims to provide a method for infrared-visible multimodal target detection and alarm in power transmission lines, which can maintain detection stability under complex lighting conditions, low-texture infrared images, small-scale targets, and the appearance of new types of defects, while reducing false alarms that do not conform to operating patterns.
[0012] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0013] A method for infrared-visible multimodal target detection and alarm for power transmission lines includes the following steps:
[0014] S1. Collect visible light images and infrared images of transmission lines, as well as corresponding meteorological information, line records, and load current information;
[0015] S2. Based on the visible light image and infrared image of S1, initial registration is performed using offline calibration parameters to obtain the initial registration result; based on the features related to the transmission line structure extracted from the visible light image, the initial registration result is locally corrected online, and the infrared image is registered to the visible light image coordinate system to generate the registered infrared image.
[0016] S3. Perform multi-layer feature extraction on the visible light image and the registered infrared image from S2 respectively to obtain visible light features and infrared features; calculate the quality reliability of the visible light mode and the infrared mode based on the visible light image and the registered infrared image; generate fusion weights based on the quality reliability, and perform weighted fusion of the visible light features and infrared features. At the same time, introduce cross-modal deformation attention during the fusion process to compensate for feature shifts caused by incomplete local registration, and obtain fused features.
[0017] S4. Input the fused features from S3 into the closed-set real-time detection network to detect preset fixed-category targets and output a first detection result including target detection box, category, confidence level, segmentation mask, and surface temperature statistics. When the confidence level of the first detection result is lower than a preset threshold, or the category of the target is unknown, trigger the open vocabulary detection model to verify the corresponding target. The open vocabulary detection model is used to identify targets not included in the preset fixed categories and output a second detection result including the possible categories and confidence levels of the target.
[0018] S5. Based on the first or second detection result obtained in S4, as well as the corresponding meteorological information, line ledger and load current information, construct the conductor heat balance equation and conductor sag calculation model, and calculate the physical residual, which includes the heat balance residual and the sag residual.
[0019] S6. Based on the first or second detection result obtained in S4 and the physical residual obtained in S5, calculate the risk score for each target in combination with the preset expert rules.
[0020] S7. Based on the risk score described in S6, schedule the corresponding operation;
[0021] When the risk score is higher than the first threshold, an alarm work order is generated;
[0022] When the risk score is in the range of below the first threshold and above the second threshold, or when the category of the target in S4 is unknown, the open vocabulary detection model is triggered for verification.
[0023] When the risk score is lower than the second threshold, and the confidence level of the first detection result is lower than the preset threshold, and / or the physical residual exceeds the preset physical reasonableness threshold, the corresponding sample is added to the labeling queue; and the manually confirmed conclusion and the samples in the labeling queue are fed back to the cloud training platform.
[0024] Compared with the prior art, the present invention has the following advantages:
[0025] 1. This solution achieves comprehensive perception of the transmission line status by simultaneously acquiring visible light images, infrared images, meteorological information, line records, and load current data from multiple dimensions. Compared to existing technologies that rely solely on a single image modality or simply overlay multi-source information, this solution provides a richer and more accurate data foundation for subsequent intelligent analysis, significantly improving the system's adaptability to complex environments.
[0026] 2. A registration strategy combining offline calibration parameters with online local correction effectively solves the image registration deviation problems caused by factors such as UAV gimbal rotation, binocular installation errors, focal length differences, and parallax. Compared with existing technologies that rely solely on hard matching of timestamps or simple geometric transformations, this solution achieves sub-pixel level registration accuracy, avoiding misjudgments such as repeated alarms for the same target, mismatched detection boxes, or hotspot position shifts.
[0027] 3. By calculating the quality reliability of the visible light and infrared modes and generating dynamic fusion weights based on this, the optimal combination of features from different modes is achieved. Simultaneously, a cross-modal deformation attention mechanism is introduced to effectively compensate for feature shifts caused by incomplete local registration. Compared to the simple feature stitching or fixed-weight fusion methods in existing technologies, this scheme maintains stable detection performance under challenging scenarios such as complex lighting, haze, and nighttime conditions, significantly reducing false negatives and missed detections caused by the limitations of a single mode.
[0028] 4. A two-stage detection mechanism combining a closed-set real-time detection network and an open-vocabulary detection model is adopted. This ensures efficient detection of pre-defined fixed-category targets while also enabling the discovery of defects in new categories. Compared with existing technologies that only use closed-set detectors or require pre-labeled training, this solution can effectively handle non-pre-defined category targets such as construction machinery, kite strings, and plastic films, achieving intelligent identification of targets that are "not seen in the training set but are indeed dangerous on-site."
[0029] 5. By constructing a conductor thermal balance equation and a conductor sag calculation model, the detection results are combined with the physical laws of line operation to calculate physical indicators such as thermal balance residuals and sag residuals. Compared with existing technologies that rely solely on image confidence for decision-making, this solution can effectively filter physically unreasonable results, avoiding misjudging cloud reflections, sun-irradiated hardware, etc., as defects, significantly improving the accuracy and reliability of alarms.
[0030] 6. By combining the detection results and physical residuals, a risk score for each target is calculated using pre-defined expert rules, enabling intelligent evaluation and classification of the detection results. Compared to the simple threshold judgment in existing technologies, this solution comprehensively considers multiple factors, providing more scientific and reasonable risk assessment results, and offering strong support for subsequent operation and maintenance decisions.
[0031] 7. Based on risk scoring, intelligent scheduling operations such as alarm work order generation, open vocabulary review, and sample feedback for annotation are implemented, constructing a complete system implementation chain. Compared with existing offline recognition solutions that lack a closed-loop mechanism, this solution achieves the organic unity of edge inference, cloud review, expert rules, work order linkage, data feedback, and model updates, forming a continuously optimized intelligent inspection system.
[0032] In summary, this method maintains detection stability under complex lighting conditions, low-texture infrared images, small-scale targets, and the emergence of new types of defects, while reducing false alarms that do not conform to operational patterns. Through the organic combination of multimodal data fusion, physical constraint modeling, expert rule evaluation, and intelligent scheduling mechanisms, this solution achieves full-process intelligent management of transmission line inspection from data acquisition to operation and maintenance decision-making, significantly improving the system's practicality, reliability, and maintainability.
[0033] Preferably, in step S2, the initial registration using offline calibration parameters includes:
[0034] Using the homography transformation matrix from infrared image to visible light image The infrared image is resampled to achieve initial registration.
[0035] ;
[0036] In the formula, This is the intrinsic parameter matrix of the visible light camera; This is the intrinsic parameter matrix of the infrared camera; This is the rotation matrix from the infrared coordinate system to the visible light coordinate system; This is the translation vector from the infrared coordinate system to the visible light coordinate system; This is the normal vector of the local reference plane; This is the distance from the camera's optical center to the local reference plane; Visible light image coordinates; For the first Frame infrared image; To register the infrared image to the visible light coordinate system;
[0037] The online local correction of the initial registration result based on features related to the transmission line structure extracted from visible light images includes:
[0038] From the visible light image, at least one of the following is extracted as the feature related to the transmission line structure: conductor edge, tower corner point, insulator string centerline, and hardware outline;
[0039] The extracted features are processed to obtain the corresponding sparse keypoints and guide curves;
[0040] Based on the spatial geometric constraints formed by the sparse keypoints and the guide curves, the initial registration result is locally deformed and corrected within the region to be corrected and its surrounding preset pixel width, as determined by the sparse keypoints and the guide curves.
[0041] Compared to existing registration methods that rely on global homography or general feature matching, this scheme significantly improves the registration accuracy and stability of infrared-visible images in power transmission line scenarios. Especially when conductors are thin, hardware textures are incomplete, backgrounds are cluttered, or thermal reflection interference exists, it can still accurately align thermal anomaly areas with corresponding structural components, avoiding misjudgments due to registration misalignment. This lays a reliable spatial consistency foundation for subsequent multimodal fusion detection and physical rationality verification. It effectively solves the image spatial mismatch problem caused by modal differences and structural characteristics in power transmission line inspection, enabling precise association of thermal anomaly localization with specific equipment components, greatly enhancing the interpretability and engineering applicability of the detection results.
[0042] Preferably, in step S3, the quality reliability of the visible light mode and the infrared mode is calculated. This can be achieved in the following ways:
[0043] ;
[0044] In the formula, σ is the Logistic function; For modal identification, v and r represent the visible light mode and infrared mode, respectively; Image frame number; This is a quality feature vector extracted based on the exposure, sharpness, infrared temperature range, registration residual, and occlusion ratio of the visible light image and the registered infrared image. This is the quality reliability weight vector; This is a quality reliability bias.
[0045] Compared to existing fusion strategies that use preset thresholds or manually designed weights, this solution can adaptively handle complex inspection scenarios. For example, it automatically reduces the weight of the visible light mode in low-light conditions at night, suppresses infrared contributions when fog and haze cause infrared blurring, and increases the infrared weight when the conductor is overheated but visible light is interfered with by strong reflections. This ensures the robustness of the fused features even when modal quality degrades, avoiding a sharp drop in detection performance caused by the failure of a single mode. This provides a more reliable foundation for subsequent high-precision target detection and physical consistency verification. It achieves a leap from static weighting to perception-driven dynamic adaptation in multimodal information fusion, improving the system's stability and generalization ability under non-ideal imaging conditions.
[0046] Preferably, in step S3, the process of obtaining the fusion features includes:
[0047] For each spatial location u of each feature layer, based on the visible light modal quality reliability and infrared modal quality reliability and visible light local confidence map and infrared partial confidence map Calculate visible light fusion weights Infrared fusion weights :
[0048] ;
[0049] Then, the fusion features are generated using the following formula. :
[0050] ;
[0051] In the formula, h is the feature level; n is the image frame number; and denoted as visible light features and infrared features at position u, respectively; Φ is a cross-modal deformation attention term used to dynamically adjust feature alignment based on the local similarity between visible light features and infrared features, in order to compensate for feature offset caused by incomplete local registration.
[0052] Compared to existing fusion methods that rely solely on global weights or simple stitching, this approach effectively mitigates feature misalignment caused by local registration errors. While maintaining the complementary advantages of multimodal approaches, it significantly suppresses fusion artifacts and positioning drift, enabling subsequent detection networks to more accurately associate thermal anomaly regions with their corresponding physical structural components. This improves the detection accuracy of small targets and the reliability of hot spot positioning. It represents a leap from coarse-grained modality selection to fine-grained spatial adaptive alignment fusion, providing crucial support for high-precision multimodal sensing in complex power transmission scenarios.
[0053] Preferably, the line ledger includes the line name, tower number, phase, span, and conductor type;
[0054] In step S5, the thermal balance equation for the conductor is:
[0055] ;
[0056] in, ;
[0057] ;
[0058] ;
[0059] ;
[0060] In the formula, The equivalent heat capacity of the conductor; This is the derivative of the conductor temperature with respect to time. For the temperature of the conductor; It is Joule fever; Absorbing heat from the sun; For convection cooling; For radiative heat dissipation; This refers to the line load current. For the wire at temperature The resistance below; Solar absorptivity; Solar radiation intensity; The convective heat transfer coefficient is related to wind speed; Ambient temperature; Emissivity of the conductor surface; It is the Stefan-Boltzmann constant;
[0061] The thermal balance residual Calculated using the following formula:
[0062] .
[0063] Compared to traditional black-box alarm logic based solely on image temperature values, this solution effectively identifies and eliminates abnormal detection results that do not conform to the thermodynamic laws of conductors. For example, high-temperature points appearing under low load and high wind conditions may actually be due to sunlight reflection or background interference. Simultaneously, it can detect hidden risks, such as atypical temperature rises caused by abnormally high resistance. This significantly improves the physical rationality and engineering reliability of alarms, avoids invalid work orders caused by environmental disturbances or imaging artifacts, and enhances the system's robust decision-making capabilities in real-world operating scenarios. It provides a traceable, explainable, and verifiable physical foundation for intelligent inspection.
[0064] Preferably, in step S5, the conductor sag calculation model is as follows:
[0065] ;
[0066] In the formula, The theoretical value of conductor sag is calculated based on the aforementioned line ledger, load current, and meteorological information. For gear distance; The total load per unit length of the conductor; The horizontal tension of the conductor; This is the load per unit length corresponding to the self-weight of the conductor; This represents the load per unit length corresponding to icing. This represents the load per unit length corresponding to the wind load.
[0067] The sag residual Calculated using the following formula:
[0068] ;
[0069] In the formula, These are observations of conductor sag obtained through image measurement or sensors.
[0070] Compared to existing technologies that only compare historical images or set fixed sag thresholds, this approach effectively distinguishes between normal sag changes caused by environmental disturbances and abnormal deformations caused by structural defects such as loose fittings or broken conductor strands. This significantly reduces false alarm rates and provides early warnings of potential mechanical hazards. For example, when the sag residual is consistently positively biased and exceeds the expected range of thermo-mechanical coupling, it indicates abnormal conductor tension or support failure risk, thereby improving the accuracy and foresight of transmission line mechanical condition assessment.
[0071] Preferably, in step S6, the risk score is calculated using the following formula:
[0072] ;
[0073] in, ;
[0074] ;
[0075] ;
[0076] In the formula, For the first Risk score for each objective; Risk scoring bias; , , , , σ represents the risk score weights; σ is the Logistic function. The confidence level for detection in step S4; For the first The temperature rise of the target relative to the ambient temperature; For expert rule hit strength; For normalized physical residuals; For the first The net distance from a target to the nearest wire or live component; For the first The surface temperature statistics of a target in an infrared image can be taken as the highest temperature, the average temperature, or the quantile temperature. A set of expert rules; For expert rule serial numbers; For the first Membership function of expert rules; This is the normalization constant for the thermal equilibrium residuals; The normalization constant for the sag residual; The category of the target.
[0077] Compared to risk assessments that rely solely on a single indicator or a black-box deep classifier, this approach significantly improves the causal rationality and anti-interference capability of risk judgment. For example, when a hotspot has a high temperature but the thermal balance residual is close to 0 and meets the expert rule of high load + low wind speed, the system assigns a high-risk score. Conversely, if the temperature is abnormal but the residual is significantly negative and there is no rule to support it, the system automatically downgrades the weighting to avoid misjudging sunlight reflection or background heat sources as defects. At the same time, by dynamically adjusting the contribution of each evidence source through learnable weights, the system adapts to different line types and operating conditions, making the risk ranking more aligned with actual maintenance priorities.
[0078] Preferably, in step S6, calculating the risk score for each target using preset expert rules includes:
[0079] Based on the target category obtained from S4, the expert rule corresponding to that category is applied for judgment, where:
[0080] For targets identified as clamps or connectors, determine whether they meet the following criteria: the overlap between the infrared hot spot segmentation mask and the target segmentation mask exceeds the preset overlap threshold, and the target temperature rise exceeds the temperature rise of normal components in the same or adjacent phases.
[0081] For targets identified as channel foreign objects, determine whether they meet the following conditions: the target is located in the conductor protection zone or crossing zone, and the net distance from the target to the nearest live component is lower than the preset safe net distance threshold.
[0082] For targets identified as insulators, determine whether they meet the following criteria: the temperature difference distribution between individual pieces within the insulator string exceeds the preset temperature difference pattern, and exclude false judgments caused by single-frame reflections.
[0083] The more judgment conditions of the target that meet its corresponding rule, or the higher the degree of satisfaction, the higher the calculated expert rule hit strength; the expert rule hit strength is used as an input parameter to calculate the risk score in step S6.
[0084] Compared to risk assessment using uniform rules or simple temperature / distance thresholds, this approach significantly improves the category specificity and engineering adaptability of defect identification. For example, it avoids misjudging uniform temperature rises in insulator strings caused by contamination as abnormal hotspots, while accurately capturing localized overheating caused by loose clamps. This maintains a high detection rate while greatly reducing cross-category false alarms, making the rule output more aligned with the cognitive logic and handling priorities of on-site maintenance personnel.
[0085] Preferably, in step S4, when the open vocabulary detection model is triggered for verification, based on the detection box of the corresponding target output in S4, a local visible light image and a local infrared image containing the target are cropped from the visible light image and the registered infrared image, respectively, and a prompt word containing the category and / or location information of the target is generated; then the local visible light image, the local infrared image and the prompt word are uploaded to the cloud, and the open vocabulary detection model deployed in the cloud is used for recognition;
[0086] In step S7, after the manually confirmed conclusions and the samples in the queue to be labeled are fed back to the cloud training platform, the cloud training platform uses the fed-back samples to retrain and update the closed set real-time detection network and / or the open vocabulary detection model, and then sends the updated model to the edge device.
[0087] Compared to statically deployed or periodically offline updated detection systems, this approach significantly improves the model's adaptability to various scenarios and its ability to cover long-tail defects during long-term operation. For example, when a novel composite insulator exhibits atypical thermal characteristics due to material aging, the initial model may miss or misjudge the problem. After manual annotation and data retransmission, the model can learn the pattern within hours and accurately identify it in subsequent inspections, preventing similar problems from recurring. Furthermore, by only transmitting manually confirmed samples rather than the full dataset, privacy and transmission efficiency are balanced, ensuring accurate, efficient, and sustainable model updates.
[0088] Preferably, the closed-set real-time detection network minimizes the total loss function. During training, the total loss function for:
[0089] ;
[0090] In the formula, To detect the loss; For cross-modal consistency loss; For registration loss; Loss due to physical constraints; Loss of expert consensus; , , , , These are the weights corresponding to the loss;
[0091] in, ;
[0092] In the formula, For the first The set of valid pixels in the frame that participate in the registration loss calculation; This represents the number of pixels in the set. Huber loss function; Operators for describing visible light edges or structures; For infrared edge or structure description operators; For the first Frame of visible light images; To register the infrared image to the visible light coordinate system; These are the coordinates of the visible light image.
[0093] ;
[0094] In the formula, The number of time samples used in physical constraint training or verification; For thermal equilibrium residuals; The weighting is the sag residual. Weighting for exceeding limits; For sag residuals; This is the maximum permissible sag for this span of conductor; This is the theoretical value for conductor sag.
[0095] This setup enables the multi-constraint joint loss mechanism to upgrade the modeling paradigm from fitting observation data to following physical laws, providing a solid and reliable technical foundation for power vision perception systems with high safety requirements. Attached Figure Description
[0096] To make the objectives, technical solutions, and advantages of the invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:
[0097] Figure 1 This is a flowchart of the method;
[0098] Figure 2 This is a system framework diagram illustrating the implementation of the method in Example 1;
[0099] Figure 3 This is a flowchart of the infrared-visible light registration process in Example 1;
[0100] Figure 4 This is a schematic diagram of the infrared-visible light reliability gating fusion network mechanism in Example 1;
[0101] Figure 5 This is a flowchart illustrating the real-time detection of closed sets and the verification of open vocabularies in Example 1.
[0102] Figure 6 This is a schematic diagram of the physical constraint verification flow in S5 of Example 1;
[0103] Figure 7 This is a flowchart of the expert model and intelligent agent closed loop in Example 1. Detailed Implementation
[0104] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0105] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to represent selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0106] It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship commonly used when the product is in use. They are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance. In addition, the terms "horizontal," "vertical," etc., do not indicate that the component is required to be absolutely horizontal or suspended, but can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but can be slightly tilted. In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0107] Example 1
[0108] like Figure 1 As shown, this invention provides a method for infrared-visible multimodal target detection and alarm for power transmission lines. In specific implementation, the system framework used in this method is as follows: Figure 2 As shown.
[0109] This method includes the following steps:
[0110] S1. Collect visible light images and infrared images of the transmission line, as well as corresponding meteorological information, line ledgers, and load current information; wherein, the line ledger includes the line name, tower number, phase, span, and conductor type.
[0111] In practice, during the same inspection task, visible light images, infrared images, shooting time, drone or fixed point position and orientation, ranging information, meteorological information, line ledgers, and load current information are collected simultaneously. The line ledger must include at least the line name, voltage level, tower number, phase, span, conductor type, hardware type, insulator string type, corridor risk zone, and historical defect records. The data acquisition equipment can be a dual-light pan-tilt unit or a separate unit consisting of a visible light camera and an infrared thermal imager.
[0112] To facilitate subsequent processing, a unified index is created for each frame of data, recording the image, temperature matrix, equipment parameters, shooting point, and line component information. This approach ensures that the detection box can be directly associated with "a specific line, tower, phase, and component," avoiding the output of merely a single image coordinate without any operational or maintenance meaning.
[0113] S2. Based on the visible light image and infrared image of S1, initial registration is performed using offline calibration parameters to obtain the initial registration result. Based on the features related to the transmission line structure extracted from the visible light image, the initial registration result is locally corrected online, and the infrared image is registered to the visible light image coordinate system to generate the registered infrared image.
[0114] The infrared-visible light registration process is as follows: Figure 3 As shown.
[0115] In specific implementation, the initial registration using offline calibration parameters includes:
[0116] Using the homography transformation matrix from infrared image to visible light image The infrared image is resampled to achieve initial registration.
[0117] ;
[0118] In the formula, This is the intrinsic parameter matrix of the visible light camera; This is the intrinsic parameter matrix of the infrared camera; This is the rotation matrix from the infrared coordinate system to the visible light coordinate system; This is the translation vector from the infrared coordinate system to the visible light coordinate system; This is the normal vector of the local reference plane; This is the distance from the camera's optical center to the local reference plane; Visible light image coordinates; For the first Frame infrared image; To register the infrared image to the visible light coordinate system;
[0119] The online local correction of the initial registration result based on features related to the transmission line structure extracted from visible light images includes:
[0120] From the visible light image, at least one of the following is extracted as the feature related to the transmission line structure: conductor edge, tower corner point, insulator string centerline, and hardware outline;
[0121] The extracted features are processed to obtain the corresponding sparse keypoints and guide curves;
[0122] Based on the spatial geometric constraints formed by the sparse keypoints and the guide curves, the initial registration result is locally deformed and corrected within the region to be corrected and its surrounding preset pixel width, as determined by the sparse keypoints and the guide curves.
[0123] In practice, if the target on site is not planar, the system uses the homography matrix as the initial value and then uses sparse key points and traverse curve constraints to perform local deformation correction.
[0124] In practice, regarding the registration method, a fixed mapping after one calibration can be used for fixed cameras; for UAVs, joint estimation of gimbal attitude, ranging and visual key points can be used; for close-up views of towers, block monography or deformable registration can be used; and for equipment with lidar, point cloud projection can be used to generate more accurate cross-modal correspondences.
[0125] S3. Perform multi-layer feature extraction on the visible light image and the registered infrared image from S2 respectively to obtain visible light features and infrared features; calculate the quality reliability of the visible light mode and the infrared mode based on the visible light image and the registered infrared image; generate fusion weights based on the quality reliability, and perform weighted fusion of the visible light features and infrared features. At the same time, introduce cross-modal deformation attention during the fusion process to compensate for feature shifts caused by incomplete local registration, and obtain fused features.
[0126] In practical implementation, the infrared-visible light reliability gating fusion network mechanism (i.e., the fusion in method S3) is as follows: Figure 4 As shown.
[0127] In practice, multi-layer features are extracted from both the visible light image and the registered infrared image. Edge devices can employ lightweight convolutional networks, lightweight visual Transformers, or detection Transformer backbone networks; the cloud can utilize self-supervised visual backbones and larger multimodal backbones. The features are represented as follows:
[0128] ;
[0129] In the formula, It has visible light characteristics; Infrared characteristics; For feature level; For the visible light feature extractor in the first Layer mapping; For the infrared feature extractor in the first Layer mapping.
[0130] In practical implementation, the quality reliability of the visible light mode and the infrared mode is calculated. This can be achieved in the following ways:
[0131] ;
[0132] In the formula, σ is the Logistic function; For modal identification, v and r represent the visible light mode and infrared mode, respectively; Image frame number; This is a quality feature vector extracted based on the exposure, sharpness, infrared temperature range, registration residual, and occlusion ratio of the visible light image and the registered infrared image. This is the quality reliability weight vector; This is a quality reliability bias.
[0133] The process of obtaining fusion features includes:
[0134] For each spatial location u of each feature layer, based on the visible light modal quality reliability and infrared modal quality reliability and visible light local confidence map and infrared partial confidence map Calculate visible light fusion weights Infrared fusion weights :
[0135] ;
[0136] Then, the fusion features are generated using the following formula. :
[0137] ;
[0138] In the formula, h is the feature level; n is the image frame number; and denoted as visible light features and infrared features at position u, respectively; Φ is a cross-modal deformation attention term used to dynamically adjust feature alignment based on the local similarity between visible light features and infrared features, in order to compensate for feature offset caused by incomplete local registration.
[0139] In practical implementation, in addition to the above-mentioned fusion methods, early channel splicing, mid-term cross-attention fusion, late-stage box-level fusion, evidence-theory fusion, or Bayesian fusion can also be used. For edge devices with very low computing power, infrared and visible light can be detected separately first, and then the results can be fused using registration relationships and expert rules. Further details will not be elaborated here.
[0140] S4. Input the fused features from S3 into the closed-set real-time detection network to detect preset fixed-category targets and output a first detection result including target detection box, category, confidence level, segmentation mask, and surface temperature statistics. When the confidence level of the first detection result is lower than a preset threshold, or the category of the target is unknown, trigger the open vocabulary detection model to verify the corresponding target. The open vocabulary detection model is used to identify targets not included in the preset fixed categories and output a second detection result including the possible categories and confidence levels of the target.
[0141] The process of real-time closed set detection and open vocabulary verification is as follows: Figure 5 As shown.
[0142] In specific implementation, when the open vocabulary detection model is triggered for verification, based on the detection box of the corresponding target output by S4, local visible light images and local infrared images containing the target are cropped from the visible light image and the registered infrared image, respectively, and prompt words containing the category and / or location information of the target are generated; then the local visible light images, local infrared images and prompt words are uploaded to the cloud, and the open vocabulary detection model deployed in the cloud is used for recognition; in specific implementation, the open vocabulary detection model is one or a combination of an open vocabulary detection network, a segmentation basic model or a visual language large model.
[0143] To facilitate a better understanding by those skilled in the art, the following explanation is provided.
[0144] The fused features are fed into two detection pipelines. The first is a closed-set real-time detection pipeline, used for fixed-category targets, suitable for continuous operation on edge devices. This pipeline can use YOLOv10, YOLO11, RT-DETR, RF-DETR, or similar real-time end-to-end detection networks, or a lightweight model after pruning, distillation, and INT8 quantization. The second is an open-vocabulary pipeline, used for verification of unknown foreign objects, new construction machinery, atypical heat-generating areas, and low-confidence samples. This pipeline can use a Grounding DINO-type open-vocabulary detection network, a SAM2-type segmentation base model, and a locally deployed large-scale visual-language model. The open-vocabulary pipeline does not run continuously on the main real-time pipeline, but is triggered when the closed-set detection confidence is low, the target category is not in the database, or the expert model determines it to be high-risk.
[0145] No. The frame detection output is:
[0146] ;
[0147] In the formula, For the first Frame target set; The target sequence number; For the first Number of frame targets; For the first The detection bounding box for each target; For the first The categories of the targets; For the first Detection confidence of each target; For the first A segmentation mask for each target; For the first The surface temperature statistics of a target in an infrared image can be taken as the highest temperature, average temperature, or quantile temperature.
[0148] In practical implementation, the closed-set detection pipeline can employ YOLO series, RT-DETR series, improved DETR networks, RF-DETR-like real-time detection Transformers, lightweight Mamba vision backbones, or existing enterprise detection networks. As long as its input is a fused feature or bimodal image, and its output includes detection boxes, categories, and confidence scores, it can be replaced. Further details are omitted here.
[0149] Open category recognition can employ Grounding DINO-type models, CLIP-type image-text matching models, regional visual-language models, local multimodal large models, or cloud-based large models. Segmentation verification can utilize SAM2-type models, lightweight instance segmentation networks, or traditional GrabCut / edge segmentation methods.
[0150] S5. Based on the first or second detection result obtained in S4, as well as the corresponding meteorological information, line ledger and load current information, construct the conductor heat balance equation and conductor sag calculation model, and calculate the physical residuals, which include heat balance residuals and sag residuals.
[0151] For ease of understanding, the following explanation is provided.
[0152] Image detection provides "what is seen," while physical constraints are used to determine "whether this alarm conforms to the line's operating rules." This invention uses conductor thermal balance, sag, and safe clearance as physical constraints, not requiring them to directly replace the detection network, but rather using them as a basis for verification models and risk scoring.
[0153] The physical constraint verification process in S5 is as follows: Figure 6 As shown. In practice, this process can be integrated into the PINN model, which will not be elaborated further here.
[0154] The system obtains span, conductor type, and allowable sag from the ledger; load current from monitoring or dispatching data; and ambient temperature, wind speed, humidity, and solar radiation from meteorological data. The environmental vector is represented as:
[0155] ;
[0156] In the formula, For the first The environment vector corresponding to the frame; Ambient temperature; Wind speed; Relative humidity; This represents the intensity of solar radiation.
[0157] In specific implementation, the thermal balance equation of the conductor is:
[0158] ;
[0159] in, ;
[0160] ;
[0161] ;
[0162] ;
[0163] In the formula, The equivalent heat capacity of the conductor; This is the derivative of the conductor temperature with respect to time. For the temperature of the conductor; It is Joule fever; Absorbing heat from the sun; For convection cooling; For radiative heat dissipation; This refers to the line load current. For the wire at temperature The resistance below; Solar absorptivity; Solar radiation intensity; The convective heat transfer coefficient is related to wind speed; Ambient temperature; Emissivity of the conductor surface; It is the Stefan-Boltzmann constant;
[0164] The thermal balance residual Calculated using the following formula:
[0165] .
[0166] The calculation model for conductor sag is as follows:
[0167] ;
[0168] In the formula, The theoretical value of conductor sag is calculated based on the aforementioned line ledger, load current, and meteorological information. For gear distance; The total load per unit length of the conductor; The horizontal tension of the conductor; This is the load per unit length corresponding to the self-weight of the conductor; This represents the load per unit length corresponding to icing. This represents the load per unit length corresponding to the wind load.
[0169] The sag residual Calculated using the following formula:
[0170] ;
[0171] In the formula, These are observations of conductor sag obtained through image measurement or sensors.
[0172] The physical constraints are:
[0173] ;
[0174] In the formula, The number of time samples used in physical constraint training or verification; For thermal equilibrium residuals; The weighting is the sag residual. Weighting for exceeding limits; For sag residuals; This is the maximum permissible sag for this span of conductor; This is the theoretical value for conductor sag.
[0175] In this method, the input is not simply an image, but rather image detection results, a temperature matrix, line records, and operational data. For example, if the detector identifies a hot spot in a clamp, but the temperature rise at that point is significantly inconsistent with the load current, ambient temperature, and heat dissipation conditions, the system marks the alarm as "requiring verification." If the infrared hot spot is located within the clamp mask, and both the thermal balance residual and expert rules point to anomalies, the alarm level is increased.
[0176] In practice, physical constraints can be implemented using simplified sag formulas, finite element models, dynamic thermal setpoint models, Kalman filter state estimation, or data-driven regression models. If there is no real-time load current on site, historical load curves or time-period load estimates from the dispatching side can be used as substitutes. Further details will not be elaborated here.
[0177] S6. Based on the first or second detection result obtained in S4 and the physical residual obtained in S5, calculate the risk score for each target in combination with the preset expert rules.
[0178] For ease of understanding, the following explanation is provided.
[0179] In practice, step S6 can be implemented based on the expert model. The expert model uses rules, thresholds, component topology relationships, and historical defects for joint modeling. The rules are not simply hardcoded temperature thresholds, but rather combine component type, temperature rise, duration, adjacent comparison, target-to-conductor clearance, sag state, and weather conditions.
[0180] In practice, the risk score is calculated using the following formula:
[0181] ;
[0182] in, ;
[0183] ;
[0184] ;
[0185] In the formula, For the first Risk score for each objective; Risk scoring bias; , , , , σ represents the risk score weights; σ is the Logistic function. The confidence level for detection in step S4; For the first The temperature rise of the target relative to the ambient temperature; For expert rule hit strength; For normalized physical residuals; For the first The net distance from a target to the nearest wire or live component; For the first The surface temperature statistics of a target in an infrared image can be taken as the highest temperature, the average temperature, or the quantile temperature. A set of expert rules; For expert rule serial numbers; For the first Membership function of expert rules; This is the normalization constant for the thermal equilibrium residuals; The normalization constant for the sag residual; The category of the target.
[0186] The calculation of the risk score for each target, based on preset expert rules, includes:
[0187] Based on the target category obtained from S4, the expert rule corresponding to that category is applied for judgment, where:
[0188] For targets identified as clamps or connectors, determine whether they meet the following criteria: the overlap between the infrared hot spot segmentation mask and the target segmentation mask exceeds the preset overlap threshold, and the target temperature rise exceeds the temperature rise of normal components in the same or adjacent phases.
[0189] For targets identified as channel foreign objects, determine whether they meet the following conditions: the target is located in the conductor protection zone or crossing zone, and the net distance from the target to the nearest live component is lower than the preset safe net distance threshold.
[0190] For targets identified as insulators, determine whether they meet the following criteria: the temperature difference distribution between individual pieces within the insulator string exceeds the preset temperature difference pattern, and exclude false judgments caused by single-frame reflections.
[0191] The more judgment conditions of the target that meet its corresponding rule, or the higher the degree of satisfaction, the higher the calculated expert rule hit strength; the expert rule hit strength is used as an input parameter to calculate the risk score in step S6.
[0192] S7. Based on the risk score described in S6, schedule the corresponding operation;
[0193] When the risk score exceeds the first threshold, an alarm work order is generated and pushed to the manual confirmation process; the alarm work order includes the target location, image, temperature, clearance, physical residual and suggested handling level;
[0194] When the risk score is in the range of below the first threshold and above the second threshold, or when the category of the target in S4 is unknown, the open vocabulary detection model is triggered for verification.
[0195] When the risk score is lower than the second threshold, and the confidence level of the first detection result is lower than the preset threshold, and / or the physical residual exceeds the preset physical reasonableness threshold, the corresponding sample is added to the labeling queue; and the manually confirmed conclusion and the samples in the labeling queue are fed back to the cloud training platform.
[0196] In practice, after the manually confirmed conclusions and the samples in the queue to be labeled are fed back to the cloud training platform, the cloud training platform uses the fed-back samples to retrain and update the closed set real-time detection network and / or the open vocabulary detection model, and then sends the updated model to the edge device.
[0197] In specific implementation, the closed-set real-time detection network minimizes the total loss function. During training, the total loss function for:
[0198] ;
[0199] In the formula, To detect the loss; For cross-modal consistency loss; For registration loss; Loss due to physical constraints; Loss of expert consensus; , , , , These are the weights corresponding to the loss;
[0200] in, ;
[0201] In the formula, For the first The set of valid pixels in the frame that participate in the registration loss calculation; This represents the number of pixels in the set. Huber loss function; Operators for describing visible light edges or structures; For infrared edge or structure description operators; For the first Frame of visible light images; To register the infrared image to the visible light coordinate system; These are the coordinates of the visible light image.
[0202] In practical implementation, the intelligent agent scheduling module can be used to schedule corresponding operations. The intelligent agent scheduling module is geared towards the operation and maintenance process, rather than replacing the detection model. Its inputs include detection results, physical verification results, expert rule results, ledgers, and historical defects. The intelligent agent only calls restricted tools, including image review, open-word dictionary detection, report generation, work order push, historical record query, and data feedback; it does not directly rewrite model parameters, nor does it bypass manual confirmation to release high-risk conclusions.
[0203] The agent's scheduling state and actions are represented as follows:
[0204] ;
[0205] In the formula, For the first The agent scheduling state corresponding to the frame; The scheduling actions output by the intelligent agent; This refers to the agent scheduling strategy. The symbols mentioned above will continue to have the same meaning.
[0206] when When the value exceeds the first-level threshold, the system generates a defect alarm and a draft work order, along with visible light local images, infrared local images, segmentation masks, temperature rise, clearance, sag, rule hit reasons, and physical residuals. When the target is in the middle range, the system triggers open vocabulary detection and large model review, requiring the model to explain the possible categories of the target, the causes of danger, and points requiring manual confirmation. When the model uncertainty is low but the model uncertainty is high, the sample enters the labeling queue for subsequent active learning.
[0207] In practical implementation, lightweight detectors, registration modules, reliability-gated fusion modules, and basic expert rules can be deployed on the edge side to meet the real-time requirements of UAV inspections or fixed-point monitoring. Large-scale model verification, open-category sample mining, PINN parameter calibration, and model retraining modules are deployed in the cloud. The edge-side model is deployed using ONNX or TensorRT, with input scales established for infrared and visible light resolutions respectively; for high-resolution visible light images, candidate regions from pole components are cropped to avoid latency caused by whole-image inference.
[0208] During data feedback, not all images are uploaded. The system only uploads low-confidence samples, high-risk samples, samples with conflicting expert rules, samples with physical residual anomalies, and manually modified samples. The cloud periodically performs annotation verification, model distillation, and quantization before distributing the data to edge devices. For new lines and new devices, parameter calibration is first performed using a small number of field samples before gradually enabling automatic alarms.
[0209] In practical implementation, the expert model and the closed-loop process of the intelligent agent are as follows: Figure 7 As shown.
[0210] It should be noted that, in practice, expert rules can be implemented as rule engines, knowledge graphs, Bayesian networks, gradient boosting trees, or lightweight classifiers. For organizations with mature management processes, they can directly interface with existing defect classification standards and work order systems.
[0211] Agent scheduling can employ deterministic state machines, process engines, large-model agents with tool calls, or multi-agent collaborative mechanisms. If an enterprise does not currently allow large models to enter the production pipeline, it can first enable state machines and expert rules, reserving interfaces for future upgrades.
[0212] In addition to infrared and visible light, this method can also incorporate ultraviolet discharge images, partial discharge acoustic signatures, lidar point clouds, weather station data, and online monitoring sensor data. Adding new modes only requires adjusting quality reliability, registration relationships, and feature extractors, without altering the overall workflow.
[0213] Compared with the prior art, the present invention has the following significant advantages:
[0214] 1. Improve detection stability in complex scenarios. Reliability-gated fusion of visible light and infrared can automatically select a more reliable modality in scenarios with strong light, shadow, night, fog, and heat reflection, reducing missed detections and false detections caused by a single image source.
[0215] 2. Enhance the ability to detect unknown risk targets. The open vocabulary detection and visual language large model are triggered only when needed, and can identify construction machinery, temporary foreign objects and atypical defects outside the training set, without sacrificing the efficiency of real-time inspection on the edge side.
[0216] 3. Reduce false alarms that do not conform to physical laws. Incorporating conductor temperature, sag, span, load current, and environmental conditions into the verification process can filter out some false alarms caused by infrared heat reflection, single-frame noise, and incorrect registration.
[0217] 4. Alarm results are easier for maintenance personnel to handle. The output includes not only the frame and category, but also the tower number, phase, component, temperature rise, clearance, sag, rule hit reason, physical residual, and suggested handling level, reducing the need for manual secondary processing.
[0218] 5. Controllable engineering deployment costs. The real-time main link is deployed on edge devices, and large models and cloud models are called on demand; the models can be quantized, distilled, and pruned, making them suitable for deployment in drone inspection terminals, fixed monitoring points, and edge computing boxes.
[0219] 6. Convenient subsequent maintenance. Different detection networks, segmentation networks, large models, and expert rules can be replaced through interfaces, allowing the use of existing enterprise models as well as the integration of subsequent network structures without requiring a complete overhaul of the entire system.
[0220] Example 2
[0221] To facilitate a better understanding by those skilled in the art, the following example is provided.
[0222] Infrared-visible multimodal target detection and physical constraint alarm in drone inspection scenarios.
[0223] This embodiment uses a UAV inspection mission of a 110kV overhead transmission line as an example. The inspection equipment includes a visible light camera, an infrared thermal imaging camera, a gimbal, a ranging module, an RTK positioning module, and an edge computing unit. The visible light image resolution can be 3840×2160, and the infrared image resolution can be 640×512. Both are acquired through the same trigger signal, and the shooting time, tower number, phase, waypoint, gimbal attitude, ranging value, ambient temperature, wind speed, and line load current are written into the same frame index.
[0224] Before the inspection begins, the system reads the line log to obtain the line name, voltage level, tower number, span, conductor type, insulator string type, hardware type, and protection zone. After the UAV reaches the designated waypoint, it first acquires visible light images and infrared temperature matrices. The edge end completes the initial coordinate transformation according to the camera calibration parameters, and then uses the conductor edge, insulator string centerline, and hardware outline to correct the registration results online. When the registration residual is less than a preset threshold, the system directly enters the detection process; when the registration residual exceeds the preset threshold, the system only performs local deformation correction near the candidate tower components to reduce the computational load of the entire image registration.
[0225] During the target detection phase, the system calculates the exposure, sharpness, and occlusion ratio of the visible light image, as well as the temperature range, thermal noise, and heat reflection suspicion level of the infrared image, and generates modal reliability weights accordingly. The edge-closed set detector performs real-time detection of fixed categories such as conductors, ground wires, insulators, clamps, spacers, vibration dampers, towers, bird nests, foreign objects in passageways, and construction machinery, outputting the detection box, category, confidence level, segmentation mask, and target temperature statistics. If the detection confidence level is lower than the verification threshold, or if the target belongs to a suspected category not stably covered in the training set, the system sends the local image, infrared cropped image, and prompt words to the cloud-based open vocabulary detection link for secondary recognition.
[0226] Taking wire clamp overheating as an example, an infrared hotspot was detected at the edge of a B-phase tension wire clamp area on a certain tower. The highest temperature of the hotspot was 74.8℃, the temperature of the adjacent normal fittings in the same phase was 45.2℃, the ambient temperature was 28℃, and the line load current was 450A. The system calculates the overlap between the hotspot mask and the wire clamp segmentation mask. When the overlap exceeds a preset threshold, the hotspot is determined to be located on the wire clamp component. Subsequently, the thermal balance and sag physical constraint models are called, and the physical residual is calculated in combination with the load current, conductor type, wind speed, and ambient temperature. If the temperature rise, duration of frames, mask overlap, and physical residual all meet the expert rules, the risk score increases, and an alarm result of "abnormal overheating of the wire clamp, it is recommended to arrange review and maintenance" is output.
[0227] Taking the intrusion of construction machinery into a passageway as an example, the closed-set detector identifies a crane boom-like target in the visible light image. The system calculates the clearance between the target and the nearest conductor based on the registered spatial relationship of the line, and determines whether it has entered the passageway risk zone by combining the tower number, span, and protection zone range. If the clearance is less than the safety threshold and the target maintains a trend of approaching the conductor in consecutive frames, the expert model classifies it as a high-risk foreign object in the passageway. If the target category does not belong to the stable category of the closed-set detector, the cloud-based open vocabulary detection further provides candidate categories such as "crane boom" and "construction machinery" and their confidence levels for display on the manual review interface.
[0228] For potential false alarm scenarios, such as localized heating of hardware surfaces due to sunlight, mismatch between background building heat sources and circuit components, or isolated high-temperature noise spots in infrared images, the system will simultaneously check whether the hot spot falls within the effective component mask, whether the cross-modal registration residual exceeds the limit, whether the temperature rise is consistent with the load current and heat dissipation conditions, and whether it persists in adjacent frames. If any of these key conditions are not met, the system will lower the alarm level, mark it as "requires verification," or filter it as a low-priority event, thereby reducing false alarms caused by single-frame thermal noise and background heat sources.
[0229] After an alarm is generated, the system records the line name, tower number, phase, component category, visible light local image, infrared local image, target mask, maximum temperature, relative temperature rise, clearance, physical residual, expert rule hit items, and suggested handling level in the work order draft. After confirmation by maintenance personnel, the system feeds back the manual conclusion, correction category, and defect level to the sample database. The cloud training platform periodically selects low-confidence samples, high-risk samples, rule conflict samples, and manually corrected samples for retraining, distillation, and quantization, and then distributes the updated edge model to the UAV inspection terminal.
[0230] The line voltage level, camera resolution, temperature value, threshold, and network model name in this embodiment are only used to illustrate one possible implementation. For scenarios involving fixed cameras, inspection robots, vehicle-mounted inspection equipment, or other modalities such as ultraviolet, voiceprint, and lidar, as long as a combination of infrared-visible multimodal detection, physical constraint verification, expert model risk scoring, and human-machine closed-loop alarm is used, the technical approach of this invention can be followed.
[0231] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Those skilled in the art should understand that any modifications or equivalent substitutions to the technical solutions of the present invention without departing from the spirit and scope of the present invention should be covered within the scope of the claims of the present invention.
Claims
1. A method for infrared-visible multimodal target detection and alarm for power transmission lines, characterized in that, Includes the following steps: S1. Collect visible light images and infrared images of transmission lines, as well as corresponding meteorological information, line records, and load current information; S2. Based on the visible light image and infrared image of S1, initial registration is performed using offline calibration parameters to obtain the initial registration result; Based on the features related to the transmission line structure extracted from the visible light image, the initial registration result is locally corrected online, and the infrared image is registered to the visible light image coordinate system to generate the registered infrared image. S3. Perform multi-layer feature extraction on the visible light image and the registered infrared image of S2 respectively to obtain visible light features and infrared features; Based on the visible light image and the registered infrared image, the quality reliability of the visible light mode and the infrared mode is calculated; based on the quality reliability, fusion weights are generated, and the visible light features and infrared features are weighted and fused. At the same time, cross-modal deformation attention is introduced during the fusion process to compensate for feature shifts caused by incomplete local registration, and the fused features are obtained. S4. Input the fused features from S3 into the closed-set real-time detection network to detect preset fixed-category targets and output a first detection result including target detection box, category, confidence level, segmentation mask, and surface temperature statistics. When the confidence level of the first detection result is lower than a preset threshold, or the category of the target is unknown, trigger the open vocabulary detection model to verify the corresponding target. The open vocabulary detection model is used to identify targets not included in the preset fixed categories and output a second detection result including the possible categories and confidence levels of the target. S5. Based on the first or second detection result obtained in S4, as well as the corresponding meteorological information, line ledger and load current information, construct the conductor heat balance equation and conductor sag calculation model, and calculate the physical residual, which includes the heat balance residual and the sag residual. S6. Based on the first or second detection result obtained in S4 and the physical residual obtained in S5, calculate the risk score for each target by combining the preset expert rules. S7. Based on the risk score described in S6, schedule the corresponding operation; When the risk score is higher than the first threshold, an alarm work order is generated; When the risk score is in the range of below the first threshold and above the second threshold, or when the category of the target in S4 is unknown, the open vocabulary detection model is triggered for verification. When the risk score is lower than the second threshold, and the confidence level of the first detection result is lower than the preset threshold, and / or the physical residual exceeds the preset physical reasonableness threshold, the corresponding sample is added to the labeling queue; and the conclusion after manual confirmation and the samples in the labeling queue are fed back to the cloud training platform.
2. The method for infrared-visible multimodal target detection and alarm of transmission lines according to claim 1, characterized in that, In step S2, the initial registration using offline calibration parameters includes: Using the homography transformation matrix from infrared image to visible light image The infrared image is resampled to achieve initial registration. ; In the formula, This is the intrinsic parameter matrix of the visible light camera; This is the intrinsic parameter matrix of the infrared camera; This is the rotation matrix from the infrared coordinate system to the visible light coordinate system; This is the translation vector from the infrared coordinate system to the visible light coordinate system; The normal vector of the local reference plane; This is the distance from the camera's optical center to the local reference plane; Visible light image coordinates; For the first Frame infrared image; To register the infrared image to the visible light coordinate system; The online local correction of the initial registration result based on features related to the transmission line structure extracted from visible light images includes: From the visible light image, at least one of the following is extracted as the feature related to the transmission line structure: conductor edge, tower corner point, insulator string centerline, and hardware outline; The extracted features are processed to obtain the corresponding sparse keypoints and guide curves; Based on the spatial geometric constraints formed by the sparse keypoints and the guide curves, the initial registration result is locally deformed and corrected within the region to be corrected and its surrounding preset pixel width, as determined by the sparse keypoints and the guide curves.
3. The transmission line infrared-visual multi-modal target detection and warning method of claim 1, wherein, In step S3, the quality reliability of the visible light mode and the infrared mode is calculated. This can be achieved in the following ways: ; In the formula, σ is the Logistic function; For modal identification, v and r represent the visible light mode and infrared mode, respectively; Image frame number; This is a quality feature vector extracted based on the exposure, sharpness, infrared temperature range, registration residual, and occlusion ratio of the visible light image and the registered infrared image. This is the quality reliability weight vector; This is a quality reliability bias.
4. The method for infrared-visible multimodal target detection and alarm of transmission lines according to claim 1, characterized in that, Step S3, the process of obtaining the fused features includes: For each spatial location u of each feature layer, based on the visible light modal quality reliability and infrared modal quality reliability and visible light local confidence map and infrared partial confidence map Calculate visible light fusion weights Infrared fusion weights : ; Then, the fusion features are generated using the following formula. : ; where h is the characteristic level; n is the image frame number; and are the visible and infrared features at location u, respectively; Φ is the cross-modal deformation attention term, which dynamically adjusts the feature alignment based on the local similarity between the visible and infrared features to compensate for feature shifts caused by local misregistration.
5. The transmission line infrared-visual multi-modal target detection and warning method of claim 1, wherein, The line ledger includes the line name, tower number, phase, span, and conductor type; In step S5, the thermal balance equation for the conductor is: ; wherein ; ; ; ; In the formula, The equivalent heat capacity of the conductor; This is the derivative of the conductor temperature with respect to time; For the temperature of the conductor; It is Joule fever; Absorbing heat from the sun; For convection cooling; For radiative heat dissipation; This refers to the line load current. For the wire at temperature The resistance below; Solar absorptivity; Solar radiation intensity; The convective heat transfer coefficient is related to wind speed; Ambient temperature; Emissivity of the conductor surface; It is the Stefan-Boltzmann constant; The thermal balance residual Calculated using the following formula: 。 6. The transmission line infrared-visible multi-modal target detection and warning method of claim 5, wherein, In step S5, the conductor sag calculation model is as follows: ; In the formula, The theoretical value of conductor sag is calculated based on the aforementioned line ledger, load current, and meteorological information. For gear distance; The total load per unit length of the conductor; The horizontal tension of the conductor; This is the load per unit length corresponding to the self-weight of the conductor; This represents the load per unit length corresponding to icing. This represents the load per unit length corresponding to the wind load. the sag residual is calculated by the formula: ; wherein is the conductor sag observation value obtained by image measurement or sensor.
7. The EHV transmission line IR-Vis multimodal target detection and warning method according to claims 5 and 6, characterized in that, In step S6, the risk score is calculated using the following formula: ; wherein ; ; ; In the formula, For the first Risk score for each objective; Risk scoring bias; , , , , σ represents the risk score weights; σ is the Logistic function. The confidence level for detection in step S4; For the first The temperature rise of the target relative to the ambient temperature; For expert rule hit strength; Normalized physical residuals; For the first The net distance from a target to the nearest wire or live component; For the first The surface temperature statistics of a target in an infrared image can be taken as the highest temperature, the average temperature, or the quantile temperature. A set of expert rules; For expert rule serial numbers; For the first Membership function of expert rules; This is the normalization constant for the thermal equilibrium residuals; The normalization constant for the sag residual; The category of the target.
8. The method for infrared-visible multimodal target detection and alarm of transmission lines according to claim 1, characterized in that, In step S6, calculating the risk score for each target using preset expert rules includes: Based on the target category obtained from S4, the expert rule corresponding to that category is applied for judgment, where: For targets identified as clamps or connectors, determine whether they meet the following criteria: the overlap between the infrared hot spot segmentation mask and the target segmentation mask exceeds the preset overlap threshold, and the target temperature rise exceeds the temperature rise of normal components in the same or adjacent phases. For targets identified as channel foreign objects, determine whether they meet the following conditions: the target is located in the conductor protection zone or crossing zone, and the net distance from the target to the nearest live component is lower than the preset safe net distance threshold. For targets identified as insulators, determine whether they meet the following criteria: the temperature difference distribution between individual pieces within the insulator string exceeds the preset temperature difference pattern, and exclude false judgments caused by single-frame reflections. The more judgment conditions of the target that meet its corresponding rule, or the higher the degree of satisfaction, the higher the calculated expert rule hit strength; the expert rule hit strength is used as an input parameter to calculate the risk score in step S6.
9. The transmission line infrared-visual multi-modal target detection and warning method of claim 1, wherein, In step S4, when the open vocabulary detection model is triggered for verification, based on the detection box of the corresponding target output in S4, a local visible light image and a local infrared image containing the target are cropped from the visible light image and the registered infrared image, respectively, and a prompt word containing the category and / or location information of the target is generated; then the local visible light image, the local infrared image and the prompt word are uploaded to the cloud, and the open vocabulary detection model deployed in the cloud is used for recognition; In step S7, after the manually confirmed conclusions and the samples in the queue to be labeled are fed back to the cloud training platform, the cloud training platform uses the fed-back samples to retrain and update the closed set real-time detection network and / or the open vocabulary detection model, and then sends the updated model to the edge device.
10. The transmission line infrared-visual multi-modal target detection and warning method of claim 1, wherein, The closed set real-time detection network minimizes a total loss function The total loss function is: ; In the formula, To detect the loss; For cross-modal consistency loss; For registration loss; Loss due to physical constraints; Loss of expert consensus; , , , , These are the weights corresponding to the loss; wherein ; In the formula, For the first The set of valid pixels in the frame that participate in the registration loss calculation; This represents the number of pixels in the set. Huber loss function; Operators for describing visible light edges or structures; For infrared edge or structure description operators; For the first Frame of visible light images; To register the infrared image to the visible light coordinate system; These are the coordinates of the visible light image. ; In the formula, is the number of time samples participating in the physical constraint training or verification; is the heat balance residual; is the sag residual weight; is the out-of-limit penalty weight; is the sag residual; is the maximum allowable sag of the section conductor; is the theoretical value of the conductor sag.