Overhead line external damage intelligent identification and early warning method, system and electronic equipment

By collecting multi-view visual and environmental sound signals, combined with target detection models and dynamic threat indices, comprehensive and multi-dimensional monitoring and proactive early warning of overhead lines are achieved. This solves the problems of blind spots and high false alarm rates in existing technologies, and improves the accuracy and response capability of external damage identification.

CN121167653BActive Publication Date: 2026-02-10STATE GRID GANSU ELECTRIC POWER RESEARCH INSTITUTE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511717497.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-02-10
Estimated Expiration
2045-11-21

AI Technical Summary

Technical Problem

Existing technologies cannot achieve comprehensive and multi-dimensional monitoring of external damage to overhead lines, resulting in blind spots, high false alarm rates, and passive responses, making it difficult to effectively prevent external damage incidents.

Method used

By collecting multi-view visual data and environmental sound signals, a target detection model is constructed for three-dimensional spatial positioning. Combined with a dynamic threat index, a hierarchical adaptive early warning system is implemented, integrating multi-dimensional information to identify and warn of potential threats.

Benefits of technology

It achieves comprehensive, blind-spot-free monitoring, reduces false alarm rates, can proactively intervene on-site, accurately identifies potential external damage threats and issues audible and visual warnings, and effectively prevents external damage incidents from occurring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121167653B_ABST
    Figure CN121167653B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of power system safety monitoring, and is an overhead line external damage intelligent identification and early warning method, system and electronic device; the system comprises: a signal acquisition unit that acquires overhead line surrounding environment data signals; a target identification unit that constructs a target detection model, processes multi-view visual data signals, and identifies key targets; and a target determination unit that, based on the spatial distance and relative motion trend of the calculated key targets and the overhead line, in combination with environmental sound signals, further identifies to obtain the real identification result of the key targets. The present application can monitor in all directions without dead angles, multi-dimensional information fusion to reduce false positives, and can actively intervene on site, can fuse multi-view visual information and acoustic information, realize accurate and rapid identification of potential external damage threats, and actively issue sound and light warnings to effectively prevent external damage events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system safety monitoring technology, and in particular to a method, system, and electronic equipment for intelligent identification and early warning of external damage to overhead lines. Background Technology

[0002] Overhead lines mainly refer to exposed overhead lines, erected above the ground. They are power transmission lines that use insulators to fix the transmission conductors to towers erected on the ground to transmit electrical energy. They are relatively easy to install and maintain, and have lower costs. However, they are susceptible to weather and environmental factors (such as strong winds, lightning strikes, pollution, and snow), which can cause faults. Furthermore, the entire transmission corridor occupies a large area of ​​land and can easily cause electromagnetic interference to the surrounding environment. The main components of an overhead line include: conductors and lightning protection wires (overhead ground wires), towers, insulators, hardware, tower foundations, guy wires, and grounding devices.

[0003] External damage to overhead power lines refers to damage caused by external factors. Common causes include collisions with construction machinery, fallen trees, and animal disturbance. Specifically: Construction machinery: Large machinery operating without prior authorization or maintaining a safe distance can easily collide with power lines or guy wires. Tree hazards: Fallen trees or branches touching the power lines can cause short circuits or line breaks. Animal disturbance: Birds nesting or squirrels gnawing on power lines can damage the lines. Equipment aging: Aging of insulators, conductors, and other components can also lead to external damage.

[0004] Currently, there are two main technologies for identifying external damage to overhead lines: one is manual inspection, which relies on regular inspections by patrol personnel. This is inefficient, cannot achieve real-time monitoring, and is slow to react to sudden external damage events, such as crane misoperation during construction, making it difficult to effectively prevent such damage. The other is single-sensor monitoring, which typically uses a single camera for monitoring. This has the following drawbacks:

[0005] 1) Limited field of view: A single camera cannot cover the entire area of ​​overhead lines, such as key areas under or behind the lines, resulting in blind spots.

[0006] 2) High false alarm rate: It is difficult to distinguish real threats, such as a rising crane, from irrelevant objects (such as stationary trees) based on images. Weather conditions such as wind, snow, and rain can also easily cause false alarms.

[0007] 3) Passive response: The system only alarms after identifying a threat, and cannot actively intervene or drive it away, thus missing the best opportunity for intervention.

[0008] Therefore, how to provide an overhead line external damage identification technology that can provide all-round, blind-spot-free monitoring and multi-dimensional information fusion to reduce false alarms is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0009] This invention provides an intelligent identification and early warning method, system, and electronic device for external damage to overhead lines, which overcomes the shortcomings of the prior art. It can effectively solve the problem that the existing methods of monitoring external damage to overhead lines by manual or single sensor cannot monitor the damage in a comprehensive and multi-dimensional manner, which leads to false alarms of external damage to overhead lines.

[0010] To address the above problems, one of the technical solutions of this invention is achieved through the following method: a method for intelligent identification and early warning of external damage to overhead lines, comprising:

[0011] Collect environmental data signals around overhead lines, including multi-view visual data signals and environmental sound signals;

[0012] A target detection model is constructed to process multi-view visual data signals, identify key targets, perform three-dimensional spatial positioning of key targets, and calculate the spatial distance and relative motion trend between key targets and overhead lines; among them, key targets are construction machinery, trees or pedestrians;

[0013] Based on the calculated spatial distance and relative motion trend between the key target and the overhead line, combined with environmental sound signals, further identification is performed to obtain the true identification results of the key target;

[0014] Based on the actual identification results, a quantitative assessment model of the threat situation is constructed, which includes the fusion of multi-dimensional information. A dynamic threat index is introduced, and a hierarchical adaptive early warning is carried out based on the dynamic threat index. The multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation degree.

[0015] The aforementioned target detection model processes multi-view visual data signals to identify key targets, performs three-dimensional spatial positioning of the key targets, and calculates the spatial distance and relative motion trend between the key targets and the overhead lines, including:

[0016] The object detection model was trained on an image dataset containing scenes of construction machinery, trees, and pedestrians.

[0017] The object detection model was trained and quantized using the YOLOv5s model.

[0018] Based on the calculated spatial distance and relative motion trend between the key target and the overhead line, combined with environmental sound signals, further identification was performed to obtain the true identification results of the key target, including:

[0019] After identifying the key target, determine whether the spatial distance between the key target and the overhead line is less than the corresponding safety threshold;

[0020] In response, acoustic analysis is triggered, which involves extracting the sound signal from the environmental sound signal in the direction of the key target and performing spectral matching on the sound signal.

[0021] Determine whether the matched sound signal is the sound signal corresponding to the identified key target;

[0022] If the response is correct, then the identification result is correct, and the key target is identified as the true identification result.

[0023] Based on the actual identification results, a quantitative assessment model of the threat situation is constructed, incorporating multi-dimensional information fusion, and a dynamic threat index is introduced. The multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation level, including:

[0024] Dynamic Threat Index DTI as follows:

[0025] DTI = f ( D , V , T , C ),

[0026] In the formula, D Spatial distance between key targets and overhead power lines, unit: meters; V Radial velocity of the critical target toward the overhead line, in meters per second; T The threat type coefficient for key targets is a weight preset through machine learning or expert experience, ranging from [0, 1]. C The audiovisual fusion confirmation score, i.e., the overall confidence multiplier, ranges from [0, 1].

[0027] The dynamic threat index DTI The spatial distance and radial velocity in the data are normalized, where the spatial distance is... D The normalization formula is as follows:

[0028] ,

[0029] In the formula, This is the normalized spatial distance threat value, ranging from [0, 1]. Distance to the midpoint of risk; This is the distance sensitivity coefficient;

[0030] Among them, radial velocity V The normalization formula is as follows:

[0031] ,

[0032] In the formula, This is the normalized radial velocity threat value, ranging from [0, 1]. The safe speed threshold; Maximum dangerous speed;

[0033] Based on the above normalized parameters, the final dynamic threat index is obtained. DTI The calculation method is as follows:

[0034] ,

[0035] In the formula, This represents the final dynamic threat index, ranging from [0, 1]. The weighting index for spatial distance. The weighting index for radial velocity, As a weighting index for threat types, and ; This refers to the audiovisual fusion confirmation degree, which is the overall confidence multiplier.

[0036] The aforementioned audiovisual fusion confirmation, i.e., the overall confidence multiplier, includes single-modal internal evaluation, cross-modal cross-validation, and spatiotemporal consistency verification.

[0037] (1) The intramodal assessment includes visual confidence and acoustic confidence, wherein:

[0038] Visual confidence The output is from the object detection model, and the calculation method is as follows:

[0039] ,

[0040] In the formula, Confidence_score_of_all_crane_objects represents the confidence score of all "construction machinery" targets in the image; This indicates how confident the model is in identifying a certain region in the image as "construction machinery";

[0041] Acoustic confidence The sound classification model outputs a probability distribution. This indicates how confident the sound classification model is in identifying the current ambient sound as "the sound of construction machinery in operation";

[0042] (2) Cross-modal cross-validation: Define a correlation score, which is a Boolean value or a soft score, including spatial consistency and state consistency, where:

[0043] Spatial consistency The system uses microphone array sound source localization technology to estimate the direction of the sound source, compares this sound direction with the location of the "construction machinery" target in the visually detected image, and determines whether the two directions are basically consistent. If so, then... If the response is no, then ;

[0044] State Consistency The system determines whether the "states" described by vision and hearing match, and responds accordingly. If the response is no, then ;

[0045] The comprehensive correlation score can be obtained from the above. :

[0046] ,

[0047] In the formula, and It's weight. ;

[0048] (3) Spatiotemporal consistency verification: using a sliding window, calculate the comprehensive correlation score within the sliding window. The continuity and stability of the values ​​are calculated as follows:

[0049] ,

[0050] In the formula, For spatiotemporal consistency score, ; For comprehensive relevance score; N is the total number of frames in the window, S i It refers to the first i Frame metrics score, It is an indicator function; its value is 1 when the condition inside the parentheses is true, and 0 otherwise.

[0051] By integrating all the above information, the final audiovisual fusion confirmation is obtained. The calculation formula is as follows: ,

[0052] In the formula, Visual confidence level; Acoustic confidence level; Visual weight; Auditory weighting; The overall relevance score; The score is based on spatiotemporal consistency.

[0053] The above-mentioned hierarchical adaptive early warning based on dynamic threat index includes:

[0054] The warning system is divided into four levels, specifically:

[0055] Level 1: Sensing and Early Warning DTI ∈[0.2, 0.4); Warning action: LED warning lights slowly turn on and off;

[0056] Level 2: Warning, DTI∈[0.4, 0.7); Warning action: LED warning light flashes rapidly in yellow or orange, and the speaker plays a short, neutral sound;

[0057] Level 3: Departure warning, DTI∈[0.7, 0.9); Warning action: LED warning light flashes red, speaker plays the corresponding preset voice command at high volume;

[0058] Level 4: Emergency Evacuation and Reporting, DTI∈[0.9, 1.0]; Warning Action: Maintain the warning action of Level 3, and remotely report "Level 1 Emergency Alarm" and related threat proof information to the backend; The threat proof information includes real-time snapshots, pre-incident video, situation snapshots and audio clips.

[0059] The second technical solution of the present invention is achieved in the following way: an intelligent identification and early warning system for external damage to overhead lines, using an intelligent identification and early warning method for external damage to overhead lines, including: a central processing and communication module, wherein the central processing and communication module includes a signal acquisition unit, a target identification unit, a target determination unit and an adaptive early warning unit;

[0060] The signal acquisition unit collects environmental data signals around the overhead line, including multi-view visual data signals and environmental sound signals.

[0061] The target recognition unit constructs a target detection model, processes multi-view visual data signals, identifies key targets, performs three-dimensional spatial positioning of key targets, and calculates the spatial distance and relative motion trend between key targets and overhead lines; among them, key targets are construction machinery, trees, or pedestrians;

[0062] The target identification unit, based on the calculated spatial distance and relative motion trend between the key target and the overhead line, combined with environmental sound signals, further identifies the key target to obtain the true identification result.

[0063] The adaptive early warning unit constructs a quantitative assessment model of the threat situation that incorporates multi-dimensional information fusion based on the actual identification results. It introduces a dynamic threat index and performs hierarchical adaptive early warning based on the dynamic threat index. The multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation degree.

[0064] The above also includes a power supply module, which is connected to the central processing and communication module.

[0065] The aforementioned adaptive warning unit also includes a high-brightness LED warning light and a high-power speaker.

[0066] The third technical solution of the present invention is achieved in the following way: an electronic device, including a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to realize an intelligent identification and early warning method for external damage to overhead lines.

[0067] This invention enables all-round, blind-spot-free monitoring, multi-dimensional information fusion to reduce false alarms, and proactive on-site intervention. It can integrate multi-view visual and acoustic information to achieve accurate and rapid identification of potential external damage threats and proactively issue audible and visual warnings to effectively prevent external damage incidents from occurring. Attached Figure Description

[0068] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0069] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention.

[0070] Figure 2 This is a system block diagram of Embodiment 6 of the present invention. Detailed Implementation

[0071] The present invention is not limited to the following embodiments, and the specific implementation can be determined according to the technical solution of the present invention and the actual situation.

[0072] Example 1: As Figure 1 As shown in the figure, an embodiment of the present invention discloses an intelligent identification and early warning method for external damage to overhead power lines, including:

[0073] S101, collects environmental data signals around the overhead line, including multi-view visual data signals and environmental sound signals;

[0074] S102, Construct a target detection model, process multi-view visual data signals, identify key targets, perform three-dimensional spatial positioning of key targets, and calculate the spatial distance and relative motion trend between key targets and overhead lines; among them, key targets are construction machinery (cranes, excavators), trees or pedestrians.

[0075] S103, based on the calculated spatial distance and relative motion trend between the key target and the overhead line, combined with environmental sound signals, further identification is performed to obtain the true identification result of the key target;

[0076] S104. Based on the actual identification results, construct a quantitative assessment model of the threat situation that includes multi-dimensional information fusion, introduce a dynamic threat index, and conduct hierarchical adaptive early warning based on the dynamic threat index; among which, the multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation degree.

[0077] In step S101 above, environmental data signals around the overhead line are collected, including multi-view visual data signals and environmental sound signals.

[0078] Among them, multi-view visual data signals can be acquired using a multi-view visual acquisition module, specifically including:

[0079] Forward-facing camera: Used to monitor the area in front of and below overhead lines, mainly to identify approaching construction vehicles, such as cranes and excavators;

[0080] Rearview camera: Used to detect the area behind and below overhead lines, eliminating blind spots behind them;

[0081] Downward-viewing camera: Used to detect the area directly below overhead power lines, accurately determining the distance to trees, the height of debris piled up under the lines, etc.

[0082] There are overlapping areas between each pair of the front, rear, and lower areas to ensure that the three cameras can cover the area where the overhead lines are located.

[0083] Among them, the ambient sound signal is collected by an acoustic sensing module composed of a microphone array, including: by analyzing the sound spectrum characteristics of the ambient sound, the operating noise of specific equipment, such as the working sound of a crane engine or hydraulic system, can be identified as a supplement and confirmation to visual recognition.

[0084] In step S102 above, a target detection model is constructed, multi-view visual data signals are processed, key targets are identified, the key targets are located in three-dimensional space, and the spatial distance and relative motion trend between the key targets and the overhead lines are calculated, including:

[0085] The object detection model was trained on an image dataset containing scenes of construction machinery (cranes, excavators), trees, and pedestrians.

[0086] The object detection model was trained and quantized using the YOLOv5s model.

[0087] The object detection model can be trained on an image dataset containing a large number of scenes such as cranes, excavators, trees and pedestrians. The object detection module can be trained and quantized using the YOLOv5s model, because the YOLOv5s model is a lightweight model that can run efficiently on embedded devices.

[0088] In step S103 above, based on the calculated spatial distance and relative motion trend between the key target and the overhead line, and combined with environmental sound signals, further identification is performed to obtain the true identification result of the key target, including:

[0089] After identifying the key target, determine whether the spatial distance between the key target and the overhead line is less than the corresponding safety threshold;

[0090] In response, acoustic analysis is triggered, which involves extracting the sound signal from the environmental sound signal in the direction of the key target and performing spectral matching on the sound signal.

[0091] Determine whether the matched sound signal is the sound signal corresponding to the identified key target;

[0092] If the response is correct, then the identification result is correct, and the key target is identified as the true identification result.

[0093] Among them, the sound signal can be analyzed using an acoustic feature model to obtain the sound type. The acoustic feature model can be trained using a classification model based on a sound dataset containing a large number of crane working or running sounds, excavator working or running sounds, pedestrian walking or talking sounds, and environmental sounds when trees are still and swaying. The labels used for training can be manually labeled. For example, the sound type of crane working or running sounds can be labeled as crane, and the labels of other sound types can be applied in the same way.

[0094] In step S104 above, based on the actual identification results, a quantitative assessment model of the threat situation that incorporates multi-dimensional information fusion is constructed, and a dynamic threat index is introduced; wherein, the multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation degree, including:

[0095] Dynamic Threat Index DTI as follows:

[0096] DTI = f ( D , V , T , C ),

[0097] In the formula, D The spatial distance between the key target and the overhead power line, in meters. The closer the distance, the higher the accuracy. DTI The higher; V Radial velocity of the critical target toward the overhead line, unit: meters per second. The faster the velocity, the lower the value. DTI The higher; TThe threat type coefficient for key targets is a weight preset through machine learning or expert experience, ranging from [0, 1], for example: crane boom ( T =0.9), excavator ( T =0.7), abnormal floating objects ( T =0.8), pedestrians ( T =0.2); C The audiovisual fusion confirmation score, i.e., the overall confidence multiplier, ranges from [0, 1]; when only visual confirmation is required... C =0.8, when both audiovisual and visual confirmation are required. C =1.0, when acoustic confirmation only C =0.6;

[0098] The dynamic threat index DTI The spatial distance and radial velocity in the data are normalized, where the spatial distance is... D The normalization formula is as follows:

[0099] ,

[0100] In the formula, This is the normalized spatial distance threat value, ranging from [0, 1]. For the distance to the midpoint of risk, when hour, This is a key adjustable parameter, for example, it can be set to 8 meters; This is the distance sensitivity coefficient. The larger the value, the steeper the curve, meaning the system is more sensitive to changes in distance. For example, it can be set to 0.5.

[0101] Among them, radial velocity V The normalization formula is as follows (using a linear function with a threshold, because the threat level of targets increases more slowly below a certain speed):

[0102] ,

[0103] In the formula, This is the normalized radial velocity threat value, ranging from [0, 1]. This is a safe speed threshold; speeds below this threshold are considered non-threatening or extremely low-threat (e.g., less than 0.5 m / s). Calculate starting from 0; This is the maximum dangerous speed; reaching or exceeding this speed... It can be directly considered as 1 (highest threat), for example, it can be set to 5 m / s;

[0104] Based on the above normalized parameters (after obtaining all normalized parameters, they need to be combined, which can be done using a weighted average or a weighted geometric mean formula; the calculation method used below in this invention is the weighted geometric mean formula), the final dynamic threat index is obtained. DTI The calculation method is as follows:

[0105] ,

[0106] In the formula, This represents the final dynamic threat index, ranging from [0, 1]. The weighting index for spatial distance. The weighting index for radial velocity, As a weighting index for threat types, and ; The audiovisual fusion confirmation factor, i.e., the overall confidence multiplier, is used when the system is unsure of what it has seen ( C If the value is low, then even if other parameters seem dangerous, the final dynamic threat index will be low. The value will also be lowered, effectively preventing false alarms caused by misjudgment. Among them, the weight index determines the relative importance of different factors in the total threat.

[0107] Among them, threat types T The parameter itself is a preset expert weight coefficient, ranging from [0, 1], and does not require normalization; that is, after identifying the type of key target, the corresponding preset value can be obtained directly based on the type of key target. Example values: Crane / tower crane boom: T = 0.9; Excavator boom: T = 0.8; Abnormal floating objects (such as dust nets, balloons): T = 0.7; Tall trees (in the wind): T = 0.4; Pedestrians: T = 0.1.

[0108] Among them, audiovisual fusion confirmation C This parameter represents the system's information about the recognition result, and its range is [0, 1], requiring no normalization. Hierarchical assignment: Dual visual and acoustic verification: C = 1.0; High-confidence visual confirmation only: C = 0.8; Acoustic confirmation only: C = 0.6; Low-confidence visual recognition: C = 0.4.

[0109] Therefore, the multi-dimensional information fusion threat situation quantitative assessment model integrates static distance information with dynamic multi-dimensional information such as speed, target type, and confirmation to form a continuous, real-time risk score, enabling the system to distinguish between "slowly approaching trees" and "high-speed cranes," even if the two are at the same distance, their threat levels are completely different.

[0110] In summary, compared with the prior art, the present invention has the following advantages:

[0111] 1. Nonlinear normalization: It uses an inverse S-shaped function to process distance, which is more in line with physical intuition and safety requirements than linear scaling.

[0112] 2. Confidence Multiplication: Using the certainty level C as a multiplier rather than an addend is a "one-vote veto" design. If the system "cannot see clearly" or "cannot hear clearly," it will not easily give a high score. This reflects a clear understanding of the system's own capability limitations and greatly improves robustness.

[0113] 3. Weighted Geometric Mean: Compared to the weighted arithmetic mean, the geometric mean is more sensitive to any low score item. This means that as long as any of the distance, speed, or type is safe (value close to 0), the overall threat index will be significantly lowered, avoiding misjudgments caused by the "weakest link effect".

[0114] The aforementioned audiovisual fusion confirmation, i.e., the overall confidence multiplier, includes single-modal internal evaluation, cross-modal cross-validation, and spatiotemporal consistency verification.

[0115] (1) The intramodal assessment includes visual confidence and acoustic confidence, wherein:

[0116] Visual confidence The confidence score is output by an object detection model (such as YOLO). Along with the bounding box and category, the object detection model provides a confidence score, calculated as follows:

[0117] ,

[0118] In the formula, Confidence_score_of_all_crane_objects represents the confidence score of all "construction machinery" targets in the image; This indicates how confident the model is in identifying a certain region in the image as "construction machinery";

[0119] Acoustic confidence The sound classification model (such as SVM or a small neural network) outputs a probability distribution, and the posterior probability of the category "construction machinery" is taken as the acoustic confidence level. , This indicates how confident the sound classification model is in identifying the current ambient sound as "the sound of construction machinery in operation";

[0120] (2) Cross-modal cross-validation: Define an association score, which is a Boolean value (0 or 1) or a soft score, including spatial consistency and state consistency, where:

[0121] Spatial consistency The system uses microphone array sound source localization technology (such as the TDOA algorithm) to estimate the direction of the sound source, compares this sound direction with the image location of the visually detected "construction machinery" target, and determines whether the two directions are basically consistent. If so, the system responds accordingly. If the response is no, then ;

[0122] State Consistency The system determines whether the "states" described by vision and hearing match, and responds accordingly. If the response is no, then ;

[0123] The comprehensive correlation score can be obtained from the above. :

[0124] ,

[0125] In the formula, and Weights (e.g.) ), ;

[0126] (3) Spatiotemporal consistency verification (the judgment of a single frame may be wrong due to accidental factors, so it is necessary to smooth it in the time dimension to eliminate instantaneous interference). Use a sliding window (e.g., the most recent 1.5 seconds, containing 15 frames of data) to calculate the comprehensive correlation score within the sliding window. The continuity and stability of the values ​​are calculated as follows:

[0127] ,

[0128] In the formula, For spatiotemporal consistency score, ; For comprehensive relevance score; N is the total number of frames in the window, S i It refers to the first i Frame metrics score, This is an indicator function; its value is 1 when the condition within the parentheses is true, and 0 otherwise. If audiovisual association is detected in multiple consecutive frames, then... A value close to 1 indicates that the judgment is stable and reliable. If it is only associated with an occasional frame, then... A low value may be due to noise interference.

[0129] By integrating all the above information, the final audiovisual fusion confirmation is obtained. The calculation formula is as follows:

[0130] ,

[0131] In the formula, Visual confidence level; Acoustic confidence level; Visual weight; Auditory weighting; The overall relevance score; The score is based on spatiotemporal consistency. For example... Visual weight is usually higher.

[0132] Among them, the final audiovisual fusion confirmation degree In the calculation formula, the weighted average within parentheses represents the overall basic confidence in "what was seen and heard"; multiplied by the overall relevance score. This means that if what is "seen" and what is "heard" do not match, even the highest level of basic confidence will be significantly undermined; multiplied by the spatiotemporal consistency score. This means that if this judgment is just a fleeting moment, the final confirmation level will be lowered, thus ensuring the robustness of the decision.

[0133] An example of the aforementioned state consistency is as follows:

[0134] If the visual system detects a "crane," but the acoustic model hears "ambient wind noise," then... ;

[0135] If the visual system detects a stationary crane, but the acoustic model hears the sound of a high-speed engine, this could mean the crane is about to start. It can be set to an intermediate value (such as 0.5) to represent "potential association";

[0136] If the visual system detects a "moving boom" and the acoustic model hears "the sound of the hydraulic system working," then Highly matched.

[0137] The present invention decomposes the process of obtaining audiovisual fusion confirmation into three levels: single-modal internal evaluation, which first gives a "preliminary conclusion" and "confidence score" for both visual and auditory perception; cross-modal cross-validation, which checks whether the conclusions of visual and auditory perception support each other; and spatiotemporal consistency verification, which judges whether the conclusion is stable and reasonable over a continuous time series. The final audiovisual fusion confirmation is the result of a comprehensive evaluation at these three levels.

[0138] In step S104 above, the hierarchical adaptive early warning based on the dynamic threat index includes:

[0139] The warning system is divided into four levels, specifically:

[0140] Level 1: Sensing and Early Warning DTI ∈[0.2, 0.4); The warning action involves the LED warning light slowly turning on and off. This non-intrusive light signal can effectively attract attention without causing panic or aversion. The triggering condition for the perception warning is: a potential threat target is detected entering the first-level alert zone (e.g., 15 meters away, which can be set as needed), but the target moves slowly or has a low threat type coefficient. During the perception warning phase, no sound is played to avoid excessive disturbance. Thus, through a kind of "presence reminder," the other party is gently informed that "you have been noticed," prompting the other party to take the initiative to check their behavior.

[0141] Level 2: Warning and alert, DTI∈[0.4, 0.7); Warning action: LED warning lights flash rapidly (e.g., 1Hz), in a more conspicuous yellow or orange color, and the speaker plays a short, neutral sound; The triggering conditions for the warning and alert are: the target continues to approach, or the threat type coefficient is high; The speaker plays an "environmental alert sound," which is not a harsh alarm, but a short, neutral sound, such as a "ding-dong" sound or a short electronic sound effect; Thus, by upgrading the warning from "notice" to "warning," the combination of rapid flashing and alert sound clearly conveys the message "Attention, there is a risk," but still leaves room for ambiguity.

[0142] Level 3: Departure Warning, DTI∈[0.7, 0.9); Warning Action: LED warning light flashes (e.g., 10Hz), color is the highest level red, speaker plays the corresponding preset voice command at high volume; The trigger condition for the departure warning is: the target enters the level 2 warning zone (e.g., within 8 meters, can be set as needed), and is moving at high speed or is of a high-risk type; The content of the voice command played by the speaker is not fixed. The system will play the most matching warning according to the type of threat identified. For example: the voice command for cranes / excavators may be: "High voltage danger! Crane operation to stop immediately! Back up!"; the voice command for construction workers below ground may be: "High voltage line above! Please stay away immediately! Danger!"; the voice command for foreign objects (e.g., kites, balloons) may be: (mute, flash only); Thus, "command-style intervention" is carried out through departure warning. The flashing red light and the directional voice command use authority and urgency to create strong psychological pressure on the operator, forcing them to stop dangerous behavior.

[0143] Level 4: Emergency Avoidance and Reporting, DTI∈[0.9, 1.0]; Warning Action: Maintain the warning action of Level 3, and simultaneously report a "Level 1 Emergency Alarm" and related threat proof information to the backend remotely via 4G / 5G module; The threat proof information includes real-time snapshots, pre-incident video, situation snapshots, and audio clips; The trigger condition for emergency avoidance and reporting is: the system determines that the collision is unavoidable within 1-2 seconds; The real-time snapshot can be a high-definition image with a target bounding box and distance information, the pre-incident video can be a short video clip 10 seconds before the trigger, the situation snapshot can be JSON data containing all parameters such as DTI, D, V, T, C, etc., and the audio clip is the audio recording at the moment of triggering; Thus, through the settings of the emergency avoidance and reporting stage, even if the active warning fails, this complete chain of evidence can provide irrefutable evidence for post-event tracing, liability determination, and insurance claims, minimizing losses.

[0144] The early warning scheme of this invention adopts an early warning approach of "situational awareness + psychological deterrence + dynamic escalation". Situational awareness: the system not only judges "whether there is a threat", but also understands "what the threat is, what the intention is, and how the risk is evolving". Psychological deterrence: the design of the early warning signal is based on human psychology, aiming to upgrade from "reminder" to "deterrence" and maximize the success rate of preventing behavior. Dynamic escalation: the intensity, content and method of the early warning are adjusted in real time according to the dynamic changes of the threat, forming a "dialogue-style" intervention process.

[0145] Therefore, in the aforementioned early warning scheme of this invention, from "static threshold" to "dynamic situation": a dynamic threat index model based on multi-dimensional information is pioneered, making risk assessment more scientific and more in line with real scenarios, solving the drawbacks of the traditional "one-size-fits-all" approach; from "passive alarm" to "active dialogue": a four-level progressive early warning strategy based on psychology is designed, realizing intelligent intervention from "gentle reminder" to "forced removal", transforming "human-machine confrontation" into "human-machine communication", greatly improving the success rate of early warning; from "single signal" to "adaptive content": the early warning content (especially voice) can be adaptively adjusted according to the threat type, making the warning more targeted and authoritative, which is an unprecedented intelligent design; from "post-event tracing" to "in-event evidence": when the highest risk level is triggered, an "evidence chain" containing multimedia information is immediately generated and reported, realizing "evidence collection upon alarm", providing a last solid digital defense line for power grid security.

[0146] In summary, this invention can provide all-round, blind-spot-free monitoring, integrate multi-dimensional information to reduce false alarms, and proactively intervene on-site. It can integrate multi-view visual and acoustic information to achieve accurate and rapid identification of potential external damage threats and proactively issue audible and visual warnings to effectively prevent external damage incidents from occurring.

[0147] Example 2: This embodiment of the invention discloses an early warning scheme one in an intelligent identification and early warning method for external damage to overhead lines, including:

[0148] For simple early warning, if the actual identification result of the key target points to the current target, which is the preset key target (such as construction machinery, people, trees), and the distance between the key target and the overhead line is less than the safety threshold, the decibel of the alarm sound will be increased and the frequency of the LED warning light will be increased; if the distance between the key target and the overhead line is not less than the safety threshold, a louder and faster alarm sound can be used, and the LED warning light can also flash at a lower frequency; if the key target is not construction machinery, the latter early warning method can be used, that is, a lower decibel alarm sound can be used, and the LED warning light can also flash at a lower frequency.

[0149] Conversely, if the actual identification result of the key target points to a target that is not the preset key target (such as construction machinery, people, or trees), no warning is required.

[0150] Example 3: This embodiment of the invention discloses a second early warning scheme in an intelligent identification and early warning method for external damage to overhead lines, which is illustrated using distance values, including:

[0151] Low-risk warning: When the distance between the critical target and the line is less than 15 meters and greater than or equal to 8 meters, the LED warning light will flash slowly.

[0152] Medium-risk warning: When the distance between the critical target and the overhead line continues to shorten and the shortest distance is less than 8 meters but greater than or equal to 3 meters, the LED warning lights will flash rapidly and a low-volume warning voice will be played through the speaker.

[0153] High-risk warning: When a critical target is determined to be less than 3 meters away, the LED lights will flash, the speaker will play an emergency warning voice message at maximum volume in a loop, and an alarm message will be immediately reported to the monitoring center via the wireless communication module. The alarm message includes on-site images, videos, and threat level.

[0154] The threat levels are categorized as low, medium, and high, and the criteria for determining the threat level have been explained in the various types of warnings mentioned above.

[0155] Example 4: This embodiment of the invention discloses a technical solution that can implement the above-mentioned target detection model and acoustic feature model with a single model. Specifically, a "dual-stream progressive" lightweight detection network is designed, whose network structure consists of three parts: a shared lightweight backbone, a dual-stream detection head, and a progressive fusion module, wherein:

[0156] 1. Shared lightweight backbone

[0157] Function: To efficiently extract multi-scale features from images;

[0158] Structure selection: MobileNetV3-Small is used as the backbone network; it utilizes depthwise separable convolution and neural architecture search techniques to achieve excellent feature extraction capabilities with extremely low computational cost, making it very suitable for embedded devices;

[0159] Improvement: This paper modifies the last stage of MobileNetV3-Small, expanding its output into two feature maps of different scales, which are used to detect small targets (such as a distant crane boom) and large targets (such as a nearby crane body).

[0160] 2. Dual-flow detector head

[0161] Traditional YOLO has only one detection head, while this design uses two parallel detection heads with different functions.

[0162] A-flow - "Coarse Screening" Detector Head:

[0163] Structure: A very lightweight detection head containing only a small number of convolutional layers;

[0164] Function: Quickly scans the entire image to identify all candidate regions that may be "threats" with high recall; it does not require high accuracy but requires extremely high speed to avoid missing any potential targets; it does not output precise bounding boxes, but rather some "regions of interest";

[0165] B-flow - "Precise Judgment" Detection Head:

[0166] Structure: A relatively complex but still lightweight detection head containing more convolutional layers and attention mechanisms (such as SE modules or CBAM);

[0167] Function: Performs detailed analysis and precise regression only on the "region of interest" output by the A stream; it is responsible for accurately classifying the target (whether it is a crane or an excavator) and precisely locating its bounding box.

[0168] 3. Progressive Integration Module

[0169] Function: A bridge connecting flow A and flow B;

[0170] Working method: The candidate region coordinates output by the A-stream detection head guide the B-stream detection head to focus its attention on the corresponding position in the original feature map and perform "cropping-enlarging-precision judgment" operation; this avoids the B-stream performing indiscriminate intensive calculations on the entire image, greatly saving computing power;

[0171] Structural advantages summary: This "coarse screening first, then fine judgment" structure concentrates computing resources where they are most needed, achieving a perfect balance between speed and accuracy, and is more efficient and targeted than directly using standard lightweight models.

[0172] The training strategy for the above-mentioned model, which combines the object detection model and the acoustic feature model into a single model, includes three training phases:

[0173] Phase 1: Large-scale pre-training and multimodal knowledge distillation

[0174] Teacher model training: On the server side, a large, high-precision but computationally intensive model (such as YOLOv8-X or Swing Transformer) is used as the "teacher model".

[0175] Key innovations: In addition to training the teacher model with image data, an "audiovisual fusion dataset" was constructed; each video in the dataset is accompanied by synchronized audio tags (such as "crane engine sound" and "ambient wind sound"); the teacher model learns the correlation between sound and visual targets while learning visual features.

[0176] Knowledge distillation of the student model (the lightweight model used in this case): The goal is to enable the "student model" (i.e., the aforementioned two-stream network) to learn the knowledge of the "teacher model";

[0177] Traditional distillation: Students imitate the teacher's classification and localization results for images;

[0178] The multimodal distillation in this case includes: visual-auditory alignment distillation and cross-modal attention distillation, among which...

[0179] Visual-auditory aligned distillation: When training the student model, not only images are input, but also the corresponding audio spectrograms are input; the intermediate layer features of the student model are required to be aligned not only with the visual features of the teacher model, but also with the audio features semantically; for example, when hearing "crane sound", even if the student model cannot see it clearly visually, its internal features must be "pulled" to the feature representation of "crane" in the teacher model.

[0180] Cross-modal attention distillation: The teacher model learns "which areas of the image should be paid more attention to when hearing the sound of a crane"; this "sound-guided attention map" is also used as a supervisory signal for the student model to learn.

[0181] In this way, the lightweight model "hears" sound early in the training process and learns the relationship between sound and images. This enables it to make more robust judgments when combined with sound information, even when visual information is blurred (such as in fog or at night), laying a solid foundation for subsequent audiovisual fusion modules.

[0182] Phase Two: Difficult Case Discovery and Focus Sampling

[0183] Objective: To enable the model to focus on learning the most difficult and easily confused samples;

[0184] Method: Use the model trained in Phase 1 to make predictions on the training set and identify all samples with incorrect predictions (hard examples); perform manual analysis and labeling on these hard examples, especially those samples that "look similar but are not" (such as tower crane and crane, green vehicle and tree).

[0185] In the next round of training, the sampling weights for these difficult examples are significantly increased, allowing the model to learn them repeatedly until it masters the ability to distinguish them.

[0186] Phase 3: Scene Adaptive Fine-tuning

[0187] Objective: To adapt the model to the specific environment in which the device will be deployed, thus solving the problem of "incompatibility".

[0188] Method: Install the device on an actual utility pole and collect real-world environmental data (including various weather conditions, lighting, and background) for 1-2 weeks.

[0189] Semi-automatic annotation of these data: initial annotation is performed using the current model, followed by rapid manual verification;

[0190] Using this "scenario-customized" small dataset, the pre-trained model can be fine-tuned. At this point, most of the parameters of the backbone network can be frozen, and only the detection head part can be trained, achieving a significant improvement in model performance at minimal cost.

[0191] In summary, the aforementioned "dual-stream progressive" detection network, through a "coarse screening-fine judgment" mechanism, achieves intelligent allocation of computing resources under the premise of lightweight design, solving the efficiency or accuracy bottlenecks caused by the "one-size-fits-all" approach of traditional models.

[0192] Furthermore, by employing multimodal knowledge distillation, "sound" is introduced as a supervisory signal into the distillation process of a pure visual object detection model for the first time, allowing the lightweight model to "learn to listen" during the training phase, thus giving it an inherent potential for audiovisual fusion.

[0193] Furthermore, through scene-adaptive fine-tuning, a rapid deployment process from general models to specific scenarios was established, ensuring the high accuracy and robustness of the models in the real world and solving the common pain point of AI models "performing well in the laboratory but poorly in the field".

[0194] Example 5: This embodiment of the invention assumes a scenario in which:

[0195] Parameter configuration: ;

[0196] Weight configuration: (Emphasizing the importance of distance).

[0197] Scene 1: A crane approaches at high speed:

[0198] ,

[0199] ,

[0200] (Crane),

[0201] (Audiovisual confirmation)

[0202] calculate:

[0203]

[0204]

[0205]

[0206] in conclusion: This triggered a Level 3: Expulsion Warning.

[0207] Scene 2: A tree slowly approaches in a gentle breeze:

[0208] ,

[0209] (Below the safety threshold)

[0210] (Trees),

[0211] (High-confidence visual)

[0212] calculate:

[0213]

[0214]

[0215]

[0216] in conclusion: If the system determines there is no threat, it will not issue an alarm.

[0217] Example 6: As Figure 2As shown, this embodiment of the invention discloses an intelligent identification and early warning system for external damage to overhead lines, using an intelligent identification and early warning method for external damage to overhead lines, including: a central processing and communication module, wherein the central processing and communication module includes a signal acquisition unit, a target identification unit, a target determination unit and an adaptive early warning unit;

[0218] The signal acquisition unit collects environmental data signals around the overhead line, including multi-view visual data signals and environmental sound signals.

[0219] The target recognition unit constructs a target detection model, processes multi-view visual data signals, identifies key targets, performs three-dimensional spatial positioning of key targets, and calculates the spatial distance and relative motion trend between key targets and overhead lines; among them, key targets are construction machinery (cranes, excavators), trees, or pedestrians.

[0220] The target identification unit, based on the calculated spatial distance and relative motion trend between the key target and the overhead line, combined with environmental sound signals, further identifies the key target to obtain the true identification result.

[0221] The adaptive early warning unit constructs a quantitative assessment model of the threat situation that incorporates multi-dimensional information fusion based on the actual identification results. It introduces a dynamic threat index and performs hierarchical adaptive early warning based on the dynamic threat index. The multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation degree.

[0222] The central processing and communication module includes an embedded AI processor, such as an NPU, which can be used to run intelligent recognition algorithms; and it also integrates a 4G / 5G wireless communication module for data transmission with the back-end monitoring center.

[0223] In addition, the intelligent recognition algorithm can be configured in the monitoring center. The central processing and communication module only needs to send the image and sound information to the monitoring center, which will then perform external intelligent recognition.

[0224] The system also includes a power supply module, which is connected to the central processing and communication module. The power supply module can be powered by a combination of solar panels and batteries, achieving self-sufficiency in system power.

[0225] The aforementioned adaptive warning unit also includes a high-brightness LED warning light and a high-power speaker.

[0226] Among them, the high-brightness LED warning light is used to emit a strong flashing warning light at night or in low light conditions;

[0227] High-power speaker: Used to play preset voice warning messages, such as "High voltage danger, please stop work immediately." High-brightness LED warning lights and high-power speakers can be selected individually or both to better enhance the warning function.

[0228] Example 7: This embodiment of the invention discloses an electronic device, including a processor and a memory. The memory stores a computer program, which is loaded and executed by the processor to realize an intelligent identification and early warning method for external damage to overhead lines.

[0229] The aforementioned electronic device also includes transmission devices and input / output devices, wherein both the transmission devices and the input / output devices are connected to the processor.

[0230] The processor described above can be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. It can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The memory can include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory, portable hard drives, magnetic disks, or optical disks.

[0231] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0232] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.

[0233] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0234] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.

Claims

1. A method for intelligent identification and early warning of external damage to overhead power lines, characterized in that, include: Collect environmental data signals around overhead lines, including multi-view visual data signals and environmental sound signals; A target detection model is constructed to process multi-view visual data signals, identify key targets, perform three-dimensional spatial positioning of key targets, and calculate the spatial distance and relative motion trend between key targets and overhead lines; among them, key targets are construction machinery, trees or pedestrians; Based on the calculated spatial distance and relative motion trend between the key target and the overhead line, combined with environmental sound signals, further identification is performed to obtain the true identification results of the key target; Based on the actual identification results, a quantitative assessment model of the threat situation is constructed, which incorporates multi-dimensional information fusion. A dynamic threat index is introduced, and a hierarchical adaptive early warning is carried out based on the dynamic threat index. The multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation degree. Based on the actual identification results, a quantitative assessment model of the threat situation is constructed, incorporating multi-dimensional information fusion, and a dynamic threat index is introduced. The multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation level, including: Dynamic Threat Index DTI as follows: DTI = f ( D , V , T , C ), In the formula, D Spatial distance between key targets and overhead power lines, unit: meters; V Radial velocity of the critical target toward the overhead line, in meters per second; T The threat type coefficient for key targets is a weight preset through machine learning or expert experience, ranging from [0, 1]. C The audiovisual fusion confirmation score, i.e., the overall confidence multiplier, ranges from [0, 1]. The dynamic threat index DTI The spatial distance and radial velocity in the data are normalized, where the spatial distance is... D The normalization formula is as follows: , In the formula, This is the normalized spatial distance threat value, ranging from [0, 1]. Distance to the midpoint of risk; This is the distance sensitivity coefficient; Among them, radial velocity V The normalization formula is as follows: , In the formula, This is the normalized radial velocity threat value, ranging from [0, 1]. The safe speed threshold; Maximum dangerous speed; Based on the above normalized parameters, the final dynamic threat index is obtained. DTI The calculation method is as follows: , In the formula, This represents the final dynamic threat index, ranging from [0, 1]. The weighting index for spatial distance. The weighting index for radial velocity, As a weighting index for threat types, and ; This refers to the audiovisual fusion confirmation degree, which is the overall confidence multiplier.

2. The intelligent identification and early warning method for external damage to overhead lines according to claim 1, characterized in that, The constructed target detection model processes multi-view visual data signals, identifies key targets, performs three-dimensional spatial positioning of key targets, and calculates the spatial distance and relative motion trend between key targets and overhead lines, including: The object detection model was trained on an image dataset containing scenes of construction machinery, trees, and pedestrians. The object detection model was trained and quantized using the YOLOv5s model.

3. The intelligent identification and early warning method for external damage to overhead lines according to claim 1, characterized in that, The calculated spatial distance and relative motion trend between the key target and the overhead line, combined with environmental sound signals, are used for further identification to obtain the true identification results of the key target, including: After identifying the key target, determine whether the spatial distance between the key target and the overhead line is less than the corresponding safety threshold; In response, acoustic analysis is triggered, which involves extracting the sound signal from the environmental sound signal in the direction of the key target and performing spectral matching on the sound signal. Determine whether the matched sound signal is the sound signal corresponding to the identified key target; If the response is correct, then the identification result is correct, and the key target is identified as the true identification result.

4. The intelligent identification and early warning method for external damage to overhead lines according to claim 1, characterized in that, The audiovisual fusion confirmation, i.e., the overall confidence multiplier, includes single-modal internal evaluation, cross-modal cross-validation, and spatiotemporal consistency verification. (1) The intramodal assessment includes visual confidence and acoustic confidence, wherein: Visual confidence The output is from the object detection model, and the calculation method is as follows: , In the formula, Confidence_score_of_all_crane_objects represents the confidence score of all "construction machinery" objects in the image; This indicates how confident the model is in identifying a certain region in the image as "construction machinery"; Acoustic confidence The sound classification model outputs a probability distribution. This indicates how confident the sound classification model is in identifying the current ambient sound as "the sound of construction machinery." (2) Cross-modal cross-validation: Define a correlation score, which is a Boolean value or a soft score, including spatial consistency and state consistency, where: Spatial consistency The system uses microphone array sound source localization technology to estimate the direction of the sound source, compares this sound direction with the location of the "construction machinery" target in the visually detected image, and determines whether the two directions are basically consistent. If so, then... If the response is no, then ;· State Consistency The system determines whether the "states" described by vision and hearing match, and responds accordingly. If the response is no, then ; The comprehensive correlation score can be obtained from the above. : , In the formula, and It's weight. ; (3) Spatiotemporal consistency verification: using a sliding window, calculate the comprehensive correlation score within the sliding window. The continuity and stability of the values ​​are calculated as follows: , In the formula, For spatiotemporal consistency score, ; For comprehensive relevance score; N is the total number of frames in the window, S i It refers to the first i Frame metrics score, It is an indicator function; its value is 1 when the condition inside the parentheses is true, and 0 otherwise. By integrating all the above information, the final audiovisual fusion confirmation is obtained. The calculation formula is as follows: , In the formula, Visual confidence level; Acoustic confidence level; Visual weight; Auditory weighting; The overall relevance score; The score is based on spatiotemporal consistency.

5. The intelligent identification and early warning method for external damage to overhead lines according to claim 4, characterized in that, The hierarchical adaptive early warning based on the dynamic threat index includes: The warning system is divided into four levels, specifically: Level 1: Sensing and Early Warning DTI ∈[0.2, 0.4); Warning action: LED warning lights slowly turn on and off; Level 2: Warning / Alert DTI ∈[0.4, 0.7); Warning action: LED warning light flashes rapidly in yellow or orange, and a speaker plays a short, neutral sound; Level 3: Deportation Warning DTI ∈[0.7, 0.9); Warning action: LED warning light flashes red, speaker plays the corresponding preset voice command at high volume in a loop; Level 4: Emergency Evacuation and Reporting DTI ∈[0.9, 1.0]; Warning action, maintain the warning action of level three, and at the same time remotely report "level one emergency alarm" to the background, as well as relevant threat proof information; among which threat proof information includes real-time snapshot, pre-incident video, situation snapshot and audio clip.

6. An intelligent identification and early warning system for external damage to overhead power lines, characterized in that, The overhead line external damage intelligent identification and early warning method as described in any one of claims 1 to 5 includes: a central processing and communication module, wherein the central processing and communication module includes a signal acquisition unit, a target identification unit, a target determination unit and an adaptive early warning unit; The signal acquisition unit collects environmental data signals around the overhead line, including multi-view visual data signals and environmental sound signals. The target recognition unit constructs a target detection model, processes multi-view visual data signals, identifies key targets, performs three-dimensional spatial positioning of key targets, and calculates the spatial distance and relative motion trend between key targets and overhead lines; among them, key targets are construction machinery, trees, or pedestrians; The target identification unit, based on the calculated spatial distance and relative motion trend between the key target and the overhead line, combined with environmental sound signals, further identifies the key target to obtain the true identification result. The adaptive early warning unit constructs a quantitative assessment model of the threat situation that incorporates multi-dimensional information fusion based on the actual identification results. It introduces a dynamic threat index and performs hierarchical adaptive early warning based on the dynamic threat index. The multi-dimensional information includes spatial distance, radial velocity, threat type, and confirmation degree.

7. The intelligent identification and early warning system for external damage to overhead power lines according to claim 6, characterized in that, It also includes a power supply module, which is connected to the central processing and communication module.

8. The intelligent identification and early warning system for external damage to overhead power lines according to claim 6, characterized in that, The adaptive warning unit also includes high-brightness LED warning lights and a high-power speaker.

9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, which is loaded and executed by the processor to implement the intelligent identification and early warning method for external damage to overhead lines as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Overhead line threat grading evaluation method and system based on dynamic trajectory tracking

    CN120851626A

  • Power transmission line external damage risk identification method fusing image and sound features

    CN120929766A