A method and system for acquiring training data for a power transmission line fault prediction model

By constructing a multi-source heterogeneous data pool and performing unified spatiotemporal benchmark processing and gridding, strong convective events are identified and their correlation credibility is calculated. This solves the problem of fragmented transmission line fault data, enables the acquisition of high-quality training data for transmission line fault prediction models, and improves the accuracy and reliability of the prediction models.

CN121256370BActive Publication Date: 2026-04-03STATE GRID JIANGXI ELECTRIC POWER CO LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, there is a lack of integration mechanisms for multi-dimensional data on factors affecting transmission line faults under severe convective weather, resulting in incomplete data coverage and fragmented information. This makes it impossible to accurately establish the correlation between line segments and the time of fault, affecting the accuracy and reliability of prediction models.

Method used

A multi-source heterogeneous raw data pool is constructed, and spatiotemporal benchmarks are uniformly processed and gridded. Strong convective events are identified and affected areas are divided. Grid exposure and terrain correction weights are calculated. The association credibility is calculated by combining time factors. High-credibility association matching data are selected and attribution labeling is performed to obtain the training dataset.

Benefits of technology

This method enables multi-dimensional quantitative matching between transmission line faults and severe convective events, ensuring the accuracy and reliability of spatiotemporal correlation, eliminating spurious correlation data, and improving the quality of training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121256370B_ABST
    Figure CN121256370B_ABST
Patent Text Reader

Abstract

This invention relates to the field of meteorological data processing technology, specifically to a method and system for acquiring training data for a transmission line fault prediction model. The invention constructs a multi-source heterogeneous raw data pool containing severe convective meteorological data, transmission line operation and fault data, and geospatial data; it performs unified spatiotemporal benchmark processing on the data and divides it into grids; it maps line segments to grid cells, identifies severe convective events and delineates affected areas based on multi-parameter collaborative thresholds using density clustering algorithms, calculates grid exposure, and calculates spatial correlation strength by combining intensity level weights and terrain correction weights; it calculates time factors by setting differentiated time correlation windows according to the dominant type of severe convection, calculates correlation credibility based on spatial correlation strength and time factors, and selects high-credibility correlation results; and it performs attribution labeling on fault information to obtain the training dataset. This invention solves the problem of low training data quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of meteorological data processing technology, specifically to a method and system for acquiring training data for a power transmission line fault prediction model. Background Technology

[0002] As power systems develop towards higher voltage, larger capacity, and wider coverage, the operational safety of transmission lines is increasingly affected by severe convective weather. To reduce the damage caused by severe convective weather to transmission lines, the industry generally builds failure rate prediction models to achieve early warning of failure risks. High-quality training data is the core prerequisite for ensuring the accuracy and reliability of prediction models.

[0003] In existing technologies, the influencing factors of line faults under severe convective weather involve multi-dimensional data including meteorological, line operation, and geospatial factors. Furthermore, the data formats and storage standards are inconsistent, and there is a lack of effective integration mechanisms, resulting in data that cannot be directly correlated, leading to incomplete data coverage and fragmented information. In addition, existing methods lack precise spatiotemporal matching mechanisms, making it difficult to accurately establish the correlation between severe convective events, line segments, and the time of fault. This also makes it difficult to effectively eliminate spurious correlations, ultimately resulting in low-quality training data and affecting the accuracy and reliability of the prediction model. Summary of the Invention

[0004] This invention provides a method and system for acquiring training data for a power transmission line fault prediction model to solve the problem of low training data quality in the prior art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] This invention provides a method for acquiring training data for a transmission line fault prediction model, comprising:

[0007] S100: Construct a multi-source heterogeneous raw data pool, which includes severe convective weather data, power transmission line operation and fault data, and geospatial data.

[0008] S200: Perform spatiotemporal benchmark unification processing on the data in the multi-source heterogeneous raw data pool, and perform gridding based on the preset spatiotemporal granularity to obtain spatiotemporal gridded data;

[0009] S300: Divide the transmission line into line segments according to the latitude and longitude of the towers and map them to the corresponding grid cells of the spatiotemporal gridded data; identify strong convective events and divide the affected areas, and assign intensity level weights to each affected area; calculate the grid exposure of the line segments, and calculate the spatial correlation intensity by combining the intensity level weights and the terrain correction weights extracted based on geospatial data.

[0010] S400: Calculate the time factor based on the time relationship between the fault time and the strong convection occurrence time; calculate the association credibility based on the spatial association strength and the time factor by weighted summation, and select association results with credibility higher than a preset threshold as association matching data;

[0011] S500: Attribution labeling is performed on the transmission line fault information in the associated matching data to obtain a training dataset. The labeling content includes the correspondence between fault type, strong convection cause type and its characteristic parameters.

[0012] As a preferred technical solution of the present invention, the identification of strong convective events and the division of the affected area include:

[0013] Multi-parameter collaborative thresholds were set based on radar reflectivity factor, satellite cloud top brightness temperature, and ground automatic station observation data.

[0014] The density clustering algorithm is used to aggregate adjacent grids that meet the multi-parameter collaborative threshold to identify strong convection events;

[0015] Based on the intensity distribution of the severe convective event decreasing outward from the core area, the core influence area, significant influence area, and peripheral influence area are divided from the inside out, and each influence area is assigned a corresponding intensity level weight.

[0016] As a preferred embodiment of the present invention, the calculation of the grid exposure of the line segment includes:

[0017] The transmission line is divided into basic segments according to the tower coordinates, and then further subdivided according to the key geographical features within the line corridor to generate several fine segments.

[0018] The fine line segment is spatially cut with the spatiotemporal gridded data to obtain a sub-line segment that falls completely into a single grid cell;

[0019] The weighted length is calculated based on the length of each sub-segment and the attribute parameters of the line to which it belongs;

[0020] The grid exposure is obtained by summing the weighted lengths of all sub-segments falling within the same grid and dividing by the grid edge length.

[0021] As a preferred embodiment of the present invention, the terrain correction weights extracted from geospatial data are obtained through the following method:

[0022] Extract the slope, aspect, and relative elevation parameters of the grid from the digital elevation model data;

[0023] The weighting coefficients for the slope, aspect, and relative elevation parameters are fitted based on historical fault data.

[0024] The slope, aspect, and relative altitude parameters are weighted and summed according to their corresponding weight coefficients, and then normalized to obtain the terrain correction weights.

[0025] As a preferred embodiment of the present invention, the computational spatial correlation strength includes:

[0026] The environmental weight of the route corridor is obtained by weighted summation of the threat level of trees, the threat level of buildings and the terrain passage index within the route corridor.

[0027] Multiply the intensity level weight, the terrain correction weight, and the route corridor environment weight to obtain the spatial comprehensive weight of the grid;

[0028] The spatial correlation strength of a line segment is obtained by multiplying the grid exposure of each grid along the route with the spatial comprehensive weight of the corresponding grid and summing the results.

[0029] As a preferred embodiment of the present invention, the calculation of the time factor based on the time relationship between the fault time and the occurrence time of strong convection includes:

[0030] Differentiated time-related windows are set based on the dominant type of severe convective events;

[0031] Calculate the time difference between the time of the line fault and the peak time of the strong convection;

[0032] The time factor is determined based on the inverse ratio of the time difference to the total duration of the time-related window.

[0033] As a preferred embodiment of the present invention, the attribution annotation includes:

[0034] Based on the operation and maintenance records, fault waveforms, or on-site investigation results in the transmission line fault data, the fault type is marked.

[0035] Based on the dominant characteristics of strong convection events in the aforementioned correlation matching data, the types of strong convection triggers are labeled;

[0036] The peak values ​​of meteorological parameters corresponding to the type of severe convection are extracted as the characteristic parameters.

[0037] This invention also proposes a training data acquisition system for a transmission line fault prediction model, comprising:

[0038] The data pool construction module is used to construct a multi-source heterogeneous raw data pool, which includes severe convective meteorological data, power transmission line operation and fault data, and geospatial data.

[0039] The data preprocessing module is used to perform spatiotemporal benchmark unification processing on the data in the multi-source heterogeneous raw data pool, and to perform gridding based on a preset spatiotemporal granularity to obtain spatiotemporal gridded data.

[0040] The spatial association matching module is used to divide the transmission line into line segments according to the tower latitude and longitude and map them to the corresponding grid cells of the spatiotemporal gridded data; identify strong convective events and divide the affected areas, and assign intensity level weights to each affected area; calculate the grid exposure of the line segments, and calculate the spatial association intensity by combining the intensity level weights and the terrain correction weights extracted based on geospatial data.

[0041] The time association filtering module is used to calculate the time factor based on the time relationship between the fault time and the strong convection occurrence time; and to calculate the association credibility based on the spatial association strength and the time factor by weighted summation, and to filter the association results with credibility higher than a preset threshold as association matching data.

[0042] The attribution annotation module is used to perform attribution annotation on the transmission line fault information in the associated matching data to obtain a training dataset. The annotation content includes the correspondence between fault type, strong convection cause type and its characteristic parameters.

[0043] The present invention also proposes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for acquiring training data for a transmission line fault prediction model.

[0044] The present invention also proposes a readable storage medium, characterized in that the readable storage medium stores a computer program, which, when executed by a processor, implements the above-described method for acquiring training data for a transmission line fault prediction model.

[0045] The beneficial effects of this invention are:

[0046] 1. This invention achieves multi-dimensional quantitative matching between railway line segments and severe convective events by dividing the impact zone of severe convection and assigning intensity level weights, and combining grid exposure, terrain correction weights, and railway corridor environmental weights to calculate spatial correlation strength. This refined spatial correlation calculation mechanism can accurately assess the actual impact of severe convection of different intensities on specific railway line segments.

[0047] 2. This invention achieves precise matching and filtering in both spatiotemporal dimensions by setting differentiated time correlation windows based on the dominant type of strong convection and calculating time factors, combined with spatial correlation strength and peak intensity of strong convection to calculate correlation credibility. The credibility threshold filtering effectively eliminates spurious correlation data, ensuring the accuracy and reliability of the correlation between "strong convection event - line segment - fault time". Attached Figure Description

[0048] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0049] Figure 1 This is a flowchart illustrating a method for acquiring training data for a transmission line fault prediction model according to the present invention.

[0050] Figure 2 This is a schematic diagram of the structure of a training data acquisition system for a power transmission line fault prediction model according to the present invention. Detailed Implementation

[0051] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0052] Example 1: As Figure 1 As shown, the present invention provides a method for acquiring training data for a transmission line fault prediction model, comprising:

[0053] S100: Construct a multi-source heterogeneous raw data pool, which includes severe convective weather data, power transmission line operation and fault data, and geospatial data.

[0054] Specifically, in this embodiment, constructing a multi-source heterogeneous raw data pool involves integrating key information from different sources, with different data types and formats, that support the training of a transmission line fault rate prediction model under severe convective weather into a unified initial data set. The multi-source heterogeneous raw data pool includes severe convective meteorological data, transmission line operation and fault data, and geospatial data.

[0055] Severe convective weather data primarily encompasses radar monitoring data, satellite remote sensing data, and field-measured data from meteorological departments. These data collectively contribute to accurately capturing the timing, extent, and intensity characteristics of severe convective weather events. Specifically, radar monitoring data includes radar reflectivity factors (used to characterize precipitation intensity and cloud structure) and Doppler radial wind speed data; satellite remote sensing data includes satellite cloud top brightness temperature data (used to identify the vertical development intensity of deep convective clouds) and infrared and visible light channel imagery; and field-measured data from automatic ground stations includes real-time observations of conventional meteorological elements such as surface wind speed, precipitation, temperature, and air pressure. Additionally, data on ground-based lightning and cloud-based lightning recorded by lightning location systems includes information such as the time, location, and intensity of lightning strikes.

[0056] Transmission line operation and fault data originate from the power grid dispatch and maintenance management system. This data is directly related to the operating status of transmission lines and the specific circumstances of fault occurrence. Specifically, basic transmission line information includes static parameters such as line number, voltage level, line route, tower number, tower latitude and longitude coordinates, tower type, conductor type, and insulator type; line operation data includes real-time operating parameters such as line load, current, voltage, and power; fault record data includes fault occurrence time (accurate to the minute or second), faulted line number, fault type (such as tripping, line breakage, flashover, etc.), fault waveform data, protection action information, and on-site investigation report after the fault.

[0057] Geospatial data, sourced from a geographic information service platform, primarily comprises digital elevation model (DEM) data, high-resolution remote sensing imagery, and specific environmental data for the railway corridor. This data is used to analyze the impact of the geographical environment on the intensity of severe convective weather and the exposure risk of the railway line under different geographical conditions. Specifically, DEM data provides elevation information of the regional terrain, including topographic parameters such as altitude, slope, and aspect; high-resolution remote sensing imagery is used to identify land cover types, including vegetation distribution, building locations, and water body distribution; and specific environmental data for the railway corridor includes environmental information directly related to railway safety, such as tree height and distribution, building height and distance, and terrain corridor characteristics.

[0058] By integrating the above three types of data into a multi-source heterogeneous raw data pool, it is ensured that the subsequent training data acquisition process can be carried out based on all key influencing factors, thereby guaranteeing the comprehensiveness and effectiveness of the final training dataset.

[0059] S200: Perform spatiotemporal benchmark unification processing on the data in the multi-source heterogeneous raw data pool, and perform gridding based on the preset spatiotemporal granularity to obtain spatiotemporal gridded data;

[0060] Specifically, the first step is to unify the time base by converting the UTC time of weather radar data, the local time of transmission line fault data, and the equipment recording time of lightning location data into Coordinated Universal Time (UTC) or local standard time. For data with different sampling frequencies, time interpolation or resampling is used to ensure consistent time resolution.

[0061] Subsequently, spatial benchmark unification was carried out, converting the Gauss-Krüger projection used in the meteorological grid, the WGS-84 latitude and longitude of the transmission line towers, and the local coordinate system of the geographic topographic data into the WGS-84 coordinate system. The positional deviation caused by different projection methods was corrected through coordinate transformation algorithms.

[0062] Based on the impact range and rate of change of severe convective weather, as well as the spatial distribution characteristics of transmission lines, a uniform time interval and spatial unit size are pre-set. The time granularity can be selected as a time slice of 5 minutes or 10 minutes, and the spatial granularity can be selected as a grid unit of 1km×1km or 2km×2km.

[0063] Data that has been unified by a spatiotemporal reference is mapped to the corresponding spatiotemporal grid according to a preset granularity, so that each grid cell contains all meteorological data, route data and geographical data within the time period and spatial range, resulting in spatiotemporal gridded data.

[0064] S300: Divide the transmission line into line segments according to the latitude and longitude of the towers and map them to the corresponding grid cells of the spatiotemporal gridded data; identify strong convective events and divide the affected areas, and assign intensity level weights to each affected area; calculate the grid exposure of the line segments, and calculate the spatial correlation intensity by combining the intensity level weights and the terrain correction weights extracted based on geospatial data.

[0065] Specifically, the transmission line is divided into segments based on the latitude and longitude of the towers. More specifically, using the latitude and longitude coordinates of two adjacent towers as the starting and ending points, the entire line is divided into several inter-tower segments.

[0066] Based on the latitude and longitude range of each line segment, determine its corresponding grid cell in the spatiotemporal gridded data. By comparing the latitude and longitude of the line segment's starting point, ending point, and key nodes in the middle of the line segment with the spatial boundaries of each grid cell, determine which grid the line segment falls into completely or partially, thus achieving a precise spatial correspondence between line segments and spatiotemporal grids.

[0067] Furthermore, the identification of strong convective events and the delineation of affected areas include:

[0068] Multi-parameter collaborative thresholds were set based on radar reflectivity factor, satellite cloud top brightness temperature, and ground automatic station observation data.

[0069] The density clustering algorithm is used to aggregate adjacent grids that meet the multi-parameter collaborative threshold to identify strong convection events;

[0070] Based on the intensity distribution of the severe convective event decreasing outward from the core area, the core influence area, significant influence area, and peripheral influence area are divided from the inside out, and each influence area is assigned a corresponding intensity level weight.

[0071] Specifically, a multi-parameter collaborative threshold is set based on radar reflectivity factor, satellite cloud top brightness temperature, and ground automatic station observation data.

[0072] Radar reflectivity factor is used to capture the concentration and size of precipitation particles within severe convective clouds, with a threshold such as Z≥45dBZ indicating the core of a severe thunderstorm. Satellite cloud top brightness temperature is used to identify the vertical development intensity of deep convective clouds, with a threshold such as TBB≤-50℃ representing severe convective clouds. Ground automatic weather station data is used to verify the actual impact of severe convection on the ground, with thresholds such as maximum wind speed V≥17.2m / s and 1-hour rainfall R≥20mm. These three factors work together to form a screening standard; only when a grid simultaneously meets multiple parameter thresholds is it determined to be a grid affected by severe convection, avoiding misjudging non-severe convective weather based on a single parameter.

[0073] Density clustering algorithms (such as DBSCAN) are used to aggregate adjacent grids that meet multiple parameter collaborative thresholds. Specifically, by setting a spatial neighborhood radius and a minimum number of cluster points, spatially continuous grids that meet the intensity criteria are grouped together, thereby identifying single, independent strong convective events and distinguishing between different, discontinuous strong convective events.

[0074] Based on the decreasing intensity distribution of severe convective events from the core area outwards, the impact zone is divided into three areas from the inside out: the core impact zone, the significant impact zone, and the peripheral impact zone. Specifically: the core impact zone is the area with the highest intensity of severe convection and the greatest threat to power lines, such as grids with a radar reflectivity factor ≥ 50 dBZ; the significant impact zone is the area with slightly lower intensity but still a significant threat, such as grids with a radar reflectivity factor between 45 and 50 dBZ; and the peripheral impact zone is the area with weaker intensity and lower threat, such as grids with a radar reflectivity factor between 40 and 45 dBZ.

[0075] Each affected area is assigned a corresponding intensity level weight. For example, the core affected area has a weight of 1.0, the significant affected area has a weight of 0.7, and the peripheral affected area has a weight of 0.4. The specific weight values ​​can be obtained based on statistical analysis of historical failure data.

[0076] Furthermore, the calculation of the grid exposure of the line segment includes:

[0077] The transmission line is divided into basic segments according to the tower coordinates, and then further subdivided according to the key geographical features within the line corridor to generate several fine segments.

[0078] The fine line segment is spatially cut with the spatiotemporal gridded data to obtain a sub-line segment that falls completely into a single grid cell;

[0079] The weighted length is calculated based on the length of each sub-segment and the attribute parameters of the line to which it belongs;

[0080] The grid exposure is obtained by summing the weighted lengths of all sub-segments falling within the same grid and dividing by the grid edge length.

[0081] Specifically, after dividing the transmission line into basic segments according to the tower coordinates, it is further subdivided according to the key geographical features within the line corridor to generate several fine segments.

[0082] Key geographical features include: topographic transition features such as transitional zones between plains and mountains, and the location of valley entrances and exits; abrupt environmental changes such as the boundaries between dense forests and open farmland, and the division between towns and suburbs; and terrain features sensitive to severe convection such as narrow passages between mountains and areas with a history of frequent thunderstorms.

[0083] By using GIS spatial overlay analysis, the identified key geographic features are spatially matched with basic line segments to determine the specific coordinates of the features on the line segments. Using the feature coordinates as the center, a preset distance (e.g., 50 meters) is extended to both ends of the line segment as subdivision endpoints, dividing the basic line segment into multiple fine-grained line segments. This ensures that each fine-grained line segment covers only a single type of key geographic feature.

[0084] The fine line segments are spatially segmented with the spatiotemporal gridded data. Using GIS spatial analysis tools, the coordinates of the intersection points between the fine line segments and the grid boundaries are calculated. Fine line segments that may span multiple grids are broken down into shorter sub-segments, ensuring that each sub-segment falls entirely within a single grid cell.

[0085] The weighted length is calculated based on the length of each sub-segment and the attribute parameters of its corresponding line. These line attribute parameters characterize the importance of the transmission line, including but not limited to factors such as line voltage level, line load level, and the importance of the power supply area. For example, based on line voltage level, a 1000kV UHV line has a weight of 1.2, a 500kV EHV line has a weight of 1.0, and a 220kV HV line has a weight of 0.8; based on line load level, a line with a super-high load has a weight of 1.2, a line with a first-high load has a weight of 1.0, and a line with a second-high load has a weight of 0.8. The formula for calculating the weighted length of each sub-segment is: Weighted length = Actual length of sub-segment × Weight coefficient of its corresponding line.

[0086] The total weighted exposure length of the lines within a grid is obtained by summing the weighted lengths of all sub-segments falling within the same grid. The grid exposure is then divided by the grid edge length. For example, if a grid edge length is 1 km and the total weighted length of its sub-segments is 800 m, the grid exposure is 0.8.

[0087] Furthermore, the terrain correction weights extracted from geospatial data are obtained in the following way:

[0088] Extract the slope, aspect, and relative elevation parameters of the grid from the digital elevation model data;

[0089] The weighting coefficients for the slope, aspect, and relative elevation parameters are fitted based on historical fault data.

[0090] The slope, aspect, and relative altitude parameters are weighted and summed according to their corresponding weight coefficients, and then normalized to obtain the terrain correction weights.

[0091] Specifically, the slope, aspect, and relative elevation parameters of the grid are extracted from the digital elevation model (DEM) data. The slope is obtained by calculating the rate of change of elevation gradient within the grid, reflecting the steepness of the terrain; the aspect is obtained by calculating the direction of the elevation gradient, reflecting the orientation of the slope; and the relative elevation is the difference between the elevation within the grid and the average elevation of the surrounding area, reflecting the prominence of the terrain.

[0092] Weighting coefficients for slope, aspect, and relative altitude parameters were fitted based on historical fault data. By statistically analyzing the line fault rates corresponding to different terrain parameters under historical severe convective weather conditions, regression analysis was used to fit the weighting coefficients for each parameter. .

[0093] The slope, aspect, and relative elevation parameters are weighted and summed according to their respective weighting coefficients:

[0094] ;

[0095] The initial values ​​are normalized by dividing them by the maximum initial value in the region to obtain terrain correction weights in the range of 0 to 1.

[0096] Furthermore, the computational spatial correlation strength includes:

[0097] The environmental weight of the route corridor is obtained by weighted summation of the threat level of trees, the threat level of buildings and the terrain passage index within the route corridor.

[0098] Multiply the intensity level weight, the terrain correction weight, and the route corridor environment weight to obtain the spatial comprehensive weight of the grid;

[0099] The spatial correlation strength of a line segment is obtained by multiplying the grid exposure of each grid along the route with the spatial comprehensive weight of the corresponding grid and summing the results.

[0100] Specifically, the environmental weight of the route corridor is obtained by weighted summation of the threat level of trees, the threat level of buildings, and the terrain access index within the route corridor.

[0101] Identify trees along the route corridor and extract the height of each tree. and horizontal distance to the line The threat level of a single tree is calculated using the following nonlinear function:

[0102] ;

[0103] in, Tree height, The horizontal distance from the tree to the line. Set the high threshold for the line safety tree (e.g., 15m). The safe distance threshold (e.g., 30m). This is an adjustment factor for the influence of tree height (e.g., 1.5). This is a distance-related adjustment factor (e.g., 2.0). It is a natural exponential function.

[0104] The total threat value is obtained by summing the threat levels of all trees in the corridor. It is then normalized based on the maximum total threat value of the same type of corridor in history to obtain the forest impact weight between 0 and 1.

[0105] Building threat level calculation involves identifying buildings within the route corridor and extracting the height and horizontal distance of each building to the route. A threshold function is set based on the distance and height; buildings with a height exceeding a certain threshold and a distance to the route less than a safe distance are assigned a higher threat weight, while those with a lower threshold are assigned a lower weight. The weights are then accumulated and normalized to obtain the building impact weights.

[0106] The terrain access index is calculated based on digital elevation model data. Cross sections are taken at fixed intervals (e.g., 500m) along the route, and the width W, relative height H, slope S, and terrain continuity C parameters of each cross section are extracted. First, each parameter is normalized.

[0107] ;

[0108] ;

[0109] ;

[0110] in, , These represent the maximum and minimum cross-sectional widths within the region. This represents the maximum relative height between the two sides. The maximum slope on both sides is determined by statistically analyzing historical data. Regarding the width... Inverse normalization is used, so that the smaller the width (the narrower the terrain), the larger the normalization value. C is the terrain continuity parameter, which is already normalized (value 0-1).

[0111] The terrain access index is calculated using a weighted summation formula:

[0112] ;

[0113] in, , , , The weighting coefficients for each parameter can be determined based on statistical analysis of the correlation between line faults and terrain features during historical severe convective weather, for example... , , , The higher the index, the more easily the terrain forms a funneling effect. Normalizing the index and mapping it to a weight range of 1.0–1.5 yields the terrain funneling effect weights.

[0114] By assigning weights to the influence of trees, buildings, and topographic funneling effects respectively... , , The weighting coefficients are calculated using a weighted summation formula to obtain the environmental weights of the route corridor, where... , , satisfy It was obtained by fitting historical data.

[0115] The spatial weight of the grid is obtained by multiplying the intensity level weight, the terrain correction weight, and the route corridor environment weight.

[0116] The spatial association contribution of each grid along the route segment is obtained by multiplying its grid exposure by the corresponding spatial weight. The spatial association strength of the route segment is then calculated by summing the contributions of all passing grids.

[0117] ;

[0118] in, Spatial correlation strength of line segments, This represents the number of grid cells traversed by the route segment.

[0119] Spatial correlation strength quantifies the comprehensive degree to which a line segment is affected by strong convection in the spatial dimension, and is one of the key innovations of this invention in achieving accurate spatiotemporal matching.

[0120] S400: Calculate the time factor based on the time relationship between the fault time and the strong convection occurrence time; calculate the association credibility based on the spatial association strength and the time factor by weighted summation, and select association results with credibility higher than a preset threshold as association matching data;

[0121] Furthermore, the calculation of the time factor based on the temporal relationship between the fault time and the occurrence of strong convection includes:

[0122] Differentiated time-related windows are set based on the dominant type of severe convective events;

[0123] Calculate the time difference between the time of the line fault and the peak time of the strong convection;

[0124] The time factor is determined based on the inverse ratio of the time difference to the total duration of the time-related window.

[0125] Specifically, differentiated time-related windows are set based on the dominant type of severe convective events. Different types of severe convective weather have significantly different mechanisms of action and temporal characteristics on transmission lines, thus requiring the setting of different time windows.

[0126] Lightning-dominant type: Lightning has a transient impact on the line, with a narrower time window set, such as ±5 minutes from the time of the fault.

[0127] Strong wind-dominated type: Strong winds need to accumulate to cause line galloping and tower deformation due to stress. A wider time window is set, such as 20 minutes before and 10 minutes after the fault.

[0128] Short-duration heavy rainfall-dominated type: Heavy rainfall needs to infiltrate and accumulate, which can cause instability in the tower foundation or geological disasters. A wider time window is set, such as 30 minutes before to 20 minutes after the fault.

[0129] The dominant type of a severe convective event is determined based on the peak values ​​of various meteorological parameters during the event. For example, if the lightning location system records the highest lightning density, it is determined to be a lightning-dominated type; if the automatic weather stations record the most prominent maximum wind speed, it is determined to be a strong wind-dominated type.

[0130] Extracting the fault time of a transmission line This moment is obtained from the fault log data, accurate to the minute or second.

[0131] Extracting the peak moment of a strong convective event This moment marks the point at which various meteorological parameters of a severe convective event (such as radar reflectivity factor, surface wind speed, and precipitation intensity) reach their maximum values.

[0132] Calculate the time difference between the time of the failure and the time of the peak of the strong convection:

[0133] ;

[0134] The time factor is determined based on the inverse ratio of the time difference to the total duration of the time association window. Specifically, the time factor... The calculation formula is:

[0135] ;

[0136] ;

[0137] in, This represents the time difference between the moment of the fault and the peak of the strong convection. This represents the total duration of the time-related window (in minutes). Time factor This indicates a very weak temporal correlation; when Time factor This indicates the strongest temporal correlation.

[0138] For example, for lightning-dominated severe convective events, the time window Minutes, if the time difference between the fault time and the peak time of strong convection... minutes, then time factor .

[0139] Based on time factor Spatial correlation strength and the ratio of peak intensity to threshold of severe convective events. The correlation confidence between strong convective events and line faults is calculated using a weighted summation formula:

[0140] ;

[0141] in: To assess the credibility of the association; The time factor is calculated in step 1; The spatial correlation strength is calculated in step S300; The ratio of peak intensity to threshold for a severe convective event is calculated using the following formula: ,in This represents the peak value of the dominant parameter in a strong convective event (such as the maximum value of the radar reflectivity factor). This is the threshold value for the parameter (e.g., 45 dBZ). , , The weighting coefficients are obtained by regression fitting using labeled data of strong convection-induced faults from historical fault samples.

[0142] Set a credibility threshold Only association results with a confidence level higher than the threshold are retained as association matching data. The confidence threshold setting needs to balance precision and recall, and the optimal value is determined through validation using historical data. Specifically, the association confidence level of historical fault samples is calculated using the above method, and precision-recall curves (PR curves) are plotted under different thresholds. The threshold corresponding to the maximum F1-Score is selected as the [value / value]. For example, by verifying and determining .

[0143] For each line fault record, calculate its correlation confidence with all candidate strong convection events. If the correlation confidence between a strong convection event and the fault is... If so, then the associated record will be retained; if If the correlation is not strong enough, the associated record will be removed, and it will be considered that the strong convective event is not strongly correlated with the fault, and may be a false correlation (such as a line fault that is actually caused by equipment aging, but is mistakenly associated with a strong convective event that occurred at a distance at the same time).

[0144] After credibility screening, the obtained correlation matching data includes the following information: line segment information, including line number, tower number, and segment latitude and longitude range; severe convective event information, including event number, dominant type, affected grid set, peak time, and peak intensity; fault information, including fault time and preliminary fault type record; and correlation parameters, including spatial correlation strength. Time factor Relevance and credibility .

[0145] The correlation and matching data enabled a precise correspondence between transmission lines, severe convective weather conditions, and fault events in the spatiotemporal dimensions, providing a high-quality correlation foundation for subsequent fault attribution labeling.

[0146] By introducing a time factor and quantitative calculation of association credibility, this invention effectively eliminates spurious associations, ensuring the accuracy and reliability of training data.

[0147] S500: Attribution labeling is performed on the transmission line fault information in the associated matching data to obtain a training dataset. The labeling content includes the correspondence between fault type, strong convection cause type and its characteristic parameters.

[0148] Furthermore, the attribution labeling includes:

[0149] Based on the operation and maintenance records, fault waveforms, or on-site investigation results in the transmission line fault data, the fault type is marked.

[0150] Based on the dominant characteristics of strong convection events in the aforementioned correlation matching data, the types of strong convection triggers are labeled;

[0151] The peak values ​​of meteorological parameters corresponding to the type of severe convection are extracted as the characteristic parameters.

[0152] Specifically, the fault type is labeled based on the operation and maintenance records, fault waveforms, or on-site investigation results in the transmission line fault data.

[0153] The system extracts maintenance records such as protection action information, tripping records, and equipment status changes at the time of the fault from the power grid dispatching system to preliminarily determine the fault type. For example, a record showing "line protection tripped, reclosing successful" indicates a transient fault, while a record showing "line protection tripped, reclosing unsuccessful" indicates a permanent fault. Simultaneously, the system analyzes the waveform characteristics of the fault recordings. If the waveform shows a sudden, large instantaneous current surge (milliseconds), it is identified as a lightning strike fault; if the waveform shows a slow current rise followed by a sudden change, it may indicate a conductor breakage fault. Combined with on-site inspection results, the system directly observes the equipment damage. If discharge marks are found on the insulator surface, it indicates a flashover fault; if a tower is found to be tilted or collapsed, it indicates a tower structural fault.

[0154] Based on the above criteria, fault types are clearly classified into line tripping (including instantaneous tripping and permanent tripping), conductor breakage (including single-strand breakage and multi-strand breakage), insulator flashover (including porcelain insulator flashover and composite insulator flashover), tower tilting or collapse, surge arrester damage, and other fault types, ensuring that each fault record has a clear and unique fault type label.

[0155] The key triggers for a fault are determined based on the correlation between the peak values ​​of various meteorological parameters and the fault type during a severe convective event. If the radar reflectivity factor reaches the standard for a severe thunderstorm (e.g., Z≥50dBZ) during a severe convective event, and the lightning location system records ground flash data, with the flash point within the influence range of the faulty line segment (e.g., less than 500 meters from the line), and the fault waveform shows instantaneous high-current impact characteristics, then the trigger type is determined to be lightning. If the maximum wind speed recorded by the automatic ground station during a severe convective event reaches the standard for a strong wind (e.g., V≥25m / s), and the fault type is a tripping caused by a broken conductor, tilted tower, or conductor galloping, and on-site investigation reveals wind-induced damage, then the trigger type is determined to be strong wind. If the 1-hour precipitation during a severe convective event reaches the standard for a rainstorm (e.g., R≥50mm), and the fault type is a tilted tower or foundation settlement, and on-site investigation reveals soil saturation or signs of landslides, then the trigger type is determined to be short-duration heavy rainfall. If meteorological records show hail during a severe convective event, and the fault type is insulator damage or conductor damage, then the cause type is determined to be hail.

[0156] Based on the above criteria, the types of severe convective triggers are classified as lightning (including direct lightning and induced lightning), strong winds (including thunderstorm winds and downbursts), short-duration heavy precipitation (including localized torrential rain and persistent heavy precipitation), hail (including small hail and large hail) and compound triggers (multiple severe convective factors acting together, such as lightning + strong wind, short-duration heavy precipitation + strong wind, etc.).

[0157] For lightning-induced causes, extract the radar reflectivity factor peak. (Unit: dBZ), peak ground flash density (Unit: times / km² / hour), peak lightning intensity (Unit: kA) Minimum cloud top brightness temperature (Unit: °C). For the causes of strong winds, the peak maximum wind speed is extracted. (Unit: m / s), duration of wind speed (Unit: minutes), peak radar reflectivity factor 10-minute average wind speed (Unit: m / s). For the triggers of short-duration heavy rainfall, the peak hourly precipitation is extracted. (Unit: mm) Peak precipitation in 10 minutes (Unit: mm) Duration of precipitation (Unit: minutes), peak radar reflectivity factor Regarding the causes of hail, the diameter of the hailstones was extracted. (Unit: cm) Minimum cloud top brightness temperature Peak radar reflectivity factor Hail duration (Unit: minutes)

[0158] Feature parameters are extracted from meteorological data of the spatiotemporal grid corresponding to the severe convective event in the associated matching data. Specifically, for radar data, the maximum reflectivity factor value of the grid within the time slice corresponding to the fault time is extracted; for automatic weather station data, the maximum wind speed or maximum precipitation recorded by the station within the time window before and after the fault time is extracted; and for lightning location data, the number of lightning strikes and the maximum lightning strike intensity within the grid range within the time window before and after the fault time are statistically analyzed.

[0159] After the above attribution annotation, each associated matching data is transformed into a complete training sample. Each sample contains two parts: input features and labels. The input features include line segment attributes (line number, voltage level, tower type, conductor type, insulator type, facility importance weight), geographical environment features (grid exposure, terrain correction weight, line corridor environment weight), severe convective event features (dominant type, impact zone classification, intensity level weight, peak time), meteorological parameter features (radar reflectivity factor peak, cloud top brightness temperature, maximum wind speed, precipitation, flashover density, etc.), and spatiotemporal association features (spatial association strength, time factor, association confidence). The labels include fault type labels, severe convective cause type labels, and feature parameter values.

[0160] All training samples are aggregated to form a training dataset, which can be directly used to train a transmission line fault rate prediction model under severe convective weather. Each sample in the dataset has a clear "input feature - fault result - cause attribution" correspondence, ensuring that the model can learn the inherent correlation between severe convective weather conditions, geographical environmental factors, and line faults. Through a systematic attribution annotation method, this invention transforms the original multi-source heterogeneous data into high-quality structured training samples, providing reliable data support for the fault prediction model.

[0161] Example 2: To improve the accuracy of its transmission line fault prediction model under severe convective weather, a power grid company needs to acquire high-quality training data. Targeting the high incidence of severe convective weather in the province from June to August 2025, a training data acquisition system for a transmission line fault prediction model according to this invention was adopted. This system includes:

[0162] The data pool construction module is used to construct a multi-source heterogeneous raw data pool, which includes severe convective meteorological data, power transmission line operation and fault data, and geospatial data.

[0163] The data preprocessing module is used to perform spatiotemporal benchmark unification processing on the data in the multi-source heterogeneous raw data pool, and to perform gridding based on a preset spatiotemporal granularity to obtain spatiotemporal gridded data.

[0164] The spatial association matching module is used to divide the transmission line into line segments according to the tower latitude and longitude and map them to the corresponding grid cells of the spatiotemporal gridded data; identify strong convective events and divide the affected areas, and assign intensity level weights to each affected area; calculate the grid exposure of the line segments, and calculate the spatial association intensity by combining the intensity level weights and the terrain correction weights extracted based on geospatial data.

[0165] The time association filtering module is used to calculate the time factor based on the time relationship between the fault time and the strong convection occurrence time; and to calculate the association credibility based on the spatial association strength and the time factor by weighted summation, and to filter the association results with credibility higher than a preset threshold as association matching data.

[0166] The attribution annotation module is used to perform attribution annotation on the transmission line fault information in the associated matching data to obtain a training dataset. The annotation content includes the correspondence between fault type, strong convection cause type and its characteristic parameters.

[0167] Specifically, the system collected relevant data on transmission lines with voltage levels of 500kV and above within the province, including radar station data, satellite cloud image data, ground automatic station observation data and lightning location system data provided by the meteorological department, transmission line operation data and fault records provided by the power grid company, and digital elevation model data and line corridor environmental data provided by the geographic information platform.

[0168] The data preprocessing module unifies the time base of all data to UTC+8 standard time, converts the spatial coordinates to the WGS-84 coordinate system, and divides the data into grids with a 5-minute time granularity and a 1km×1km spatial granularity. The spatial association matching module performs fine-grained segmentation of the transmission lines, further subdividing them by identifying terrain transformation features, environmental abrupt changes, and strong convection-sensitive terrain features within the line corridor. It identifies strong convection events and divides them into core impact areas, significant impact areas, and peripheral impact areas, assigning intensity level weights of 1.0, 0.7, and 0.4, respectively. Combining grid exposure, terrain correction weights, and line corridor environmental weights, the spatial association intensity of each line segment is calculated.

[0169] The time-related filtering module sets differentiated time windows based on the dominant type of severe convection: ±5 minutes for lightning-dominated events, 20 minutes before and 10 minutes after strong winds, and 30 minutes before and 20 minutes after short-duration heavy rainfall. By calculating the time factor and correlation confidence, and setting a confidence threshold of 0.6, it successfully filtered out highly reliable correlation matching data, effectively eliminating false correlations between faults caused by non-severe convective factors such as aging line equipment and severe convective events, as well as weak correlation data with excessively large time differences.

[0170] The attribution annotation module systematically annotates the associated matching data, clearly distinguishing between lightning-induced faults (main fault types being instantaneous tripping and insulator flashover), strong wind-induced faults (main fault types being conductor breakage and tower tilting), short-duration heavy rainfall-induced faults (main fault type being tower foundation settlement), and faults caused by combined factors. It also annotates corresponding feature parameters for each sample, such as the peak radar reflectivity factor and peak flash intensity for lightning-induced samples, and the peak maximum wind speed and wind speed duration for strong wind-induced samples.

[0171] This invention achieves refined spatial matching between railway line segments and severe convective events through multi-dimensional weighted comprehensive calculations. It achieves accurate spatiotemporal correlation through differentiated time window settings and a correlation credibility filtering mechanism. Furthermore, it establishes a clear correspondence between fault types and severe convective triggers through systematic attribution labeling. Compared to traditional methods that rely solely on simple time window matching without considering spatial correlation strength and credibility filtering, this invention significantly improves the quality of the training dataset, accurately reflecting the true correlation between severe convective weather conditions, geographical environmental factors, and railway line faults. This provides reliable data support for constructing high-precision fault prediction models.

[0172] Example 3: In the third embodiment of the present invention, based on the same inventive concept, the present invention proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps of the training data acquisition method for a transmission line fault prediction model of the above embodiment.

[0173] Example 4: The fourth embodiment of the present invention, based on the same inventive concept, proposes a computer device comprising: a processor and a memory; the processor and the memory communicate with each other; the memory is used to store instructions; the processor is used to execute the instructions in the memory to execute the training data acquisition method for a transmission line fault prediction model of the above embodiment.

[0174] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0175] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for acquiring training data for a transmission line fault prediction model, characterized in that, include: S100: Construct a multi-source heterogeneous raw data pool, which includes severe convective weather data, power transmission line operation and fault data, and geospatial data. S200: Perform spatiotemporal benchmark unification processing on the data in the multi-source heterogeneous raw data pool, and perform gridding based on the preset spatiotemporal granularity to obtain spatiotemporal gridded data; S300: Divide the transmission line into line segments according to the latitude and longitude of the towers and map them to the corresponding grid cells of the spatiotemporal gridded data; identify strong convective events and divide the affected areas, and assign intensity level weights to each affected area; Calculate the grid exposure of the line segment, and combine the intensity level weights and the terrain correction weights extracted based on geospatial data to calculate the spatial association strength; S400: Calculate the time factor based on the time relationship between the time of the fault and the time of the strong convection; The association credibility is calculated by weighted summation based on the spatial association strength and time factor, and association results with credibility higher than a preset threshold are selected as association matching data. S500: Attribution labeling is performed on the transmission line fault information in the associated matching data to obtain a training dataset. The attribution labeling includes the fault type, strong convection cause type and the correspondence of their characteristic parameters.

2. The method for acquiring training data for a transmission line fault prediction model according to claim 1, characterized in that, The identification of severe convective events and the delineation of affected areas include: Multi-parameter collaborative thresholds were set based on radar reflectivity factor, satellite cloud top brightness temperature, and ground automatic station observation data. The density clustering algorithm is used to aggregate adjacent grids that meet the multi-parameter collaborative threshold to identify strong convection events; Based on the intensity distribution of the severe convective event decreasing outward from the core area, the core influence area, significant influence area, and peripheral influence area are divided from the inside out, and each influence area is assigned a corresponding intensity level weight.

3. The method for acquiring training data for a transmission line fault prediction model according to claim 1, characterized in that, The calculation of the grid exposure of the line segment includes: The transmission line is divided into basic segments according to the tower coordinates, and then further subdivided according to the key geographical features within the line corridor to generate several fine segments. The fine line segment is spatially cut with the spatiotemporal gridded data to obtain a sub-line segment that falls completely into a single grid cell; The weighted length is calculated based on the length of each sub-segment and the attribute parameters of the line to which it belongs; The grid exposure is obtained by summing the weighted lengths of all sub-segments falling within the same grid and dividing by the grid edge length.

4. The method for acquiring training data for a transmission line fault prediction model according to claim 1, characterized in that, The terrain correction weights extracted from geospatial data are obtained in the following way: Extract the slope, aspect, and relative elevation parameters of the grid from the digital elevation model data; The weighting coefficients for the slope, aspect, and relative elevation parameters are fitted based on historical fault data. The slope, aspect, and relative altitude parameters are weighted and summed according to their corresponding weight coefficients, and then normalized to obtain the terrain correction weights.

5. The method for acquiring training data for a transmission line fault prediction model according to claim 1, characterized in that, The computational spatial correlation strength includes: The environmental weight of the route corridor is obtained by weighted summation of the threat level of trees, the threat level of buildings and the terrain passage index within the route corridor. Multiply the intensity level weight, the terrain correction weight, and the route corridor environment weight to obtain the spatial comprehensive weight of the grid; The spatial correlation strength of a line segment is obtained by multiplying the grid exposure of each grid along the route with the spatial comprehensive weight of the corresponding grid and summing the results.

6. The method for acquiring training data for a transmission line fault prediction model according to claim 1, characterized in that, The calculation of the time factor based on the temporal relationship between the time of the fault and the time of the strong convection includes: Differentiated time-related windows are set based on the dominant type of severe convective events; Calculate the time difference between the time of the line fault and the peak time of the strong convection; The time factor is determined based on the inverse ratio of the time difference to the total duration of the time-related window.

7. The method for acquiring training data for a transmission line fault prediction model according to claim 1, characterized in that, The attribution annotations include: Based on the operation and maintenance records, fault waveforms, or on-site investigation results in the transmission line fault data, the fault type is marked. Based on the dominant characteristics of strong convection events in the aforementioned correlation matching data, the types of strong convection triggers are labeled; The peak values ​​of meteorological parameters corresponding to the type of severe convection are extracted as the characteristic parameters.

8. A training data acquisition system for a transmission line fault prediction model, characterized in that, include: The data pool construction module is used to construct a multi-source heterogeneous raw data pool, which includes severe convective meteorological data, power transmission line operation and fault data, and geospatial data. The data preprocessing module is used to perform spatiotemporal benchmark unification processing on the data in the multi-source heterogeneous raw data pool, and to perform gridding based on a preset spatiotemporal granularity to obtain spatiotemporal gridded data. The spatial association matching module is used to divide the transmission line into line segments according to the tower latitude and longitude and map them to the corresponding grid cells of the spatiotemporal gridded data; identify strong convective events and divide the affected areas, and assign intensity level weights to each affected area; Calculate the grid exposure of the line segment, and combine the intensity level weights and the terrain correction weights extracted based on geospatial data to calculate the spatial association strength; The time-related filtering module is used to calculate the time factor based on the time relationship between the time of the fault and the time of the strong convection. The association credibility is calculated by weighted summation based on the spatial association strength and time factor, and association results with credibility higher than a preset threshold are selected as association matching data. The attribution annotation module is used to perform attribution annotation on the transmission line fault information in the associated matching data to obtain a training dataset. The attribution annotation content includes the correspondence between fault type, strong convection cause type and its characteristic parameters.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements a method for acquiring training data for a transmission line fault prediction model as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, which, when executed by a processor, implements a method for acquiring training data for a transmission line fault prediction model as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Regional icing thickness distribution estimation method based on multi-source data

    CN115062860A

  • Power distribution network line fault prediction method and device

    CN115270965A