A power grid weather disaster judgment weight updating method and system
By introducing a cause-of-fault matching mechanism that combines time window constraints and device key matching into the power grid meteorological disaster fault assessment system, the closed-loop control problem of cause-of-fault data utilization and weight update is solved, enabling continuous self-learning and controllable release of weights, and improving the system's assessment accuracy and stability.
Patent Information
- Application Number
- CN202610805938.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-25
AI Technical Summary
Existing power grid meteorological disaster fault assessment systems face difficulties in utilizing root cause data and updating parameters, making it difficult to achieve continuous adaptive iteration. This leads to a decline in assessment accuracy and stability, especially when low-frequency disaster samples are scarce and categories are imbalanced, resulting in model bias and a lack of highly reliable matching mechanisms and closed-loop control for weight updates.
A highly reliable cause matching mechanism is established by using time window constraints and device key matching. A candidate cause set is constructed and sorted and filtered. Highly reliable labels are generated by combining threshold filtering. Weight self-learning and versioned release are realized. A closed-loop mechanism for online access control, gray-scale and rollback is constructed to ensure that the update process is traceable, evaluable and rollbackable.
It achieves continuous self-learning updates of weights without changing the interpretable judgment framework, improving the accuracy and stability of judgment, possessing regional and temporal adaptation capabilities, reducing training noise and improving interpretability and traceability.
Smart Images

Figure CN122634071A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power grid meteorological disaster assessment technology, and more specifically, relates to a power grid meteorological disaster assessment weight update method and system. Background Technology
[0002] Overhead transmission lines and distribution network equipment operate outdoors for extended periods, making them susceptible to meteorological factors such as lightning, strong convection, strong winds, low temperatures, rain and snow, icing, and the resulting galloping. This can lead to faults such as tripping, flashovers, line breaks, tower collapses, and insulation damage. These faults are characterized by their suddenness, wide impact, and short response window, potentially resulting in decreased power supply reliability and increased maintenance costs. Rapidly identifying the most likely causes of a fault within a short period after its occurrence is crucial for selecting repair routes, dispatching personnel and resources, conducting subsequent hazard investigations, and revising early warning strategies. Compared to relying solely on protection actions and alarm information, the fault mechanisms in meteorological disaster scenarios are often highly correlated with environmental conditions in the temporal and spatial context of the fault; therefore, it is essential to incorporate meteorological environmental information into the analysis.
[0003] Current engineering practices typically extract judgment factors from multi-source data, including surface meteorological observations and numerical forecasts (temperature, humidity, precipitation, wind speed and direction, air pressure, etc.), remote sensing products such as weather radar and satellites, lightning location data (number of lightning strikes, lightning current amplitude, lightning strike distance, etc.), online monitoring or forecasting information on icing and galloping, as well as the topography, altitude, and equipment protection configuration of the line corridor. After temporal and spatial alignment and standardization, the above data can be transformed into interpretable feature indicators to output disaster categories or risk probabilities, supporting operation and maintenance decisions.
[0004] However, such assessments still face common challenges in practice: First, disasters such as icing and glacial dancing are infrequent but highly destructive, and the historical sample size of the root causes is limited and the categories are unbalanced; Second, multiple data sources have time delays, inconsistent granularity, and missing data, which can easily lead to difficulties in aligning "root cause labels - factor features"; Third, regional differences and seasonal changes can cause the distribution and contribution of factors to drift over time, making it difficult for fixed weights or fixed model parameters to maintain stable results in the long term.
[0005] Existing meteorological disaster fault assessment systems for power grids typically employ interpretable models based on "factor scoring + weighted averaging." However, the weights are often preset by expert experience or calibrated offline in a one-time manner, making it difficult to continuously adapt to regional differences and time-varying characteristics, leading to decreased accuracy and stability. Specifically: ① Significant differences exist in terrain, line structure, operation and maintenance strategies, and meteorological characteristics across different regions, making it difficult to transfer fixed weights; ② Seasonal changes, equipment status changes, and data source updates cause factor contribution drift, preventing fixed weights from adapting; ③ Although the operation and maintenance system contains records of historically verified true fault causes (root causes), the lack of automatic association, training, and release loops prevents continuous calibration of assessment weights as new root cause samples accumulate; ④ Weight updates lack version management, gray-scale release, and rollback control, resulting in high update risks in the production system and low update frequency.
[0006] Existing technologies still face engineering challenges in "utilizing root cause data" and "updating parameters," making continuous adaptive iteration difficult. While existing solutions can output disaster categories or risk probabilities, they generally suffer from the following shortcomings in practical engineering applications: (1) The lack of a reliable alignment mechanism for real cause labels makes it difficult to automatically accumulate training / calibration data. In power grid business systems, the causes of faults are usually confirmed and recorded by maintenance personnel after the fact, based on the on-site situation, forming a historical real cause database (true cause database). However, existing judgment systems often do not establish a stable "fault sample - true cause record" association mechanism. Especially in cases where there are inconsistencies in coding across multiple systems, data delays, differences in location granularity, and multiple events occurring concurrently within the same time window, simple time matching or name matching can easily lead to mislabeling or conflicting labels, thus making it impossible to automatically form highly reliable training samples.
[0007] (2) Weight or model parameter updates rely on manual adjustment or offline one-time calibration, lacking versioned release and rollback, resulting in high update risk and low frequency. Existing weighted or fuzzy evaluation schemes usually deploy weights as fixed configurations, and weight updates often rely on manual experience or a small amount of offline analysis, lacking an engineering mechanism that includes "evaluation access control - version management - gray-scale activation - online monitoring - anomaly rollback". The above defects make it difficult to carry out weight updates routinely, thus failing to effectively cope with regional differences and time-varying drift.
[0008] (3) The scarcity of low-frequency disaster samples and the imbalance of categories are prominent problems, causing the model to be biased towards high-frequency disasters in the long term. In the assessment of multiple disasters, lightning events are usually more frequent, while disasters such as icing and dancing occur relatively less frequently and are more affected by regional factors. If there is a lack of training data management strategies for the scarcity of samples and the imbalance of categories (such as conflict sample removal, minority class protection and access thresholds, etc.), the model or weight update results are likely to be biased towards high-frequency disasters, affecting the ability to assess minority classes.
[0009] Therefore, the technical problems to be solved by this invention are: (1) how to use historical real fault cause records to automatically and reliably match them with fault samples to be evaluated and construct training data. (2) to achieve weight self-learning while maintaining an interpretable weighted structure, and to release the learned new weights to the online evaluation in a closed loop in a versioned manner. (3) the system has engineering control capabilities for evaluation, grayscale and rollback.
[0010] Therefore, a technical solution is needed that can address the "credibility of cause alignment" and the "controllability of closed-loop weight update" without disrupting the existing interpretable judgment framework, thereby enabling continuous self-learning updates of model weights and ensuring that the update process is traceable, evaluable, and rollbackable. Summary of the Invention
[0011] To address the shortcomings of existing technologies, this invention provides a root cause matching and weight closed-loop update mechanism for meteorological disaster fault assessment. A highly reliable root cause matching mechanism is established: based on time window constraints and device key matching, spatial / attribute constraints are introduced to calculate matching confidence, ranking candidate root causes, and generating unique, highly reliable labels through threshold screening and conflict resolution. This automatically accumulates a sample set suitable for supervised training, reducing label noise and improving sample utilization. Weight self-learning and versioned release are achieved: while maintaining the interpretable structure of "factor scoring + weighted average," the weights are updated through self-learning using the matched root cause labels, and the new weights are stored in a "version + metadata" format. The system ensures reproducible training and traceable modifications; it constructs a closed-loop mechanism for deployment access control, gray-scale deployment, monitoring, and rollback: setting access control strategies for new weights based on sample size, offline indicators, and stability thresholds; supporting gray-scale implementation by region / line / proportion; continuously monitoring the effect based on newly added root cause data after deployment; and automatically rolling back to the previous stable version when degradation conditions are triggered, thus making weight updates controllable, auditable, and sustainable in the production environment; and providing training and deployment protection for low-frequency disasters: through minority class sample size thresholds, training set balancing strategies, and disaster-level access control conditions, it avoids excessive bias of weight learning towards high-frequency disasters, ensuring that the judgment capability of low-frequency, high-impact disasters such as icing and dancing can be gradually improved through continuous iteration. Through the achievement of the above objectives, this invention can form a closed-loop self-learning process of "root cause alignment—training—evaluation—deployment—monitoring—rollback—retraining" in engineering, enabling the judgment model to have regional and temporal adaptive capabilities while maintaining interpretability and traceability.
[0012] The present invention adopts the following technical solution.
[0013] The first aspect of the present invention provides a method for updating the weights for assessing meteorological disasters affecting power grids, comprising the following steps: Obtain the set of fault samples to be analyzed and the set of historical root causes from power grid business data; Candidate causes are selected from the historical cause set to form a candidate cause set; the matching confidence of each candidate cause in the candidate cause set with each fault sample in the fault sample set is calculated; based on the descending order of the matching confidence results, combined with the preset confidence threshold and the preset conflict difference threshold, the labeled sample set is determined. The labeled sample set is divided into a training set and a validation set, and a factor weight matrix is constructed according to network type and disaster type. Based on the factor scores of fault samples in the training set and the corresponding weight vectors of network type and disaster type in the factor weight matrix, the judgment score of each disaster type is calculated to predict the probability. The error between the predicted probability and the actual disaster type label is calculated to iteratively update the factor weight matrix. The judgment performance of the current factor weight matrix is evaluated using the validation set. When the judgment performance reaches the convergence condition, candidate weights are determined according to the factor weight matrix. Candidate weights are evaluated using access control. Candidate weights that pass the access control evaluation are then given a gray-scale effective range determined by a preset strategy. Within the gray-scale effective range, the status of the candidate weights that pass the access control evaluation is assessed based on online monitoring indicators. If the evaluation status is deteriorated, the candidate weights are automatically rolled back to the previous stable version. If the evaluation status is stable, all candidate weights are released, and the current candidate weights are updated to the stable version.
[0014] Preferably, the step of determining the labeled sample set includes: Based on each fault sample, candidate causes that meet the time window matching and device key matching are selected from the historical cause set to form a candidate cause set. For each fault sample in the fault sample set, calculate the matching confidence of each candidate cause in the candidate cause set with the fault sample, and sort the matching confidence in descending order. If the highest matching confidence is less than the preset confidence threshold, the fault sample is classified as an unlabeled sample; if the difference between the highest matching confidence and the second highest matching confidence is less than the conflict difference threshold and cannot be resolved, the fault sample is classified as a conflict sample; otherwise, the real disaster type label carried by the candidate true cause corresponding to the highest matching confidence is assigned to the fault sample, forming a labeled sample and included in the labeled sample set.
[0015] Preferably, the device key includes device type, device identifier, device topology affiliation, and device line location; If the device key of the fault sample is the same as the device type and device identifier of the device key of the historical cause, it is determined that the device key match is satisfied. If the device key of the fault sample is different from the device key of the historical cause, but the device key of the fault sample and the device key of the historical cause can be mapped to the same power grid object through the predefined topology relationship table, then it is determined that the device key matching is satisfied. If any of the following conditions are met: the device topology of the device key of the fault sample is the same as the device identifier of the device key of the historical cause, the device topology of the device key of the historical cause is the same as the device identifier of the device key of the fault sample, or the device topology of the device key of the fault sample is the same as the device topology of the device key of the historical cause, the device key matching is deemed to be satisfied. All other cases are judged as not meeting the device key matching requirement.
[0016] Preferably, the matching confidence score between each candidate cause in the candidate cause set and the fault sample is calculated, expressed by the following formula:
[0017] In the formula, This represents the matching confidence score between the i-th fault sample and the m-th candidate cause. , , and Let represent the weights of temporal similarity, spatial similarity, device attribute similarity, and record quality factor, respectively, and all of them be no less than 0, satisfying the condition that... , express and Time similarity, This represents the failure occurrence time of the m-th candidate cause. This represents the time of failure occurrence for the i-th fault sample. express and Spatial similarity, This indicates the fault location of the m-th candidate cause. This indicates the fault location of the i-th fault sample. express and The similarity of device attributes, This represents the device key in the m-th candidate true factor. The device key representing the i-th fault sample. The quality factor representing the candidate true cause Indicates the m-th candidate true cause The quality factor.
[0018] Preferably, The time similarity is expressed by the following formula:
[0019] In the formula, This represents the allowed matching time window for the d-th disaster type.
[0020] Preferably, Spatial similarity is expressed by the following formula:
[0021] In the formula, This represents the distance between the m-th candidate root cause and the i-th fault sample. This represents the spatial matching radius of the d-th disaster type.
[0022] Preferably, determining the candidate weights includes: The labeled sample set is divided into a supervised training set and a validation set, and a factor weight matrix is constructed according to network type and disaster type. The factor scores based on the samples in the supervised training set are weighted with the weight vectors of the corresponding network type and disaster type in the factor weight matrix to calculate the judgment score of each disaster type, and the prediction probability is output through the Softmax function. The error between the predicted probability and the true label is calculated using the cross-entropy loss function with category weights, and the factor weight matrix is iteratively updated using the gradient descent method. After each iteration, a non-negative normalization constraint is applied to the weight vectors in the updated factor weight matrix to normalize the sum of the weight components. The performance index of the current factor weight matrix is evaluated using the validation set. If the performance index does not meet the preset convergence condition, the updated factor weight matrix is used to continue training. If the preset convergence condition is met, the factor weight matrix is output as a candidate weight.
[0023] Preferably, the assessment score for each type of disaster is calculated and expressed by the following formula:
[0024] in, This represents the judgment score of the t-th sample in the supervised training set belonging to the d-th disaster type. Indicate network type The factor weight vector corresponding to the d-th disaster type, This represents the network class of the t-th sample in the supervised training set. This represents the factor score of the t-th sample in the supervised training set. Indicate network type The bias term corresponding to the d-th disaster type.
[0025] Preferably, the cross-entropy loss function is expressed by the following formula:
[0026] in, Indicates training loss, This represents the set of weights to be trained. This represents the true disaster type of the t-th sample in the supervised training set. The corresponding category weights, This indicates that the t-th sample in the supervised training set corresponds to its true disaster type. The predicted probability, This represents the stability constraint coefficient. Indicate network type The factor weight corresponding to the d-th disaster type Indicates the network type of the previous iteration. The factor weight corresponding to the d-th disaster type To sum all samples in the supervised training set, To sum the results for all networks, To sum up all types of disasters, It is the square of the L2 norm.
[0027] A second aspect of the present invention provides a power grid meteorological disaster-causing assessment weight update system, which, according to the power grid meteorological disaster-causing assessment weight update method described in the first aspect, includes: The data acquisition module is used to obtain the set of fault samples to be analyzed and the set of historical root causes from power grid business data; The labeling module is used to filter candidate causes from the historical cause set to form a candidate cause set; calculate the matching confidence of each candidate cause in the candidate cause set with each fault sample in the fault sample set; based on the descending order of the matching confidence results, combined with the preset confidence threshold and the preset conflict difference threshold, determine the labeled sample set. The weight update module is used to divide the labeled sample set into a training set and a validation set, and construct a factor weight matrix according to network type and disaster type. Based on the factor scores of fault samples in the training set and the weight vectors of the corresponding network type and disaster type in the factor weight matrix, the judgment score of each disaster type is calculated to predict the probability. The error between the predicted probability and the actual disaster type label is calculated to iteratively update the factor weight matrix. The judgment performance of the current factor weight matrix is evaluated using the validation set. When the judgment performance reaches the convergence condition, candidate weights are determined according to the factor weight matrix. The stable weight acquisition module is used to perform access control evaluation on candidate weights. The candidate weights that pass the access control evaluation are given a gray-scale effective range through a preset strategy. Within the gray-scale effective range, the status of the candidate weights that pass the access control evaluation is evaluated based on online monitoring indicators. If the evaluation status is deteriorated, it will automatically roll back to the previous stable version of the candidate weights. If the evaluation status is stable, the candidate weights will be released in full, and the current candidate weights will be updated to the stable version of the candidate weights.
[0028] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded onto the processor, it implements a power grid meteorological disaster assessment weight update method according to the first aspect.
[0029] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a power grid meteorological disaster assessment weight update method according to the first aspect.
[0030] Compared with the prior art, the beneficial effects of the present invention include at least the following: This invention revolves around a closed-loop mechanism of "root cause matching—weight self-learning—versioned release—online monitoring—automatic rollback," enabling weights to be continuously calibrated and controllably deployed based on historical real causes without altering the interpretable framework of "factor scoring + weighting." Because this invention improves the credibility of root cause label alignment from the source and engineers weight updates into an assessable, traceable, and rollbackable process, it enhances the credibility of root cause label alignment, reduces training noise from the source, and improves the stability and reproducibility of weight learning. It forms a continuously accumulating pool of supervisory samples and closed-loop iteration capabilities, allowing the model to evolve by continuously calibrating the judgment weights as new root cause samples accumulate. It achieves versioned management and controllable release of weights, reducing deployment risks and improving traceability and auditability. Using the "factor scoring + weighting" structure as the basis for online judgment preserves the interpretability of factor weights, facilitating operational review and auditing. In summary, this invention achieves a technical path of "reliable samples → stable training → controllable deployment → continuous evolution" through systematic improvements to the cause alignment mechanism and weight update closed-loop mechanism. This not only improves the accuracy and stability of the analysis, but also meets the requirements of the production environment for risk control, traceability and explainability. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the power grid meteorological disaster assessment weight update method provided in accordance with the embodiments of the present invention; Figure 2 This is a schematic diagram of the cause matching and sample construction process provided in accordance with the embodiments of the present invention; Figure 3 This is a schematic diagram of the weight self-learning training process provided in accordance with an embodiment of the present invention; Figure 4 This is a schematic diagram of weighted versioning closed-loop release and rollback provided in accordance with the embodiments of the present invention; Figure 5 This is a schematic diagram of the overall system architecture provided in accordance with an embodiment of the present invention. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this invention. The described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the spirit of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of this invention.
[0033] like Figure 1 As shown, Embodiment 1 of the present invention provides a method for updating the weights for assessing meteorological disasters affecting power grids, comprising the following steps: Step 1: Obtain the set of fault samples to be analyzed and the set of historical root causes from the power grid business data.
[0034] Set of fault samples to be analyzed In, each sample Fault occurrence time including fault samples Device key of faulty sample Fault location of the fault sample Factor scores .in, The original physical quantities are mapped to scores using existing scoring rules to maintain interpretability. In Let represent the j-th factor score of the i-th fault sample, and n represent the number of factors.
[0035] The factor set includes the original physical quantities of each type of disaster, including lightning disaster, icing disaster, and galloping disaster. Each original physical quantity is a factor within the factor set.
[0036] For example, the raw physical quantities for lightning disasters may include lightning current amplitude, lightning distance, lightning density, and precipitation intensity; the raw physical quantities for icing disasters may include temperature, humidity, wind speed, icing thickness, and altitude; and the raw physical quantities for galloping disasters may include wind speed, wind direction angle, icing condition, and span. By statistically analyzing historical fault data and collecting deviation standardization parameters, a linear transformation is performed on the raw physical quantities, mapping different raw physical quantity scores to... between.
[0037] The factor score of this invention is obtained by standardizing or grading the original physical quantity based on deviation.
[0038]
[0039] In the formula, This represents the j-th factor score of the i-th fault sample. This represents the actual value of the j-th factor in the i-th fault sample. and Let represent the lower and upper bounds of the standardization for the j-th factor, respectively. Positively correlated factors are raw physical quantities whose larger values indicate higher disaster risk, such as, but not limited to, lightning current amplitude, wind speed, icing thickness, and precipitation intensity. When a factor is positively correlated... The negative correlation factor is the original physical quantity whose smaller value indicates a higher disaster risk, such as, but not limited to, icing temperature. When the factor is a negative correlation factor, .
[0040] Through the above processing, raw data of different dimensions such as lightning current amplitude, lightning strike distance, wind speed, temperature, humidity, icing thickness, and altitude can be uniformly converted into interpretable factor scores.
[0041] Data sources for power grid operations include: fault master record database, factor score detail database, and historical true cause database.
[0042] The fault master record database includes the fault occurrence time of the fault samples. Device key of faulty sample Fault location of the fault sample Fault samples, factor score detail library including factor scores Detailed scoring information is obtained by extracting a set of historical true causes from a database of historical true causes. , This represents the k-th historical truth cause. Each historical truth cause contains at least one true disaster label for the historical truth cause. The time of occurrence of the historical cause of the failure The device key of historical truth The location of the fault due to historical reasons .
[0043] Step 2: Select candidate root causes from the historical root cause set to form a candidate root cause set; calculate the matching confidence score between each candidate root cause in the candidate root cause set and each fault sample in the fault sample set; based on the descending order of the matching confidence scores, and combined with a preset confidence threshold and a preset conflict difference threshold, determine the labeled sample set, such as... Figure 2 As shown.
[0044] In a preferred but non-limiting embodiment of the present invention, step 2 includes: Step 2.1: Using each fault sample as a benchmark, candidate root causes that satisfy both time window matching and device key matching are selected from the historical root cause set to form a candidate root cause set, expressed by the following formula:
[0045] In the formula, Let i represent the set of candidate root causes for the i-th fault sample. This represents the k-th historical truth. , Represents the set of historical truths , This represents the time when the fault occurred in the k-th historical cause. This represents the time of failure occurrence for the i-th fault sample. This indicates the allowed matching time window for the d-th disaster type. The time window is configured according to the disaster type. For example, lightning disasters can be set to 15 minutes to 1 hour before and after the fault, icing disasters can be set to 6 hours to 24 hours before and after the fault, and galloping disasters can be set to 1 hour to 6 hours before and after the fault. This represents the device key in the k-th historical truth. The device key representing the i-th fault sample. This represents a device key matching function used to determine the device key in the k-th historical cause. Does the device key of the i-th fault sample satisfy device key matching? If the device key in the k-th historical cause... If the device key of the i-th fault sample satisfies any of the following matching relationships: device matching, mapping matching, or hierarchical matching, the value is 1; otherwise, the value is 0.
[0046] The device key includes the device type, device identifier, device topology affiliation, and device line location.
[0047] Equipment types include one of the following: poles and towers, line sections, switches, distribution transformers, and substations; Equipment identification includes one of the following: equipment ID, pole ID, switch ID, transformer ID, or line segment ID; The topology affiliation of a device identifies the superior affiliation of the device in the power grid topology, including one of the following: the ID of the line to which it belongs, the ID of the feeder to which it belongs, the ID of the distribution area to which it belongs, or the ID of the substation to which it belongs. The equipment's line location identifier indicates the specific location of the equipment within the line, including one of the following: pole number, pole number, or station number.
[0048] Device key matching includes exact matching, mapping matching, and hierarchical matching. A device key match is determined to be satisfied if any one of these three types of matching is met.
[0049] If the device key of the fault sample is the same as the device type and device identifier of the device key of the historical cause, it is determined to be an exact match, and the device key match is satisfied.
[0050] If the device key of the fault sample is different from the device key of the historical cause, but the device key of the fault sample and the device key of the historical cause can be mapped to the same power grid object (such as unified asset number or GIS element ID) through a predefined asset ledger mapping table or topology relationship table, then it is determined that the mapping match is satisfied and the device key match is satisfied.
[0051] If the device topology of the device key of the fault sample is the same as the device identifier of the device key of the historical cause (e.g., the tower belongs to a certain line), the device topology of the device key of the historical cause is the same as the device identifier of the device key of the fault sample (e.g., a certain line contains the tower), or the device topology of the device key of the fault sample is the same as the device topology of the device key of the historical cause (e.g., they belong to the same line or the same substation), it is determined that the hierarchical matching is satisfied and the device key matching is satisfied.
[0052] If none of the above conditions are met, it is determined that the device key matching is not satisfied.
[0053] Equipment key matching includes equipment matching, mapping matching, and hierarchical matching. Equipment matching refers to two items corresponding to the same equipment, pole, switch, transformer, or line segment. Mapping matching refers to two items that, although from different business systems and with different codes, can be mapped to the same power grid object through asset ledgers, equipment mapping tables, or topology relationship tables. Hierarchical matching refers to two items that, although not the same object, have hierarchical topology or management relationships such as pole and line, switch and feeder, transformer and distribution area, or line and substation. Through the above equipment key matching, the problem of incorrect labeling of faults of different lines and different equipment within the same time window due to relying solely on time proximity can be avoided.
[0054] When a pending fault record appears, the fault sample of the pending fault record is read from the fault master record library, and the factor score of the fault sample is pulled from the factor score detail library using the fault number of the fault sample as the key; at the same time, the candidate true cause set is pulled from the historical true cause library using the time window and equipment key conditions, forming a unified input structure containing fault samples, factor scores and candidate true cause sets.
[0055] Step 2.2: For each fault sample in the fault sample set, calculate the matching confidence of each candidate cause in the candidate cause set with the fault sample.
[0056] When the candidate root cause set is not empty, constraints of time, space, device attributes, and record quality factors are introduced, and a matching confidence score is calculated for each candidate root cause. It can be expressed by the following formula:
[0057] In the formula, This represents the matching confidence between the i-th fault sample and the m-th candidate cause, and its value ranges from 0 to 1 in the closed interval. , , , These represent the weights of temporal similarity, spatial similarity, device attribute similarity, and record quality factor, respectively, and all are not less than 0. Preferably, they satisfy the following conditions: ; The time similarity is expressed by the following formula:
[0058] In the formula, express and Time similarity is used to measure and The degree of closeness, This represents the failure occurrence time of the m-th candidate cause. This represents the time of failure occurrence for the i-th fault sample. This indicates the allowed matching time window for the d-th disaster type. The time window is configured according to the disaster type. For example, lightning disasters can be set to 15 minutes to 1 hour before and after the fault, icing disasters can be set to 6 hours to 24 hours before and after the fault, and galloping disasters can be set to 1 hour to 6 hours before and after the fault.
[0059] express and Spatial similarity is expressed by the following formula:
[0060] In the formula, express and Spatial similarity is used to measure and The degree of closeness, This indicates the fault location of the m-th candidate cause. This indicates the fault location of the i-th fault sample. This represents the distance between the m-th candidate root cause and the i-th fault sample. The distance is preferably a latitude-longitude distance, a line corridor distance, or a distance along the power grid topology path. This represents the spatial matching radius of the d-th disaster type.
[0061] express and Device attribute similarity is used to measure how close two devices are in the power grid topology or management level. This represents the device key in the m-th candidate true factor. The device key represents the i-th fault sample.
[0062] like and The equipment types and equipment identifiers are all the same, or, and The devices may have different identifiers, but they can be linked through a predefined topology table. and Mapped to the same power grid object, retrieve ; like and If either the equipment type or the equipment identifier is different, further judgment is needed. Device topology attribution and The device identifiers are the same, or, Device topology attribution and The device identifiers are the same, take ; like Device topology attribution and Equipment identification, Device topology attribution and The device identifications are all different. Further investigation reveals that... Device topology attribution and The devices belong to the same topology, take ; like Device topology attribution and The device topologies are also different. Further investigation reveals that... Equipment line location and The equipment lines are in the same location, so take ; If none of the above conditions are met, take .
[0063] Indicates the m-th candidate true cause The quality factor is used to represent the m-th candidate true cause. The credibility of data can be weighted by factors such as whether it was manually verified, the field completeness rate, and the data source level, and can be expressed by the following formula:
[0064] in, , , For the weighting coefficients, the preferred one is to satisfy... , This indicates a manually confirmed indicator; if the m-th candidate root cause... After manual confirmation, The value is 1, otherwise The value is 0. The field completeness rate metric is expressed by the following formula:
[0065] In the formula, This represents the set of required fields. Indicates to Each field ,like exist It exists in ,otherwise, , This indicates the number of fields in the set of required fields.
[0066] Indicates the level of data source. It originated from the determination of responsibility during operation and maintenance, and was judged to be of the highest level. Takes the value 1. Based on professional monitoring, it was determined to be of medium to high level. The value is 0.7. It was determined to be of medium level during the planned maintenance. The value is 0.5. When the data was imported manually, it was classified as a lower level. The value is 0.3.
[0067] Step 2.3: Sort the matches in descending order of confidence; if the highest match confidence is less than the preset confidence threshold... If the faulty sample is not identified, it will be classified as an unlabeled sample; if the highest matching confidence score is 1, the faulty sample will be classified as an unlabeled sample. The second highest matching confidence The difference is less than the conflict difference threshold. If the fault cannot be resolved, the fault sample is classified as a conflict sample; otherwise, the real disaster type label carried by the candidate true cause corresponding to the highest matching confidence is assigned to the fault sample to form a labeled sample and included in the labeled sample set.
[0068] Conflict resolution is then performed in the following order: most recent time, finer equipment granularity, and higher record quality. Fineer equipment granularity prioritizes tower level over line level, switch level over feeder level, and transformer level over distribution area level. If a unique true disaster type label still cannot be determined after following these rules, the fault sample is marked as a conflict sample, temporarily excluded from supervised training, and can be added to the manual review queue.
[0069] Labeled samples are added to the supervised training queue; conflicting samples can be added to the review queue or temporarily excluded from training; unlabeled samples can be used for statistical monitoring or semi-supervised extension.
[0070] After cause matching is completed, the fault samples are divided into labeled samples, unlabeled samples, and conflict samples, with only labeled samples used as supervised training samples. Enter the training queue. For labeled samples, data cleaning is performed first, including handling missing values, outliers, and duplicate records. When the proportion of missing factors in a single sample exceeds the preset missing threshold, the sample is removed. When the missing proportion does not exceed the preset missing threshold, it can be filled using the median or mean of historical samples from the same region, type of disaster, and season, and a missing value is recorded. Outliers that exceed the physically reasonable range or historical statistical quantile range can be truncated. For duplicate records, deduplication can be performed based on the fault number, equipment identifier, and fault occurrence time window, retaining the sample with the highest record quality. To address the class imbalance where there are many lightning samples and few icing and galloping samples, class reweighting or minority class oversampling can be used during training to ensure that minority disaster types have sufficient influence during training and avoid excessive bias in training results towards high-frequency disaster types.
[0071] Step 3: Divide the labeled sample set into a training set and a validation set, and construct a factor weight matrix according to network type and disaster type. Based on the factor scores of fault samples in the training set and the corresponding weight vectors of network type and disaster type in the factor weight matrix, calculate the judgment score of each disaster type to predict the probability. Calculate the error between the predicted probability and the actual disaster type label to iteratively update the factor weight matrix. Use the validation set to evaluate the judgment performance of the current factor weight matrix. When the judgment performance reaches the convergence condition, determine the candidate weights based on the factor weight matrix, such as... Figure 3 As shown.
[0072] In a preferred but non-limiting embodiment of the present invention, step 3 includes: Step 3.1: Divide the labeled sample set into a supervised training set and a validation set, and construct a factor weight matrix according to network type and disaster type.
[0073] More preferably, step 3.1 includes: Step 3.1.1, label the sample set , This indicates the network type of the s-th labeled sample, where g=1 for the main network and g=2 for the distribution network. This represents the true disaster type label of the s-th labeled sample, such as lightning, icing, or dancing. This represents the factor score of the s-th labeled sample.
[0074] Step 3.1.2, from Samples were randomly selected from the dataset according to a preset ratio (8:2) and divided into supervised training sets. and verification set t represents the index of a sample in the supervised training set, and v represents the index of a sample in the validation set. This represents the network class of the t-th sample in the supervised training set. This represents the true disaster label of the t-th sample in the supervised training set. This represents the factor score of the t-th sample in the supervised training set. This represents the network type of the v-th sample in the validation set. This represents the true disaster label of the v-th sample in the validation set. This represents the factor score of the v-th sample in the validation set.
[0075] Step 3.1.3: Construct factor weight matrices by network type and disaster type. For each network type g, construct the factor weight matrix. , This represents the factor weight corresponding to the d-th disaster type under network type g. This indicates the number of disaster categories. Disaster categories include lightning, icing, and dancing, while network categories include the main network and distribution network.
[0076] Step 3.2: Based on the factor scores of the samples in the supervised training set, perform a weighted operation with the weight vectors of the corresponding network type and disaster type in the factor weight matrix to calculate the assessment score for each disaster type, expressed by the following formula:
[0077] in, This represents the judgment score of the t-th sample in the supervised training set belonging to the d-th disaster type. Indicate network type The factor weight corresponding to the d-th disaster type This represents the network class of the t-th sample in the supervised training set. This represents the factor score of the t-th sample in the supervised training set. Indicate network type The bias term corresponding to the d-th disaster type is optional; it can also be used directly without setting a bias term. Based on the above definition, even if each fault sample corresponds to only one real disaster type label, it is possible to calculate its score as belonging to multiple disaster types such as lightning, icing, and dancing.
[0078] The predicted probability output by the Softmax function includes: the judgment score. By normalizing using the Softmax function, the predicted probability of a sample belonging to each disaster type is obtained, expressed by the following formula:
[0079] in, Let represent the predicted probability that the t-th sample in the supervised training set belongs to the d-th disaster type. Indicates the number of disaster categories. This represents the judgment score of the t-th sample in the supervised training set belonging to the r-th disaster type.
[0080] The disaster type with the highest predicted probability is taken as the final assessment result for this sample. , This represents the predicted disaster label for the t-th sample in the supervised training set. It is used for the engineering management of training data for labeled / unlabeled / conflicting samples to improve training stability and reproducibility.
[0081] Step 3.3: The error between the predicted probability and the true label is calculated using the cross-entropy loss function with class weights, and the factor weight matrix is iteratively updated using the gradient descent method.
[0082] The predicted probability and the true label in step 3.2 are calculated using the loss function. The error between the predicted and actual disaster labels is used to output a scalar loss value. The weight training module takes the training set consisting of labeled samples as input and the difference between the predicted probability and the actual disaster label as the training objective. Preferably, a cross-entropy loss function with class weights is used, with a weight stability constraint term superimposed. The loss function can be expressed as:
[0083] in, Indicates training loss, This represents the set of weights to be trained. This represents the true disaster type of the t-th sample in the supervised training set. The corresponding category weights, This indicates that the t-th sample in the supervised training set corresponds to its true disaster type. The predicted probability, This represents the stability constraint coefficient. Indicate network type The factor weight corresponding to the d-th disaster type Indicates the network type of the previous iteration. The factor weight corresponding to the d-th disaster type To sum all samples in the supervised training set, To sum the results for all networks, To sum up all types of disasters, This is the square of the L2 norm. During training, gradient descent, batch gradient descent, or adaptive optimization methods can be used to update the weights along the descent direction of the loss function. After each update, a non-negative normalization constraint is applied to the weights to maintain their interpretability and stability.
[0084] The gradient of the scalar loss value with respect to the factor weights is calculated using the gradient descent method, and the factor weights are updated in the opposite direction of the gradient.
[0085] Step 3.4: After each iteration, apply a non-negative normalization constraint to the weight vectors in the updated factor weight matrix to normalize the sum of the weight components. Use the validation set to evaluate the performance index of the current factor weight matrix. If the performance index does not meet the preset convergence condition, continue training using the updated factor weight matrix. If the preset convergence condition is met, output the factor weight matrix as a candidate weight.
[0086] More preferably, step 3.4 includes: Step 3.4.1: After each round of gradient update, apply a non-negative normalization constraint to the weight vector obtained from the current training to ensure the physical interpretability of the weights (i.e., each factor weight represents its contribution ratio to the disaster type assessment).
[0087] In the formula, Indicate network type Factor weights corresponding to the d-th disaster type To maintain the interpretability of the "factor score + weighted average" structure, this invention applies a non-negative normalization constraint to the weight vector for the j-th factor weight component. After gradient update, the weights can be processed using projection normalization, specifically as follows: When the denominator is 0, the weight vector can be restored to the previous stable version weights or the initial expert weights. Through the above constraints, each factor weight can be interpreted as the relative contribution of that factor to the corresponding disaster assessment score.
[0088] Step 3.4.2: Use the validation set to evaluate the performance index of the current factor weight matrix; if the performance index does not meet the preset convergence condition, continue training using the updated factor weight matrix; if the preset convergence condition is met, output the factor weight matrix as a candidate weight.
[0089] For each validation sample v, the weight matrix Wgv is called according to its network class gv to calculate the predicted probability pvd. Precision, recall, F1 score, and minority class metric are also calculated. When the F1 score is greater than a set convergence threshold, the preset convergence condition is met, and the final factor weight matrix is set as the candidate weight matrix. .
[0090] Step 4: Perform access control evaluation on candidate weights. For candidate weights that pass the access control evaluation, determine their gray-scale effective range using a preset strategy. Within this gray-scale effective range, evaluate the status of the candidate weights based on online monitoring indicators. If the evaluation status is deteriorated, automatically roll back to the previous stable version of the candidate weights; if the evaluation status is stable, release all candidate weights and update the current candidate weights to the stable version. Figure 4 As shown.
[0091] In a preferred but non-limiting embodiment of the present invention, step 4 includes: Step 4.1: Evaluate the access control system for the candidate weighted versions.
[0092] When the sample size of any minority class in the supervisory data sample set of the candidate weight version is less than the sample size threshold, the access control evaluation is deemed to fail and the candidate weight is archived as unpublishable. Calculate the metrics on the validation set and compare them with the previous stable version to obtain the difference result. When the difference result is less than the metric threshold, the access control evaluation is deemed unsuccessful and the candidate weights are archived as unpublishable. When the change in weight exceeds the stability threshold, the access control evaluation is deemed unsuccessful, and the candidate weights are archived as unpublishable. Only when the sample size threshold, indicator threshold, and stability threshold are all met, the candidate weight version passes the gate and is archived as publishable, proceeding to step 4.2; if any one of the conditions is not met, it is archived as unpublishable.
[0093] Step 4.2: The weights passed through the access control system are written into the weight version library as candidate weight versions, and their version status is marked as "candidate" or "in grayscale". The release and rollback module gradually expands the scope of effectiveness of the candidate weight versions according to a preset grayscale strategy, which includes gradually implementing them according to region, line set, device set, and disaster type ratio. The online analysis module loads the currently effective version in real time or periodically, and records the weight version number used each time the analysis result is output, so as to align and evaluate it with newly added root cause data in the future.
[0094] The newly added root cause data refers to the actual fault cause data of newly added equipment fault records after the candidate weight version is launched, which has been confirmed manually or through an authoritative process, and aligned with the judgment results by the root cause matching module. Based on the newly added root cause data, the release and rollback module continuously monitors the online operation performance of the candidate weight version within the grayscale range. Online monitoring indicators include hit rate, macro average F1 score, minority class recall rate, conflict rate, unlabeled rate, and critical path misjudgment rate. Specifically, the hit rate represents the proportion of disaster type judgment results output by the candidate weight version that match the newly added actual disaster type labels; the macro average F1 score represents the comprehensive judgment effect of each disaster type; the minority class recall rate represents the ability to identify low-frequency disaster types such as icing and galloping; the conflict rate and unlabeled rate represent the stability of the root cause alignment process; and the critical path misjudgment rate represents the judgment risk of key production operation objects.
[0095] Online monitoring results can be categorized into four states: stable, improved, deteriorated, and abnormal. If key indicators such as the candidate version's hit rate, macro average F1 score, and minority class recall rate are within the preset allowable fluctuation range compared to the previous stable version, the candidate version is considered to have stable online performance. If the candidate version's hit rate, macro average F1 score, or minority class recall rate exceeds the previous stable version by more than a set threshold, the candidate version's online performance is considered to have improved. If the candidate version's hit rate, macro average F1 score, or minority class recall rate is lower than the previous stable version over K consecutive monitoring periods, and the degree of decline exceeds the preset deterioration threshold, the candidate version's online performance is considered to have deteriorated. If critical path misjudgment rates exceed the threshold, or version loading fails, it is considered abnormal. The various thresholds can be configured according to region, network type, disaster type, and production operation requirements.
[0096] When the online monitoring results are stable or improved, the release and rollback module continues to expand its effective scope according to the gray-scale strategy; when the candidate weight version continues to remain stable or improved within the preset observation period, and the gray-scale range reaches the preset release ratio, the candidate weight version will be fully released and its version status will be updated to stable version.
[0097] When the online monitoring results show deterioration or anomalies, or when preset rollback conditions are triggered, the release and rollback module automatically stops the continued release of candidate weight versions and rolls back the currently effective version to the previous stable version. At the same time, it records the version number before rollback, the version number after rollback, the rollback reason, the triggering metric, the triggering time, and the version status in the weight version library, thereby achieving traceability, auditability, and rollback capability of the weight update process.
[0098] Through the above steps, the present invention forms a self-learning closed loop of "cause matching → training → access control → gray release → online monitoring → rollback → retraining", which enables the model weights to continuously adapt in the regional and temporal dimensions, while maintaining interpretability and engineering controllability.
[0099] (1) True cause / historical true cause: refers to the classification result of fault cause in the operation and maintenance or fault management system that has been manually confirmed or confirmed by an authoritative process.
[0100] (2) Fault sample: refers to a single fault record, including information such as time, equipment identification, location and related factor scores.
[0101] (3) Factor scoring: The process of mapping raw meteorological or monitoring factors into discrete or continuous scores according to rules, in order to maintain interpretability.
[0102] (4) Device key matching (Match): A function used to determine whether the device identifier in the historical record and the device identifier in the analysis sample belong to the same object or the same level of object.
[0103] (5) Confidence (γ): A numerical value that measures the degree of confidence in the association between a historical cause record and a fault sample, used for candidate ranking and threshold screening.
[0104] (6) Conflicting samples: The same fault sample has multiple root cause candidates that are difficult to resolve, or samples that still do not meet the threshold requirements after resolution.
[0105] (7) Access control: The set of admission criteria before the release of new weights, including sample size, offline indicators and stability thresholds.
[0106] (8) Gray release: A release method in which new weights are first tested in a limited area and then gradually expanded.
[0107] (9) Rollback: When the online monitoring metric triggers the condition, the effective weight version is restored to the previous stable version.
[0108] Embodiment 2 of the present invention provides a power grid meteorological disaster-causing assessment weight update system, which runs the power grid meteorological disaster-causing assessment weight update method described in Embodiment 1, including: The data acquisition module is used to acquire the set of fault samples to be analyzed and the set of historical root causes from power grid business data; The labeling module is used to filter candidate causes from the historical cause set to form a candidate cause set; for each fault sample in the fault sample set, the matching confidence of each candidate cause in the candidate cause set with the fault sample is calculated, and the matching confidence is sorted in descending order. The results are combined with a preset confidence threshold and a conflict difference threshold to determine the labeled sample set. The weight update module is used to divide the labeled sample set into a training set and a validation set, and construct a factor weight matrix according to network type and disaster type. Based on the factor scores of fault samples in the training set and the weight vectors of the corresponding network type and disaster type in the factor weight matrix, the judgment score of each disaster type is calculated to predict the probability. The error between the predicted probability and the actual disaster type label is calculated to iteratively update the factor weight matrix. The judgment performance of the current factor weight matrix is evaluated using the validation set. When the judgment performance reaches the convergence condition, candidate weights are determined according to the factor weight matrix. The stable weight acquisition module is used to perform access control evaluation on candidate weights. The candidate weights that pass the access control evaluation are given a gray-scale effective range through a preset strategy. Within the gray-scale effective range, the status of the candidate weights that pass the access control evaluation is evaluated based on online monitoring indicators. If the evaluation status is deteriorated, it will automatically roll back to the previous stable version of the candidate weights. If the evaluation status is stable, the candidate weights will be released in full, and the current candidate weights will be updated to the stable version of the candidate weights.
[0109] In this invention, the weights obtained after training are called candidate weights, which are not yet directly used for full-scale online evaluation. The candidate weights are packaged with the version number, training time window (the time period covered by the training samples), total number of supervisory data samples, number of samples for each disaster type, and metadata of the verification indicators to form a candidate weight version. After passing the evaluation and access control module's judgment, the candidate weight version can enter the canary release phase. If the online indicators are stable or improve during the canary release, it is archived as a stable version; if the online indicators deteriorate or become abnormal, it is marked as a rolled-back version or an unreleaseable version and reverts to the previous stable version. During initial deployment, the weights configured based on expert experience can be used as the initial version. Each subsequent new version formed by training retains its parent version number or the previous stable version number to achieve traceability and rollback of the weight update process.
[0110] like Figure 5 As shown, this invention can also be written as a data access module M1: used to extract the basic data of "fault samples - factor scores - true cause candidates" required for training from multi-source business data and to encapsulate them uniformly. Combined with... Figure 1 The data sources for M1 include at least: a fault master record database (providing fault occurrence time, equipment / line identifier, network identifier, location, etc.), a factor score detail database (providing factor scores or score details), and a historical true cause database (providing true disaster type labels, occurrence time, equipment identifier, and location, etc.). When a record to be processed appears in the fault master record database, M1 reads the master information of the fault and retrieves the corresponding score from the factor score detail database using the fault number as the key; simultaneously, it retrieves candidate true cause records from the true cause database using time window and equipment key conditions, forming a unified input structure containing fault samples, feature vectors, and candidate sets.
[0111] Root Cause Matching Module M2: Reliably aligns historical root cause records with fault samples, outputting highly reliable labels or conflict / unlabeled states suitable for supervised training. M2 includes a candidate selection unit (based on...) and The system generates a candidate set, a confidence calculation unit (which calculates and sorts the confidence scores of the candidates), and a threshold and conflict resolution unit, which determines the candidates to be unlabeled. It also resolves the conflicts based on rules such as prioritizing candidates with the closest execution time, finer granularity, and higher quality.
[0112] M2 receives the "fault sample + candidate root cause" output by M1, first generates a set of candidate factors, then calculates and sorts the confidence scores; if the top 1 candidate passes the threshold and there is no conflict, it outputs a unique label and attaches matching metadata (candidate ID, γ value, hit rule, etc.); if the conflict cannot be resolved, it marks the sample as a conflict sample for removal or manual review.
[0113] Sample Management Module M3: Used for engineering management of training data output from M2 to improve training stability and reproducibility. M3 may include a sample cleaning unit (handling missing values / outliers / duplicate records), a sample stratification and balancing unit (minority class protection, resampling or class weighting, etc.), and a dataset splitting unit (splitting the training set / validation set / backtesting set by time to avoid information leakage).
[0114] M3 receives three types of data: labeled samples, unlabeled samples, and conflicting samples. Only labeled samples are included in the supervised training queue; conflicting samples can be included in the review queue or temporarily excluded from training; and unlabeled samples can be used for statistical monitoring or semi-supervised extension (optional).
[0115] The weight training module M4 is used to learn candidate weight versions based on training data while maintaining the interpretable structure of "factor scoring + weighted inference". M4 can train weight vectors separately by network type (main network / distribution network) and disaster type, and output training indicators and factor contribution for interpretation and auditing.
[0116] M4 obtains the training and validation sets from M3, generates candidate weights after training, and outputs validation metrics (accuracy, recall, minority class metrics, etc.) and training metadata (training time window, sample size statistics, etc.).
[0117] The evaluation and access control module M5 is used to evaluate and control the admission of candidate weights before they go live. M5 may include an indicator evaluation unit (which calculates indicators on the validation / backtesting set and compares them with the previous stable version) and an access control decision unit (which generates pass / reject conclusions based on sample size thresholds, indicator thresholds, and stability thresholds) to avoid unstable weights directly affecting online evaluation.
[0118] When the sample size of minority classes such as icing / dancing is insufficient or the key indicators deteriorate, M5 will fail the judgment and archive the candidate weight version as "unpublishable"; when the access control conditions are met, it will output "publishable" and proceed to M6.
[0119] The M6 deployment and rollback module is used to deploy weights that pass through the access control system to the online system in a versioned manner, and to implement canary release, online monitoring, and rollback control. M6 includes at least a versioned storage unit (which writes candidate weight versions along with metadata such as training windows, sample statistics, indicators, and applicable scope into the weight version library and maintains the version status), a canary release unit (which gradually activates weights by region, line set, time period, or proportion), an online monitoring unit (which monitors online effects based on the comparison of newly added root cause feedback and analysis results), and a rollback control unit (which reverts to the previous stable version when degradation is triggered and records the reason and timestamp).
[0120] After M6 writes the weights into the weight version library, the online analysis module reads the "current effective version" and records the output results and version number in the analysis result library. When new root cause data arrives and the alignment assessment shows continuous deterioration, M6 automatically rolls back to the previous stable version to achieve a traceable closed-loop update.
[0121] Embodiment 3 of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements a power grid meteorological disaster assessment weight update method according to Embodiment 1.
[0122] Embodiment 4 of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a power grid meteorological disaster assessment weight update method according to Embodiment 1.
[0123] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.
[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A method for updating weights in power grid meteorological disaster assessment, characterized by the following steps: include: Obtain the set of fault samples to be analyzed and the set of historical root causes from power grid business data; Candidate true causes are selected from the set of historical true causes to form a set of candidate true causes; Calculate the matching confidence of each candidate cause in the candidate cause set with each fault sample in the fault sample set. Based on the descending order of the matching confidence, and combined with the preset confidence threshold and the preset conflict difference threshold, determine the labeled sample set. The labeled sample set is divided into a training set and a validation set, and a factor weight matrix is constructed according to network type and disaster type. Based on the factor scores of fault samples in the training set and the corresponding weight vectors of network type and disaster type in the factor weight matrix, the judgment score of each disaster type is calculated to predict the probability. The error between the predicted probability and the actual disaster type label is calculated to iteratively update the factor weight matrix. The judgment performance of the current factor weight matrix is evaluated using the validation set. When the judgment performance reaches the convergence condition, candidate weights are determined according to the factor weight matrix. Candidate weights are evaluated using access control. Candidate weights that pass the access control evaluation are then given a gray-scale effective range determined by a preset strategy. Within the gray-scale effective range, the status of the candidate weights that pass the access control evaluation is assessed based on online monitoring indicators. If the evaluation status is deteriorated, the candidate weights are automatically rolled back to the previous stable version. If the evaluation status is stable, all candidate weights are released, and the current candidate weights are updated to the stable version.
2. The method for updating the weights for assessing meteorological disasters affecting power grids according to claim 1, characterized in that: The steps for determining the labeled sample set include: Based on each fault sample, candidate causes that meet the time window matching and device key matching are selected from the historical cause set to form a candidate cause set. For each fault sample in the fault sample set, calculate the matching confidence of each candidate cause in the candidate cause set with the fault sample, and sort the matching confidence in descending order. If the highest matching confidence is less than the preset confidence threshold, the fault sample is classified as an unlabeled sample; if the difference between the highest matching confidence and the second highest matching confidence is less than the conflict difference threshold and cannot be resolved, the fault sample is classified as a conflict sample; otherwise, the real disaster type label carried by the candidate true cause corresponding to the highest matching confidence is assigned to the fault sample, forming a labeled sample and included in the labeled sample set.
3. The method for updating the weights for assessing meteorological disasters in power grids according to claim 2, characterized in that: The device key includes the device type, device identifier, device topology affiliation, and device line location; If the device key of the fault sample is the same as the device type and device identifier of the device key of the historical cause, it is determined that the device key match is satisfied. If the device key of the fault sample is different from the device key of the historical cause, but the device key of the fault sample and the device key of the historical cause can be mapped to the same power grid object through the predefined topology relationship table, then it is determined that the device key matching is satisfied. If any of the following conditions are met: the device topology of the device key of the fault sample is the same as the device identifier of the device key of the historical cause, the device topology of the device key of the historical cause is the same as the device identifier of the device key of the fault sample, or the device topology of the device key of the fault sample is the same as the device topology of the device key of the historical cause, the device key matching is deemed to be satisfied. All other cases are judged as not meeting the device key matching requirement.
4. The method for updating the weights for assessing meteorological disasters affecting power grids according to claim 2, characterized in that: The matching confidence score between each candidate cause in the candidate cause set and the fault sample is calculated using the following formula: In the formula, This represents the matching confidence score between the i-th fault sample and the m-th candidate cause. , , and Let represent the weights of temporal similarity, spatial similarity, device attribute similarity, and record quality factor, respectively, and all of them be no less than 0, satisfying . , express and Time similarity, This represents the failure occurrence time of the m-th candidate cause. This represents the time of failure occurrence for the i-th fault sample. express and Spatial similarity, This indicates the fault location of the m-th candidate cause. This indicates the fault location of the i-th fault sample. express and The similarity of device attributes, This represents the device key in the m-th candidate true factor. The device key representing the i-th fault sample. The quality factor representing the candidate true cause Indicates the m-th candidate true cause The quality factor.
5. The method for updating the weights for assessing meteorological disasters affecting power grids according to claim 4, characterized in that: The time similarity is expressed by the following formula: In the formula, This represents the allowed matching time window for the d-th disaster type.
6. The method for updating the weights for assessing meteorological disasters affecting power grids according to claim 4, characterized in that: Spatial similarity is expressed by the following formula: In the formula, This represents the distance between the m-th candidate root cause and the i-th fault sample. This represents the spatial matching radius of the d-th disaster type.
7. The method for updating the weights for assessing meteorological disasters affecting power grids according to claim 1, characterized in that: Determining candidate weights includes: The labeled sample set is divided into a supervised training set and a validation set, and a factor weight matrix is constructed according to network type and disaster type. The factor scores based on the samples in the supervised training set are weighted with the weight vectors of the corresponding network type and disaster type in the factor weight matrix to calculate the judgment score of each disaster type, and the prediction probability is output through the Softmax function. The error between the predicted probability and the true label is calculated using the cross-entropy loss function with category weights, and the factor weight matrix is iteratively updated using the gradient descent method. After each iteration, a non-negative normalization constraint is applied to the weight vectors in the updated factor weight matrix to normalize the sum of the weight components. The performance index of the current factor weight matrix is evaluated using the validation set. If the performance index does not meet the preset convergence condition, the updated factor weight matrix is used to continue training. If the preset convergence condition is met, the factor weight matrix is output as a candidate weight.
8. The method for updating the weights for assessing meteorological disasters in power grids according to claim 7, characterized in that: The assessment score for each type of disaster is calculated and expressed by the following formula: in, This represents the judgment score of the t-th sample in the supervised training set belonging to the d-th disaster type. Indicate network type The factor weight vector corresponding to the d-th disaster type. This represents the network class of the t-th sample in the supervised training set. This represents the factor score of the t-th sample in the supervised training set. Indicate network type The bias term corresponding to the d-th disaster type.
9. The method for updating the weights for assessing meteorological disasters affecting power grids according to claim 7, characterized in that: The cross-entropy loss function is expressed by the following formula: in, Indicates training loss, This represents the set of weights to be trained. This represents the true disaster type of the t-th sample in the supervised training set. The corresponding category weights, This indicates that the t-th sample in the supervised training set corresponds to its true disaster type. The predicted probability, This represents the stability constraint coefficient. Indicate network type The factor weight corresponding to the d-th disaster type Indicates the network type of the previous iteration. The factor weight corresponding to the d-th disaster type To sum all samples in the supervised training set, To sum the results for all networks, To sum up all types of disasters, It is the square of the L2 norm.
10. A power grid meteorological disaster-causing assessment weight update system, comprising a power grid meteorological disaster-causing assessment weight update method according to any one of claims 1-9, characterized in that, include: The data acquisition module is used to obtain the set of fault samples to be analyzed and the set of historical root causes from power grid business data; The tagging module is used to filter candidate true causes from the historical true cause set to form a candidate true cause set; Calculate the matching confidence of each candidate cause in the candidate cause set with each fault sample in the fault sample set. Based on the descending order of the matching confidence, and combined with the preset confidence threshold and the preset conflict difference threshold, determine the labeled sample set. The weight update module is used to divide the labeled sample set into a training set and a validation set, and construct a factor weight matrix according to network type and disaster type. Based on the factor scores of fault samples in the training set and the weight vectors of the corresponding network type and disaster type in the factor weight matrix, the judgment score of each disaster type is calculated to predict the probability. The error between the predicted probability and the actual disaster type label is calculated to iteratively update the factor weight matrix. The judgment performance of the current factor weight matrix is evaluated using the validation set. When the judgment performance reaches the convergence condition, candidate weights are determined according to the factor weight matrix. The stable weight acquisition module is used to perform access control evaluation on candidate weights. The candidate weights that pass the access control evaluation are given a gray-scale effective range through a preset strategy. Within the gray-scale effective range, the status of the candidate weights that pass the access control evaluation is evaluated based on online monitoring indicators. If the evaluation status is deteriorated, it will automatically roll back to the previous stable version of the candidate weights. If the evaluation status is stable, the candidate weights will be released in full, and the current candidate weights will be updated to the stable version of the candidate weights.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements a power grid meteorological disaster assessment weight update method according to any one of claims 1-9.
12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements a power grid meteorological disaster assessment weight update method according to any one of claims 1-9.