Source network operation risk monitoring and early warning method and system
By obtaining source network line information for signal feature extraction and analysis, and combining the random forest model and least squares optimization, the problem of misjudgment in power grid source network operation risk monitoring and early warning is solved, and more efficient fault detection and diagnosis is achieved.
Patent Information
- Application Number
- CN202510608169.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing methods for monitoring and warning of power grid source network operation risks are prone to misjudgment, and a more accurate risk monitoring and warning solution is needed.
By obtaining source network line information, performing signal peak detection and waveform analysis, and combining Fourier transform to extract time domain and frequency domain features, a random forest model is used for prediction and a decision boundary is constructed. A comprehensive objective function is constructed by combining fault occurrence time and load data, and the error is optimized through the least squares method to achieve accurate calculation of the fault probability.
It improves the accuracy of power grid source network operation risk monitoring and fault diagnosis capabilities, reduces misjudgments, enhances the sensitivity and adaptability of the model, and can better reflect the complexity of line information and the accuracy of fault location.
Smart Images

Figure CN120151102B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power grid technology, and in particular to a source network operation risk monitoring and early warning method and system. Background Art
[0002] In the prior art, CN112580961A discloses a method for early warning of operation risks based on a power grid information system, comprising: obtaining operation and maintenance data at multiple time nodes at a preset time of the power grid information system; performing linear regression analysis on the data of the same indicator and the same time node under the normal operation of the power grid information system, setting a threshold or threshold range for the operation and maintenance data indicator, and judging whether the operation and maintenance data exceeds the normal range based on the threshold or threshold range; analyzing the cause of the abnormality and determining the type of operation risk; and determining the fault level and issuing an early warning if the detection indicator is in a risk state.
[0003] In summary, although the existing technology can perform risk monitoring and early warning based on power grid source network operation information, it is easy to make misjudgments by judging only by thresholds. Therefore, a solution is needed to solve some of the problems existing in the existing technology. Summary of the Invention
[0004] The embodiments of the present invention provide a source network operation risk monitoring and early warning method and system, which can at least solve some of the problems existing in the prior art.
[0005] A first aspect of an embodiment of the present invention provides a source network operation risk monitoring and early warning method, comprising:
[0006] Obtaining line information of the source network, determining signal peaks based on the line information and performing waveform analysis, selecting time domain features through correlation analysis, converting the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extracting the frequency domain features corresponding to the line information through spectrum analysis;
[0007] Integrating the time domain features and the frequency domain features to obtain a feature input, adding the feature input to a preset random forest model for prediction, combining the prediction outputs of each tree in the random forest model to obtain a comprehensive prediction result, and determining the decision boundary through a sequential minimum optimization algorithm based on the comprehensive prediction result with the goal of maximizing the boundary interval;
[0008] Based on the decision boundary, a comprehensive objective function is constructed in combination with the fault occurrence time and load data. The fault data is processed by the least squares method. The error is minimized by combining the measurement weights corresponding to different measurement points to obtain the fault probability.
[0009] In an optional embodiment,
[0010] The obtaining of line information of the source network, determining a signal peak based on the line information and performing waveform analysis, selecting time domain features through correlation analysis, converting the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extracting the frequency domain features corresponding to the line information through spectrum analysis includes:
[0011] Acquiring an analog signal from a source network line through a pre-installed sensor, recording the signal as line information of the source network, performing waveform analysis on the line information and drawing a corresponding waveform graph, observing the waveform graph and determining basic characteristics of the signal, determining peak points in the waveform based on the waveform graph through a peak detection algorithm, determining similarity between the line information in the waveform graph and a reference signal through a correlation analysis algorithm, and determining time domain characteristics;
[0012] The time domain features corresponding to the line information are converted into frequency domain signals through Fourier transform, and spectrum analysis is performed on the frequency domain signals to extract the frequency domain features.
[0013] In an optional embodiment,
[0014] The step of integrating the time domain features and the frequency domain features to obtain a feature input, adding the feature input to a preset random forest model for prediction, and synthesizing the prediction output of each tree in the random forest model to obtain a comprehensive prediction result includes:
[0015] The time domain features and the frequency domain features are concatenated to generate a feature vector, and the feature vector is recorded as a feature input, and the feature input is added as an input to a preset random forest model;
[0016] Initialize the parameters of the random forest model, set the number of trees, maximum depth and number of leaf nodes, and for each tree in the random forest model, generate a training subset by sampling with replacement from the original data set, randomly select a training feature in the training subset and build a tree based on the current training feature;
[0017] For each tree in the random forest model, at each node, the information gain corresponding to the training feature as input is calculated, the training feature with the largest information gain is selected for splitting, and the splitting is repeated until a preset stopping condition is reached;
[0018] After each tree is constructed, the leaf nodes of the tree are traversed, and the importance of each leaf node in the current tree is calculated respectively. For non-leaf nodes in the tree, the significance corresponding to the non-leaf nodes is determined by a chi-square test. Based on the importance and significance, the confidence corresponding to each node is calculated, and the confidence is compared with a preset confidence threshold. If the confidence of the current node is less than the preset confidence threshold, it is considered that the current node needs to be pruned and a pruning operation is performed; otherwise, the current node is retained;
[0019] Repeatedly construct trees until a preset number is reached, add the feature input to each tree of the random forest model, obtain the prediction results corresponding to each tree, and combine the prediction results to obtain the comprehensive prediction result.
[0020] In an optional embodiment,
[0021] At each node, the information gain corresponding to the training feature as input is calculated as shown in the following formula:
[0022] ;
[0023] in, IG() represents information gain, F Represents a given attribute, S represents a dataset, p i Indicates that the data set belongs to the category i The proportion of samples c represents the total number of categories, n Representation attributes F The number of values of v j Representation attributes F No. j A value, P() represents the probability, p ij Representation attributes F The value is v j Under the conditions, the data set S Analogy i The proportion of samples.
[0024] In an optional embodiment,
[0025] The confidence corresponding to each node is calculated based on the importance and significance as shown in the following formula:
[0026] ;
[0027] in, C R Representation characteristicsR At the confidence level of the current tree, T R Representation characteristics R The number of times it is used in the current tree, T Indicates the total number of features used in the current tree. S R Indicates the significance of the current node, α Represents the weight factor.
[0028] In an optional embodiment,
[0029] The step of determining the decision boundary by a sequence minimum optimization algorithm based on the comprehensive prediction result with the goal of maximizing the boundary interval includes:
[0030] Initializing a first multiplier and a boundary threshold, selecting a second multiplier by a heuristic method based on the comprehensive prediction result, and constructing an optimization function with the goal of maximizing the boundary interval;
[0031] Calculating the corresponding upper bound and lower bound for the first multiplier and the second multiplier, respectively, solving the optimization function analytically to obtain an error buffer, and updating the first multiplier, the second multiplier, and the boundary threshold. For the updated first multiplier and the second multiplier, respectively, determining whether they are within the range corresponding to the upper bound and the lower bound, if so, saving the values; if not, rejecting the update and recalculating using the first multiplier and the second multiplier before the update;
[0032] For each retained first multiplier and second multiplier, determine whether the equality constraints and inequality constraints in the nonlinear programming are satisfied. If so, retain the corresponding decision boundary.
[0033] In an optional embodiment,
[0034] Based on the decision boundary, a comprehensive objective function is constructed in combination with the fault occurrence time and load data. The fault data is processed by the least squares method, and the error is minimized by combining the measurement weights corresponding to different measurement points. The fault probability is obtained, which includes:
[0035] Obtain the fault occurrence time and corresponding load data, determine the measurement weights corresponding to different measurement points, and for each fault, determine the fault path based on the topology of the source network and calculate the time difference between each measurement point receiving the fault signal, which is recorded as fault data;
[0036] Based on the fault data and the load data, a comprehensive objective function is constructed and an initial fault location estimate is set. Based on the initial fault location estimate, the sum of squared residuals between the function prediction value and the actual observation value is minimized by adjusting the parameters in the comprehensive objective function. Combined with the measurement weight, the fault probability is calculated through a probability density function.
[0037] A second aspect of an embodiment of the present invention provides a source network operation risk monitoring and early warning system, including:
[0038] A first unit is configured to obtain line information of a source network, determine a signal peak value based on the line information and perform waveform analysis, select time domain features through correlation analysis, convert the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extract the frequency domain features corresponding to the line information through spectrum analysis;
[0039] A second unit is configured to integrate the time domain features and the frequency domain features to obtain a feature input, add the feature input to a preset random forest model for prediction, integrate the prediction output of each tree in the random forest model to obtain a comprehensive prediction result, and determine the decision boundary based on the comprehensive prediction result by a sequential minimum optimization algorithm with the goal of maximizing the boundary interval;
[0040] The third unit is used to construct a comprehensive objective function based on the decision boundary, combined with the fault occurrence time and load data, process the fault data through the least squares method, combine the measurement weights corresponding to different measurement points, minimize the error, and obtain the fault probability.
[0041] According to a third aspect of the embodiments of the present invention,
[0042] An electronic device is provided, comprising:
[0043] processor;
[0044] a memory for storing processor-executable instructions;
[0045] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0046] According to a fourth aspect of the embodiments of the present invention,
[0047] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0048] In the present invention, by obtaining the line information of the source network and performing signal peak detection and waveform analysis, the signal characteristics of the line can be better understood, including the position of the peak point and the shape of the waveform, providing a basis for subsequent feature extraction. By selecting time domain features through correlation analysis, the relationship between signals at different times can be explored, thereby capturing time domain features related to faults, which helps to improve the sensitivity and accuracy of the model. Integrating the features of the time domain and frequency domain to obtain more comprehensive feature inputs helps to improve the model's ability to represent the system state and better reflect the complexity of line information. The decision boundary is determined by the sequential minimum optimization algorithm to maximize the boundary interval, thereby improving the classification performance of the model and enhancing the ability to distinguish the occurrence of faults. Combining the decision boundary with the fault occurrence time and load data to construct a comprehensive objective function helps the model better adapt to actual data and improve the accuracy of fault location and diagnosis. In summary, the present invention can achieve accurate fault detection in the source network line through multi-level feature extraction, model training and optimization, and the construction of a comprehensive objective function. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Schematic diagram of the process of the source network operation risk monitoring and early warning method according to an embodiment of the present invention;
[0050] Figure 2 This is a relationship diagram between the decision tree node pruning effect and model performance of the source network operation risk monitoring and early warning method according to an embodiment of the present invention;
[0051] Figure 3 This is a comparison chart of the fault location accuracy of the source network operation risk monitoring and early warning method according to an embodiment of the present invention;
[0052] Figure 4 This is a comparison chart of the fault location performance of the source network operation risk monitoring and early warning method according to an embodiment of the present invention;
[0053] Figure 5 This is a schematic diagram of the structure of a source network operation risk monitoring and early warning system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0055] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0056] Figure 1 FIG. 1 is a flow chart of a method for monitoring and warning of source network operation risk according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0057] S1. Obtaining line information from the source network, determining signal peaks based on the line information and performing waveform analysis, selecting time domain features through correlation analysis, converting the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extracting the frequency domain features corresponding to the line information through spectrum analysis;
[0058] The line information refers to the path of signal transmission, including the propagation path of the signal from the transmitter to the receiver, the devices passed through in the middle, the connection method, etc. The signal peak refers to the maximum amplitude value in the signal waveform. The waveform analysis is the process of observing and interpreting the signal waveform. The correlation analysis is used to measure the degree of correlation between two signals. The time domain characteristics refer to the characteristics of the signal on the time axis, including the signal's mean, variance, time domain waveform, peak value, etc. The spectrum analysis is the process of converting the signal to the frequency domain. The frequency domain characteristics include the signal's frequency, spectrum shape, and spectrum peak value, which are used to identify the frequency components of the signal.
[0059] In an optional embodiment,
[0060] The obtaining of line information of the source network, determining a signal peak based on the line information and performing waveform analysis, selecting time domain features through correlation analysis, converting the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extracting the frequency domain features corresponding to the line information through spectrum analysis includes:
[0061] Acquiring an analog signal from a source network line through a pre-installed sensor, recording the signal as line information of the source network, performing waveform analysis on the line information and drawing a corresponding waveform graph, observing the waveform graph and determining basic characteristics of the signal, determining peak points in the waveform based on the waveform graph through a peak detection algorithm, determining similarity between the line information in the waveform graph and a reference signal through a correlation analysis algorithm, and determining time domain characteristics;
[0062] The time domain features corresponding to the line information are converted into frequency domain signals through Fourier transform, and spectrum analysis is performed on the frequency domain signals to extract the frequency domain features.
[0063] The peak point refers to the maximum amplitude reached by the signal at a certain moment. The analog signal is a continuous signal and can take any value at any time. The correlation analysis is used to measure the similarity between two signals.
[0064] The analog signal of the source network line is acquired through pre-installed sensors and collected as digital data. The waveform of the collected analog signal is analyzed. The overall shape, amplitude, periodicity, etc. of the signal are observed by drawing a waveform graph. The peak point in the waveform is identified using a peak detection algorithm. For example, a threshold can be set. When the signal exceeds the threshold, it is considered that a peak has occurred. A reference signal is selected and the similarity between the reference signal and the line information is compared using a correlation analysis algorithm to extract time domain features, such as the mean, variance, and peak value of the signal.
[0065] The line information in the time domain is converted into a frequency domain signal through Fourier transform, and the frequency domain signal is subjected to spectrum analysis to extract frequency domain features. For example, the spectrum graph can be observed to identify the main frequency components and measure the width of the spectrum.
[0066] For example, suppose a sensor acquires an analog signal from a source network line over a period of time. The sensor collects the analog signal once per second, resulting in one minute of data, totaling 60 data points. A waveform graph of these 60 data points is plotted, revealing that the signal exhibits large amplitude fluctuations in the first half of the period, while the amplitude gradually decreases in the second half. Using a peak detection algorithm, three distinct peaks corresponding to the peaks of the signal are detected in the first half of the period. A known reference signal is selected, and a correlation analysis algorithm is used to calculate the correlation between the line information and the reference signal. The resulting correlation value indicates that the line information is similar to the reference signal in some respects. Time-domain characteristics of the line information, such as mean, variance, and peak value, are calculated. Assuming the signal has a mean of 0.5, a variance of 0.2, and a maximum peak value of 1.2, the line information is Fourier transformed to obtain a spectrum. The spectrum is then observed and analyzed to extract the primary frequency components and the spectrum width. For example, if the primary frequency is 10 Hz and the spectrum width is 2 Hz, the correlation value is 10 Hz.
[0067] In this embodiment, by drawing a waveform graph, the changing trend, amplitude and possible periodicity of the signal over the entire time period can be intuitively observed, which provides a basis for subsequent signal analysis. By performing correlation analysis with a reference signal, the similarity between the line information and the reference signal can be quantitatively evaluated, which helps to understand the relationship between the signal and the expected signal or the baseline signal. By performing spectrum analysis on the frequency domain signal, the main frequency components and spectrum characteristics can be extracted, which helps to further understand the frequency domain characteristics of the signal. In summary, this embodiment helps to identify key information, abnormal events or periodic changes in the signal, and provides a strong foundation for subsequent signal processing and analysis.
[0068] S2. Integrating the time domain features and the frequency domain features to obtain a feature input, adding the feature input to a preset random forest model for prediction, combining the prediction outputs of each tree in the random forest model to obtain a comprehensive prediction result, and determining the decision boundary using a sequential minimum optimization algorithm based on the comprehensive prediction result with the goal of maximizing the boundary interval;
[0069] The feature input refers to attributes or features used to describe various aspects of a data sample. The random forest model is an integrated learning method consisting of multiple decision trees. Each decision tree is trained based on randomly selected features and samples. Classification or regression is performed by voting or averaging the outputs of each decision tree. The boundary interval refers to the distance between the decision boundary (or hyperplane) and the training sample closest to it. The decision boundary is the boundary or hyperplane that the model uses to demarcate areas of different categories in the feature space.
[0070] In an optional embodiment,
[0071] The step of integrating the time domain features and the frequency domain features to obtain a feature input, adding the feature input to a preset random forest model for prediction, and synthesizing the prediction output of each tree in the random forest model to obtain a comprehensive prediction result includes:
[0072] The time domain features and the frequency domain features are concatenated to generate a feature vector, and the feature vector is recorded as a feature input, and the feature input is added as an input to a preset random forest model;
[0073] Initialize the parameters of the random forest model, set the number of trees, maximum depth and number of leaf nodes, and for each tree in the random forest model, generate a training subset by sampling with replacement from the original data set, randomly select a training feature in the training subset and build a tree based on the current training feature;
[0074] For each tree in the random forest model, at each node, the information gain corresponding to the training feature as input is calculated, the training feature with the largest information gain is selected for splitting, and the splitting is repeated until a preset stopping condition is reached;
[0075] After each tree is constructed, the leaf nodes of the tree are traversed, and the importance of each leaf node in the current tree is calculated respectively. For non-leaf nodes in the tree, the significance corresponding to the non-leaf nodes is determined by a chi-square test. Based on the importance and significance, the confidence corresponding to each node is calculated, and the confidence is compared with a preset confidence threshold. If the confidence of the current node is less than the preset confidence threshold, it is considered that the current node needs to be pruned and a pruning operation is performed; otherwise, the current node is retained;
[0076] Repeatedly construct trees until a preset number is reached, add the feature input to each tree of the random forest model, obtain the prediction results corresponding to each tree, and combine the prediction results to obtain the comprehensive prediction result.
[0077] The training subset refers to the part of data extracted from the entire data set for training the model, the training features refer to the input variables used to learn the model during the model training process, the chi-square test is a statistical method used to test whether there is independence between two categorical variables, the confidence level refers to the confidence level of an estimate or hypothesis in statistical inference, and the significance generally refers to the significance level of the statistical test results, that is, the degree of rejection of the null hypothesis.
[0078] Based on the pre-acquired time domain features and frequency domain features, the time domain features and the frequency domain features are concatenated to obtain a feature input;
[0079] Initializing the parameters of the random forest model, setting the number of trees, maximum depth, number of leaf nodes, etc., performing sampling with replacement on the original data set to generate a training subset. The sampling process can repeatedly select the same sample, so that the same sample may be selected multiple times. Adding the feature input to the random forest model, for each tree in the random forest model, randomly selecting a training feature from the training subset, and constructing a tree in the random forest model based on the training feature;
[0080] For each tree in the random forest model, at each node, the information gain corresponding to all possible features is calculated, and the feature with the largest information gain is selected for splitting, and the splitting is repeated until a preset stopping condition is reached;
[0081] After each tree is built, the leaf nodes of the tree are traversed to calculate the importance of each leaf node in the current tree. For example, the importance of the node can be determined by calculating the purity of the node. A chi-square test is performed on each non-leaf node to determine its significance. For each node, its confidence is calculated. Considering the importance and significance of the node, if the confidence of the node is lower than the preset confidence threshold, it is considered that the node needs to be pruned and a pruning operation is performed;
[0082] Repeat the above steps to build the next tree until the preset number of trees is reached. Add the feature input to each tree and obtain the prediction results corresponding to each tree. Combine the prediction results of all trees to obtain the final comprehensive prediction result by averaging.
[0083] Figure 2This figure shows the relationship between the effect of decision tree node pruning and model performance for the source network operation risk monitoring and early warning method according to an embodiment of the present invention. This figure analyzes the relationship between the effect of decision tree node pruning and model performance. The horizontal axis represents the confidence threshold, the left vertical axis represents the prediction accuracy, and the right vertical axis represents the average number of nodes. The data clearly shows that as the confidence threshold increases, the accuracy of this technical solution initially increases and then slightly decreases, reaching a peak of 98.7% at a confidence threshold of 0.5. The accuracy of the traditional random forest without pruning remains roughly between 93% and 94%, a significant difference. Furthermore, as the confidence threshold increases, the average number of nodes increases from 85 nodes at a confidence threshold of 0.1 to 460 nodes at a confidence threshold of 0.6, a 5.4-fold increase. The figure shows that the optimal balance point is at a confidence threshold of 0.4, at which point the accuracy of this technical solution reaches 98.2% with an average number of nodes of 295, a 4.2 percentage point improvement compared to the accuracy of the unpruned method with a similar number of nodes.
[0084] Figure 2 This clearly demonstrates that the chi-square test and confidence-based pruning strategy employed in this technical solution effectively removes redundant nodes from the model, reducing the risk of overfitting while maintaining high prediction accuracy. Specifically, in the confidence threshold range from 0.2 to 0.4, accuracy increased by 2.0 percentage points (from 96.2% to 98.2%), while only increasing the number of nodes by 138 (from 157 to 295). This demonstrates that the pruning strategy in this range achieves the optimal performance-complexity balance, improving the model's generalization while significantly reducing computational complexity and storage requirements.
[0085] In this embodiment, time domain features and frequency domain features are spliced into feature vectors, which helps to comprehensively consider the time domain and frequency domain information of the signal, more comprehensively describe the characteristics of the source network line analog signal, and improve the model's ability to understand the signal. At each node of each tree, the information gain of the input feature is calculated, and the feature with the largest information gain is selected for splitting, which helps the decision tree select the feature that contributes most to the prediction target at each node, thereby improving the accuracy of the model. After each tree is constructed, the importance of each leaf node in the current tree is calculated, which helps to understand which features contribute more to the prediction of the overall model, thereby performing more in-depth feature analysis. The significance of non-leaf nodes is determined by the chi-square test, and the confidence of each node is calculated based on the importance and significance of the node, which helps to generalize and simplify the model and reduce overfitting. In summary, this embodiment establishes a random forest model with both predictive accuracy and generalization ability, which is used to classify source network line analog signals or perform other prediction tasks.
[0086] In an optional embodiment,
[0087] At each node, the information gain corresponding to the training feature as input is calculated as shown in the following formula:
[0088] ;
[0089] in, IG() represents information gain, F Represents a given attribute, S represents a dataset, p i Indicates that the data set belongs to the category i The proportion of samples c represents the total number of categories, n Representation attributes F The number of values of v j Representation attributes F No. j A value, P() represents the probability, p ij Representation attributes F The value is v j Under the conditions, the data set S Analogy i The proportion of samples.
[0090] In this function, information gain, as an indicator of decision tree node splitting, can help select the best features for data set division. When the information gain is large, it means that selecting this feature can better classify the data and improve the prediction accuracy of the model. Information gain can be used to determine the contribution of the feature to the classification task and reduce the uncertainty of the data set. In summary, this function provides an effective feature selection method that helps to build a decision tree model with good generalization performance.
[0091] In an optional embodiment,
[0092] The confidence corresponding to each node is calculated based on the importance and significance as shown in the following formula:
[0093] ;
[0094] in, C R Representation characteristics R At the confidence level of the current tree, T R Representation characteristics R The number of times it is used in the current tree, T Indicates the total number of features used in the current tree. S R Indicates the significance of the current node, αRepresents the weight factor.
[0095] In this function, by using confidence, nodes can be evaluated during the model construction process, improving the model's interpretability and generalization capabilities. By identifying and pruning nodes that contribute less to the model, the efficiency and generalization of the model can be further improved. In summary, this function considers the frequency of feature usage, the significance of the feature at the node, and the influence of the weight factor to obtain the confidence of each node, which helps to build a high-performance decision tree and generate accurate prediction results.
[0096] S3. Based on the decision boundary, a comprehensive objective function is constructed in combination with the fault occurrence time and load data. The fault data is processed using the least squares method. The error is minimized by combining the measurement weights corresponding to different measurement points to obtain the fault probability.
[0097] The least squares method is a method for estimating model parameters, and is particularly suitable for linear regression problems. The estimated value of the parameter is obtained by minimizing the sum of squares of the residuals between the observed value and the model predicted value. The measurement weight represents the credibility of the measurement value.
[0098] In an optional embodiment,
[0099] The step of determining the decision boundary by a sequence minimum optimization algorithm based on the comprehensive prediction result with the goal of maximizing the boundary interval includes:
[0100] Initializing a first multiplier and a boundary threshold, selecting a second multiplier by a heuristic method based on the comprehensive prediction result, and constructing an optimization function with the goal of maximizing the boundary interval;
[0101] Calculating the corresponding upper bound and lower bound for the first multiplier and the second multiplier, respectively, solving the optimization function analytically to obtain an error buffer, and updating the first multiplier, the second multiplier, and the boundary threshold. For the updated first multiplier and the second multiplier, respectively, determining whether they are within the range corresponding to the upper bound and the lower bound, if so, saving the values; if not, rejecting the update and recalculating using the first multiplier and the second multiplier before the update;
[0102] For each retained first multiplier and second multiplier, determine whether the equality constraints and inequality constraints in the nonlinear programming are satisfied. If so, retain the corresponding decision boundary.
[0103] The boundary threshold usually refers to a critical value. Exceeding or falling below this threshold may lead to different results or decisions. The heuristic method is a problem-solving strategy, usually a simplified method based on experience or rules. The upper bound is the maximum value that the objective function may reach, and the lower bound is the minimum value that may be reached. Understanding the upper and lower bounds helps to evaluate the difficulty of the problem and the performance of the optimization algorithm. The equality constraint refers to restricting the value of an equation to be equal to a constant, and the inequality constraint refers to restricting the value of an equation to be within a certain range.
[0104] Initialize the first multiplier and boundary threshold, set the parameters of the heuristic method, such as the learning rate, the number of iterations, etc., and select the second multiplier using the heuristic method based on the comprehensive prediction results, with the goal of maximizing the boundary interval. For example, the second multiplier can be obtained by traversing the samples and selecting the multiplier corresponding to the sample with the largest boundary interval. The prediction result is calculated by using the kernel function and the sample label, and the optimization function is constructed by combining the first multiplier and the second multiplier, with the goal of maximizing the boundary interval.
[0105] Calculate the corresponding upper and lower bounds for the first and second multipliers respectively, which will be used to determine the feasibility of the update. Use analytical methods such as gradient descent to solve the constructed optimization function to obtain the optimal first and second multipliers, as well as the corresponding error buffer. Based on the results obtained, update the first and second multipliers and boundary thresholds. For the updated first and second multipliers, determine whether they are within the corresponding upper and lower bounds. If so, save the update; otherwise, reject the update and recalculate using the parameters before the update.
[0106] For each retained first multiplier and second multiplier, determine whether the equality constraints and inequality constraints in the nonlinear programming are satisfied. If so, retain the corresponding decision boundary and repeat the operation until the preset number of iterations or convergence conditions are reached.
[0107] For example, the first multiplier and the second multiplier are initialized to 0, the boundary threshold is also 0, the second multiplier is selected by a heuristic method, and a sample point that does not meet the KKT condition is selected. Assume that sample 1 is selected, and construct an optimization function with the goal of maximizing the boundary interval. Assume that the upper and lower bounds are [0,10]. Assume that the optimal first multiplier is 5 and the second multiplier is 3 through analytical solution, and update the corresponding optimal values according to the numerical values. Since the first and second multipliers are between the upper and lower bounds, the updated first and second multipliers are saved.
[0108] In this embodiment, by selecting the second multiplier and constructing an optimization function, the goal is to maximize the boundary interval, which helps to improve the generalization ability of the model and make it more robust to the classification of new samples. The constructed optimization function is solved analytically to obtain the optimal first multiplier and second multiplier, which can efficiently find the global optimal solution and avoid the complexity of finding the solution through an iterative method. For the updated first multiplier and second multiplier, respectively determining whether they are within the range corresponding to the upper and lower bounds helps to maintain a reasonable range of parameters and avoid excessive adjustment. In summary, this embodiment can improve the generalization ability of the model, thereby making the model more robust and reliable in practical applications.
[0109] In an optional embodiment,
[0110] Based on the decision boundary, a comprehensive objective function is constructed in combination with the fault occurrence time and load data. The fault data is processed by the least squares method, and the error is minimized by combining the measurement weights corresponding to different measurement points. The fault probability is obtained, which includes:
[0111] Obtain the fault occurrence time and corresponding load data, determine the measurement weights corresponding to different measurement points, and for each fault, determine the fault path based on the topology of the source network and calculate the time difference between each measurement point receiving the fault signal, which is recorded as fault data;
[0112] Based on the fault data and the load data, a comprehensive objective function is constructed and an initial fault location estimate is set. Based on the initial fault location estimate, the sum of squared residuals between the function prediction value and the actual observation value is minimized by adjusting the parameters in the comprehensive objective function. Combined with the measurement weight, the fault probability is calculated through a probability density function.
[0113] The load data refers to the load consumption of each node or area in the power system, including real-time load, historical load and forecast data of future load. The topology structure represents the connection relationship between various components (generators, transformers, lines, etc.) in the power system. The initial fault location estimation refers to the preliminary determination of the location of the fault through measurement and monitoring data when a fault occurs. The residual sum of squares is the sum of the squares of the differences between the measured values and the estimated values. By minimizing the residual sum of squares, a more accurate estimate of the power system state can be obtained. The probability density function describes the probability distribution of a random variable near a certain value point.
[0114] Obtain the timestamp of the fault occurrence and the associated load data from the monitoring system or other data sources. Based on the system topology and the location of the measurement points, determine the importance of each measurement point for fault location, and then determine the measurement weight. For each fault, determine the possible fault path based on the system topology. By considering the speed of electromagnetic waves propagating in the power system and the length of the path, calculate the time difference between each measurement point receiving the fault signal and record it as fault data.
[0115] Using fault data and load data, a comprehensive objective function is constructed. The objective function includes the residual sum of squares and other influencing factors. Prior knowledge and historical data are used to set an initial location estimate for each fault. By adjusting the parameters in the comprehensive objective function, the residual sum of squares between the function's predicted value and the actual observed value is minimized. Combined with the determined measurement weights and using a probability density function, the probability of each fault is calculated by considering the weights of the measurement points and the probability distribution of the residuals to obtain the fault probability.
[0116] For example, assume there is a simplified power system, including three nodes (A, B, C) and two measurement points (M1, M2). Fault 1 occurs at time t1=10 seconds, and fault 2 occurs at time t2=20 seconds. The corresponding load data includes the loads of nodes A, B, and C at these two time points. The measurement weight is determined based on the distance between the measurement point and the node. Assuming that M1 is closer to node A, the measurement weight for node A is larger. Assuming that the measurement weight of point A is 0.6 and the measurement weight of point B is 0.4, for each fault, the possible paths of fault signal propagation are calculated, and the time difference of receiving the fault signal at each measurement point is estimated. An objective function is constructed, considering factors such as the residual sum of squares and measurement weights, and an initial location estimate is set for each fault. Prior knowledge or simple rules can be used. For example, if the fault occurs at node A, the parameters in the objective function are adjusted through optimization algorithms such as gradient descent to minimize the residual sum of squares. The probability of each fault is calculated using the measurement weight and the probability distribution of the residual. Assuming the fault probability is 0.5, the probability of a fault at node A is 0.5.
[0117] In this embodiment, by considering the time difference between the load data of multiple measurement points and the fault path, the system can more accurately locate the location of the fault, adjust the parameters to minimize the sum of squared residuals in the objective function, thereby improving the degree of fit of the model to the observed data, and considering the information of the measurement weight and probability density function, the calculation of the fault probability is more credible and can better reflect the system status. In summary, this embodiment combines the calculation of measurement weight and probability density function to obtain a reliable estimate of the fault probability, which helps system operation and maintenance personnel make more informed decisions in fault diagnosis and maintenance.
[0118] Figure 3 This is a comparison chart of the fault location accuracy of the source network operation risk monitoring and early warning method according to an embodiment of the present invention. Figure 3 As shown, Figure 3 The fault location accuracy comparison of this technical solution, the Bayesian estimation method, and the traditional least squares method under different system load rate conditions is shown. It can be clearly seen from the figure that this technical solution has significant advantages under all load rate conditions. Under light load conditions with a load rate of 20%, the accuracy of this technical solution reaches 92.5%, while the Bayesian estimation method and the traditional least squares method are only 89.8% and 85.0%, respectively. As the load rate increases, the accuracy of the three methods increases, but the advantage of this technical solution is more obvious, reaching a peak accuracy of 97.7% at a load rate of 100%, which is 2.2 percentage points higher than the Bayesian estimation method and 8.5 percentage points higher than the traditional least squares method. It is worth noting that when the load factor exceeds 100% and enters an overload state, the accuracy of the three methods all decreases slightly, but this technical solution still maintains a high accuracy of 97.4% at a load factor of 120%, indicating that it has stronger adaptability and stability under extreme load conditions. This is mainly due to the fact that this solution combines the fault occurrence time and load data in the fault location process. Through the comprehensive objective function and measurement weight optimization, the accuracy and reliability of fault location are greatly improved.
[0119] Figure 4 This is a comparison chart of the fault location performance of the source network operation risk monitoring and early warning method according to an embodiment of the present invention. Figure 4A comprehensive comparison of the fault location performance of this technical solution with that of the adaptive impedance method and the improved traveling wave method under various system conditions was conducted. The data in the table shows that under all system conditions, this technical solution significantly outperformed the other two methods in terms of three key metrics: average location error, location time, and accuracy. Under normal load conditions, the average location error of this technical solution was only 0.87%, compared to 1.53% for the adaptive impedance method and 1.78% for the improved traveling wave method. Under heavy load conditions, the location errors of all three methods increased, but this technical solution maintained a low error level of 0.92%. Under light load conditions, the error of this technical solution increased slightly to 1.05%, but still significantly outperformed the other methods. Particularly noteworthy is that under complex system conditions with distributed power generation, the location error of this technical solution was only 1.18%, while the errors of the adaptive impedance method and the improved traveling wave method reached as high as 2.34% and 2.92%, respectively, demonstrating that this technical solution has greater adaptability to system complexity. Furthermore, this technical solution maintains a location time of less than 9ms under all system conditions, approximately 45% faster than the adaptive impedance method and 65% faster than the improved traveling wave method, respectively. Furthermore, its accuracy remains above 92% under all system conditions, reaching a peak accuracy of 96.8% under heavy load conditions, demonstrating its outstanding performance in power system fault location.
[0120] Figure 5 FIG. 1 is a schematic diagram of the structure of a source network operation risk monitoring and early warning system according to an embodiment of the present invention. Figure 5 As shown, the system includes:
[0121] A first unit is configured to obtain line information of a source network, determine a signal peak value based on the line information and perform waveform analysis, select time domain features through correlation analysis, convert the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extract the frequency domain features corresponding to the line information through spectrum analysis;
[0122] A second unit is configured to integrate the time domain features and the frequency domain features to obtain a feature input, add the feature input to a preset random forest model for prediction, integrate the prediction output of each tree in the random forest model to obtain a comprehensive prediction result, and determine the decision boundary based on the comprehensive prediction result by a sequential minimum optimization algorithm with the goal of maximizing the boundary interval;
[0123] The third unit is used to construct a comprehensive objective function based on the decision boundary, combined with the fault occurrence time and load data, process the fault data through the least squares method, combine the measurement weights corresponding to different measurement points, minimize the error, and obtain the fault probability.
[0124] According to a third aspect of the embodiments of the present invention,
[0125] An electronic device is provided, comprising:
[0126] processor;
[0127] a memory for storing processor-executable instructions;
[0128] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0129] According to a fourth aspect of the embodiments of the present invention,
[0130] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0131] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A source network operation risk monitoring and early warning method, characterized in that: include: Obtaining line information of the source network, determining signal peaks based on the line information and performing waveform analysis, selecting time domain features through correlation analysis, converting the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extracting the frequency domain features corresponding to the line information through spectrum analysis; The time domain features and the frequency domain features are integrated to obtain feature inputs, the feature inputs are added to a preset random forest model for prediction, the prediction outputs of each tree in the random forest model are combined to obtain a comprehensive prediction result, and based on the comprehensive prediction result, a decision boundary is determined by a sequential minimum optimization algorithm with the goal of maximizing the boundary interval, including: The time domain features and the frequency domain features are concatenated to generate a feature vector, and the feature vector is recorded as a feature input, and the feature input is added as an input to a preset random forest model; Initialize the parameters of the random forest model, set the number of trees, maximum depth and number of leaf nodes, and for each tree in the random forest model, generate a training subset by sampling with replacement from the original data set, randomly select a training feature in the training subset and build a tree based on the current training feature; For each tree in the random forest model, at each node, the information gain corresponding to the training feature as input is calculated, the training feature with the largest information gain is selected for splitting, and the splitting is repeated until a preset stopping condition is reached; After each tree is constructed, the leaf nodes of the tree are traversed, and the importance of each leaf node in the current tree is calculated respectively. For non-leaf nodes in the tree, the significance corresponding to the non-leaf nodes is determined by a chi-square test. Based on the importance and significance, the confidence corresponding to each node is calculated, and the confidence is compared with a preset confidence threshold. If the confidence of the current node is less than the preset confidence threshold, it is considered that the current node needs to be pruned and a pruning operation is performed; otherwise, the current node is retained; Repeatedly constructing trees until a preset number is reached, adding the feature input to each tree of the random forest model, obtaining the prediction results corresponding to each tree, and combining the prediction results to obtain the comprehensive prediction result; The confidence corresponding to each node is calculated based on the importance and significance as shown in the following formula: ; in, C R Representation characteristics R At the confidence level of the current tree, T R Representation characteristics R The number of times it is used in the current tree, T Indicates the total number of features used in the current tree. S R Indicates the significance of the current node, α represents the weight factor; Initializing a first multiplier and a boundary threshold, selecting a second multiplier by a heuristic method based on the comprehensive prediction result, and constructing an optimization function with the goal of maximizing the boundary interval; Calculating the corresponding upper bound and lower bound for the first multiplier and the second multiplier, respectively, solving the optimization function analytically to obtain an error buffer, and updating the first multiplier, the second multiplier, and the boundary threshold. For the updated first multiplier and the second multiplier, respectively, determining whether they are within the range corresponding to the upper bound and the lower bound, if so, saving the values; if not, rejecting the update and recalculating using the first multiplier and the second multiplier before the update; For each retained first multiplier and second multiplier, determine whether the equality constraints and inequality constraints in the nonlinear programming are satisfied. If so, retain the corresponding decision boundary. Based on the decision boundary, a comprehensive objective function is constructed in combination with the fault occurrence time and load data. The fault data is processed by the least squares method. The error is minimized by combining the measurement weights corresponding to different measurement points to obtain the fault probability.
2. The method according to claim 1, characterized in that The obtaining of line information of the source network, determining a signal peak based on the line information and performing waveform analysis, selecting time domain features through correlation analysis, converting the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extracting the frequency domain features corresponding to the line information through spectrum analysis includes: Acquiring an analog signal from a source network line through a pre-installed sensor, recording the signal as line information of the source network, performing waveform analysis on the line information and drawing a corresponding waveform graph, observing the waveform graph and determining basic characteristics of the signal, determining peak points in the waveform based on the waveform graph through a peak detection algorithm, determining similarity between the line information in the waveform graph and a reference signal through a correlation analysis algorithm, and determining time domain characteristics; The time domain features corresponding to the line information are converted into frequency domain signals through Fourier transform, and spectrum analysis is performed on the frequency domain signals to extract the frequency domain features.
3. The method according to claim 1, characterized in that At each node, the information gain corresponding to the training feature as input is calculated as shown in the following formula: ; in, IG() represents information gain, F Represents a given attribute, S represents a dataset, p i Indicates that the data set belongs to the category i The proportion of samples c represents the total number of categories, n Representation attributes F The number of values of v j Representation attributes F No. j A value, P() represents the probability, p ij Representation attributes F The value is v j Under the conditions, the data set S Analogy i The proportion of samples.
4. The method according to claim 1, wherein Based on the decision boundary, a comprehensive objective function is constructed in combination with the fault occurrence time and load data. The fault data is processed by the least squares method, and the error is minimized by combining the measurement weights corresponding to different measurement points. The fault probability is obtained, which includes: Obtain the fault occurrence time and corresponding load data, determine the measurement weights corresponding to different measurement points, and for each fault, determine the fault path based on the topology of the source network and calculate the time difference between each measurement point receiving the fault signal, which is recorded as fault data; Based on the fault data and the load data, a comprehensive objective function is constructed and an initial fault location estimate is set. Based on the initial fault location estimate, the sum of squared residuals between the function prediction value and the actual observation value is minimized by adjusting the parameters in the comprehensive objective function. Combined with the measurement weight, the fault probability is calculated through a probability density function.
5. A source network operation risk monitoring and early warning system, configured to implement the source network operation risk monitoring and early warning method according to any one of claims 1 to 4, characterized in that: include: A first unit is configured to obtain line information of a source network, determine a signal peak value based on the line information and perform waveform analysis, select time domain features through correlation analysis, convert the time domain signal corresponding to the line information into a frequency domain signal through Fourier transform, and extract the frequency domain features corresponding to the line information through spectrum analysis; A second unit is configured to integrate the time domain features and the frequency domain features to obtain a feature input, add the feature input to a preset random forest model for prediction, integrate the prediction output of each tree in the random forest model to obtain a comprehensive prediction result, and determine the decision boundary based on the comprehensive prediction result by a sequential minimum optimization algorithm with the goal of maximizing the boundary interval; The third unit is used to construct a comprehensive objective function based on the decision boundary, combined with the fault occurrence time and load data, process the fault data through the least squares method, combine the measurement weights corresponding to different measurement points, minimize the error, and obtain the fault probability.
6. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 4.
7. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Method and device for analyzing VoLTE network fault reasons based on random forest
CN110474786A
Power grid fault detecting and positioning method and system for intelligent power grid system
CN119492958A
Power transmission line forest fire risk assessment and early warning system and method
CN119811051A