Electric power data analysis method and system based on AI model
Through the power data analysis method based on AI model, the problem of traditional methods being difficult to deal with complex power data and accurately predicting power loads is solved, and the safety and efficiency of power grid operation is improved. By identifying potential safety threats and automatic response mechanisms, risks and losses are reduced, and long-term stability is ensured through continuous learning.
Patent Information
- Application Number
- CN202510339495.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional power data analysis methods are difficult to process large-scale and high-dimensional complex power data, and cannot accurately predict power loads and timely detect faults, resulting in low grid operation safety and inefficiency.
Power data analysis method based on AI model is adopted, by collecting and preprocessing the data of power grid monitoring points, calculating data stability baseline values, training AI models to identify behaviors that deviate from the normal range, judging potential security threats, and calculating risk scores through in-depth analysis, automatically triggering response mechanisms and learning strategies for updating AI models.
It significantly improves the safety and efficiency of power grid operations, quickly and accurately identify potential security threats, reduces false alarm rates and missed alarm rates, provides multi-dimensional risk assessments, helps decision makers take the most appropriate response measures, reduce losses caused by emergencies, and ensures long-term stability and reliability through continuous learning and self-optimization.
Smart Images

Figure CN120197104A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence, power systems, and data processing and analysis, and particularly to a power data analysis method and system based on an AI model. Background Art
[0002] With the continuous development of power systems and the advancement of the intelligentization process, the demand for power data analysis is increasing day by day. Traditional power data analysis methods often rely on manual experience and simple statistical models, and it is difficult to process large-scale and high-dimensional complex power data.
[0003] At the same time, the operating state of the power system is affected by various factors, such as meteorological conditions, user electricity consumption behavior, etc., making power data have characteristics such as non-linearity, time-variation, and uncertainty. Traditional methods face challenges in accurately predicting power loads and timely detecting faults. The rapid development of artificial intelligence technology provides new ideas and methods for power data analysis. AI models have powerful learning and generalization abilities, and can automatically extract features and rules from a large amount of data, providing more accurate and efficient decision-making support for the optimal operation and management of power systems. Summary of the Invention
[0004] A power data analysis method and system based on an AI model, comprising:
[0005] S10. The power data analysis system collects and preprocesses data of grid monitoring points, and calculates the data stability baseline value;
[0006] S20. Based on the data stability baseline value, the power data analysis system trains an AI model to identify behaviors deviating from the normal range and judge potential security threats;
[0007] S30. For the data marked by the AI model as having security threats, the power data analysis system conducts in-depth analysis and calculates a risk score that comprehensively considers the degree of abnormality, duration, and scope of influence;
[0008] S40. According to the risk score, the system automatically triggers the corresponding response mechanism;
[0009] S50. The power data analysis system continuously updates the learning strategy of the AI model by evaluating the effects of processed events.
[0010] The power data analysis method based on an AI model as described above, wherein the power data analysis system collects and preprocesses data of grid monitoring points and calculates the data stability baseline value, including the following sub-steps:
[0011] Collect power operation data from grid monitoring points and store it in a database;
[0012] Clean and standardize the collected data;
[0013] Determine the stability baseline value by calculating the mean and standard deviation of historical data.
[0014] The power data analysis method based on the AI model as described above, wherein, based on the data stability baseline value, the power data analysis system trains the AI model to identify behaviors deviating from the normal range and judge potential security threats, including the following sub-steps:
[0015] Divide the preprocessed historical data into a training set, a validation set and a test set, and extract relevant features to form a feature vector;
[0016] Select a suitable AI model architecture, train it using the training set, adjust the parameters through the validation set to prevent overfitting, and finally evaluate the model performance using the test set.
[0017] The power data analysis method based on the AI model as described above, wherein, for the data marked as having security threats by the AI model, the power data analysis system conducts in-depth analysis and calculates the risk score that comprehensively considers the degree of anomaly, duration and scope of influence, including the following sub-steps:
[0018] Extract key power parameters from the data marked as having security threats by the AI model, calculate the degree of deviation, and obtain the quantified score of the degree of anomaly according to the weight;
[0019] Determine the duration of the abnormal data, divide it into intervals and assign scores, analyze the scope of influence to obtain scores, and synthesize them to provide a basis for the subsequent risk score;
[0020] Sort out and summarize the quantified scores of the degree of anomaly, duration and scope of influence respectively;
[0021] Perform weighted summation on the scores of the degree of anomaly, duration and scope of influence by setting weights to obtain the final risk score.
[0022] The power data analysis method based on the AI model as described above, wherein, according to the risk score, the system automatically triggers corresponding response mechanisms, including the following sub-steps:
[0023] Divide the risk score into different levels to correspond to the requirements of different response mechanisms;
[0024] Trigger corresponding alarms and handling measures according to different risk levels of low, medium and high.
[0025] The power data analysis method based on the AI model as described above, wherein, the power data analysis system continuously updates the learning strategy of the AI model by evaluating the effects of processed events, including the following sub-steps:
[0026] Collect relevant data on processed events, quantitatively evaluate the effectiveness of event processing according to evaluation metrics, and analyze the reasons for poor effectiveness;
[0027] According to the analysis results, adjust the model training parameters or optimize the response mechanism.
[0028] A power data analysis system based on an AI model, including:
[0029] Data acquisition and preprocessing module: The power data analysis system collects and preprocesses data from grid monitoring points and calculates the baseline value of data stability;
[0030] Model training and behavior recognition module: Based on the baseline value of data stability, the power data analysis system trains an AI model to identify behaviors deviating from the normal range and judge potential security threats;
[0031] In-depth analysis and risk assessment module: For data marked as having security threats by the AI model, the power data analysis system conducts in-depth analysis and calculates a risk score that comprehensively considers the degree of abnormality, duration, and scope of influence;
[0032] Response mechanism trigger module: According to the risk score, the system automatically triggers the corresponding response mechanism;
[0033] Model optimization and learning strategy update module: The power data analysis system continuously updates the learning strategy of the AI model by evaluating the effectiveness of processed events.
[0034] A computer storage medium, characterized by including: at least one memory and at least one processor;
[0035] The memory is used to store one or more program instructions;
[0036] The processor is used to run one or more program instructions to execute the power data analysis method based on the AI model described in any one of the above.
[0037] The beneficial effects achieved by the present invention are as follows:
[0038] By introducing advanced AI technology into power data analysis, the safety and efficiency of power grid operation have been significantly improved. First of all, this method can quickly and accurately identify potential security threats. Compared with traditional manual monitoring methods, it not only improves the detection speed but also greatly reduces the false alarm rate and missed alarm rate, ensuring the effectiveness of power grid security monitoring. Secondly, based on a multi-dimensional risk assessment mechanism, the system can provide specific risk scores for each abnormal situation, which helps decision-makers take the most appropriate countermeasures according to the actual situation, thus more effectively controlling the development of the situation and reducing losses caused by emergencies. Moreover, through continuous learning and self-optimization processes, the AI model continuously adapts to new data patterns and threat types, ensuring stability and reliability in long-term applications. This method promotes the rational allocation and utilization of resources, reduces unnecessary maintenance costs, and is of great significance for promoting the intelligent transformation of the power industry. In summary, the present invention not only enhances the safety of the power system but also promotes the development of the entire industry towards a more efficient and intelligent direction. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0040] Figure 1 It is a flowchart of a power data analysis method based on an AI model provided by an embodiment of the present application.
[0041] Figure 2 It is a schematic diagram of a power data analysis system based on an AI model provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0043] Embodiment 1
[0044] As Figure 1 shown, the power data analysis method based on an AI model in the embodiment of the present application includes:
[0045] Step S10: The power data analysis system collects and preprocesses the data of power grid monitoring points, and calculates the data stability baseline value, which specifically includes the following sub-steps:
[0046] Step S11: Collect power operation data from grid monitoring points and store it in a database;
[0047] Through specific data acquisition devices or sensors, power operation data is obtained in real time at grid monitoring points. These data include various parameters such as voltage, current, and power. The collected data is transmitted in a specific data format and then stored in a dedicated database system. This database system needs to have efficient data storage and retrieval capabilities to ensure data integrity and availability. During the storage process, data security and backup strategies also need to be considered to prevent data loss or tampering.
[0048] Step S12: Clean and standardize the collected data;
[0049] Since the data collected from grid monitoring points may contain noise, outliers, or inconsistencies, it needs to be cleaned and standardized. The cleaning process includes removing obvious outliers, such as overly high or low values, which are caused by sensor failures or data transmission errors. Then, the data is standardized to have a unified format and unit for subsequent analysis and calculation. Standardization can be achieved by converting the data into a specific range or adopting a unified unit system.
[0050] Step S13: Determine the stability baseline value by calculating the mean and standard deviation of historical data;
[0051] Extract historical data from the database. This historical data has a certain representativeness and time span. Through statistical analysis of the historical data, the mean and standard deviation are calculated. The mean represents the central tendency of the data, and the mean is calculated using the following formula:
[0052]
[0053] where, represents the mean; a is a correction factor with a value range between 0 and 1. It is used to adjust the proportion of the weighted average in the final result, reflecting the relative importance attached to weighted average and geometric average; n is the number of data points, representing the total number of data points in the data set participating in the calculation; ω i is the weight of the i-th data point, usually satisfying and 0 ≤ ω i ≤ 1; x i is the i-th data point, representing each data value in the original data set participating in the calculation; is the n-th root of the geometric mean, reflecting the average growth or proportional relationship of the data.
[0054] The standard deviation reflects the degree of dispersion of the data and is calculated using the following formula:
[0055]
[0056] Among them, σ is the standard deviation; k represents the number of data groups; i is an index variable used to represent the group involved in the current calculation; f i represents the frequency of the data in the i-th group, that is, the number of times the data in the i-th group appears; x i is the midpoint of the i-th group of data. In grouped data, each group is usually represented by a representative value, and this representative value is the midpoint of the group; represents the weighted average, which is the average calculated by considering the frequencies of each group of data.
[0057] The stability baseline value can be determined according to the mean and standard deviation. The upper and lower limits of stability can be set by adding and subtracting a certain multiple of the standard deviation from the mean, so as to determine a reasonable range. When the newly collected data exceeds this range, it is considered that the power grid operation has an abnormal situation and further analysis and processing are required.
[0058] Suppose n key features x1, x2,..., x n have been determined, and each feature has its corresponding weight ω1, ω2,..., ω n . These weights reflect the importance of each feature to the stability of the power system. In addition, a set of eigenvalue c1, c2,..., c n under the ideal stable state is defined as the reference point. The data stability baseline value S can be calculated by the following formula:
[0059]
[0060] Among them, S is the data stability baseline value, which is used to measure the gap between the current data point and the ideal stable state; n is the number of features, representing the key features used to calculate the stability score; i represents the i-th feature; ω i represents the weight of the i-th feature. The larger the weight, the greater the influence of the feature on the stability score; x i represents the actual observed value of the i-th feature, which is the standardized feature value collected from the power grid monitoring point and preprocessed; c i represents the ideal stable state value of the i-th feature, which is the value that the feature should reach under the ideal stable state; α represents the weight coefficient of the cross term, which is used to adjust the influence degree of the cross term on the stability score; ω ij represents the interaction weight between the i-th feature and the j-th feature, which reflects the interaction strength between these two features. If ω ij is larger, it means that the interaction between these two features has a greater influence on the stability score; x jRepresents the actual observed value of the j-th feature; c j Represents the ideal stable state value of the j-th feature.
[0061] S20. Based on the data stability baseline values, the power data analysis system trains an AI model to identify behaviors deviating from the normal range and judge potential security threats, specifically including the following sub-steps:
[0062] Step S21. Divide the preprocessed historical data into a training set, a validation set, and a test set, and extract relevant features to form feature vectors;
[0063] This step aims to prepare for model training and evaluation. First, randomly divide the preprocessed historical data into a training set, a validation set, and a test set according to a certain proportion. Usually, the training set accounts for a relatively large proportion and is used for the model to learn parameters; the validation set is used to monitor the model performance during training and adjust hyperparameters; the test set is independent of the training and validation processes and is used to finally evaluate the generalization ability of the model. Then, extract relevant features from the data and form them into feature vectors to transform the original data into a more informative and interpretable form, so that the model can better understand and process the data.
[0064] Step S22. Select a suitable AI model architecture, train it using the training set, adjust the parameters through the validation set to prevent overfitting, and finally evaluate the model performance using the test set;
[0065] First, according to the characteristics of power data and the problem requirements, comprehensively consider factors such as computational complexity, interpretability, and training time, and select a suitable AI model architecture. Then, train the model using the training set, set appropriate hyperparameters, and adjust the model parameters by continuously optimizing the loss function to enable the model to gradually learn the patterns and rules in the data. During the training process, regularly evaluate the model performance using the validation set. Once it is found that the performance of the model on the validation set decreases, it indicates that overfitting may occur. At this time, measures such as increasing the data volume and using regularization techniques can be taken to adjust the hyperparameters to prevent overfitting. Finally, evaluate the trained model using an independent test set, and measure the generalization ability and practical application value of the model through evaluation metrics such as accuracy, recall rate, F1 value, and mean squared error. If the model performance does not meet the requirements, it is necessary to reconsider the model architecture, adjust the parameters, or perform more data preprocessing and feature engineering.
[0066] Step S30. For the data marked as having security threats by the AI model, the power data analysis system conducts in-depth analysis and calculates the risk score that comprehensively considers the degree of abnormality, duration, and scope of influence, specifically including the following sub-steps:
[0067] Step S31. Extract key power parameters from the data marked as having security threats by the AI model, calculate the degree of deviation, and obtain the quantitative score of the degree of abnormality according to the weights;
[0068] First, determining the key power parameters is an important start to this step. Key parameters can be selected by analyzing the parameters in historical data that have a greater impact on the safety of the power system or by combining expert experience. Then, for each key parameter, calculate the degree of deviation from its normal range. Suppose there are n key parameters, namely x1, x2,..., x n , and the corresponding data stability baseline values are S1, S2,..., S n , with weights ω1, ω2,..., ω n .
[0069] First, calculate the degree of deviation for each parameter:
[0070]
[0071] where d i represents the quantified value of the degree of deviation of the i-th key power parameter; x i represents the actual measured value of the i-th key power parameter; S i represents the baseline value of the i-th key power parameter.
[0072] Then, calculate the quantified score of the degree of abnormality:
[0073]
[0074] where E i represents the quantified score of the degree of abnormality; n represents the number of key power parameters; i is an index variable used to iterate through all key power parameters, from 1 to n; ω i represents the weight of the i-th key power parameter; S i represents the baseline value of the i-th key power parameter.
[0075] Step S32: Determine the duration of abnormal data, divide it into intervals and assign scores, analyze the scope of influence to obtain a score, and synthesize it to provide a basis for subsequent risk scoring;
[0076] When determining the duration of abnormal data, obtain the start and end times from the timestamps of the data, and then calculate the duration. After that, divide the duration into different intervals and set corresponding scores for each interval. The interval division can be based on experience or data analysis. For example, it can be divided into several time period intervals with corresponding different scores. At the same time, analyze the scope of influence of the abnormal data, which may include the number of affected devices, the area of the region, etc. Similarly, divide the scope of influence into different intervals and assign scores.
[0077] Step S33: Organize and summarize the quantified scores of the degree of abnormality, duration, and scope of influence respectively;
[0078] Sort out the quantification scores of the anomaly degree, duration scores, and influence scope scores respectively to prepare for the next risk scoring, ensuring that each score accurately corresponds to its corresponding indicator.
[0079] Step S34: Perform a weighted sum of the scores of the anomaly degree, duration, and influence scope by setting weights to obtain the final risk score.
[0080] The system counts the duration of the abnormal data. The duration refers to the length of time elapsed from the moment the abnormal data appears to the moment it disappears. A longer duration may indicate a higher severity of the problem. At the same time, the system will evaluate the possible scope affected by the abnormal data; it also includes the area of the affected region, whether it is a problem in a local area or a larger area is affected. Then, by synthesizing these two important aspects of the duration and influence scope of the abnormal data, the following formula is used to calculate the risk score:
[0081]
[0082] Where R is the risk score; a is the overall coefficient, which is used to adjust the size of the calculation result of the entire formula. According to the operating conditions of the actual power system and the experience of data analysis, it can be set to an appropriate value to ensure that the risk score is within a reasonable range; T is the duration of the abnormal data; γ is the exponent of the duration T in the trigonometric function part. Different γ values will affect the contribution of this part to the risk score; θ is the angle parameter in the trigonometric function sin(θT), which is combined with the duration T to adjust the period and phase of the trigonometric function, thereby affecting the fluctuation characteristics of this part in the risk score calculation; δ is the exponent of the duration T in the logarithmic function part; b is the base of the logarithmic function log b (T). Different bases will change the growth characteristics of the logarithmic function, thereby affecting the role of this part in the risk score; S is the quantification index of the influence scope of the abnormal data; ε is the exponent of the influence scope S in the trigonometric function part; φ is the angle parameter in the trigonometric function cos(φS); is the exponent of the influence scope S in the square root function part.
[0083] Step S40: According to the risk score, the system automatically triggers the corresponding response mechanism;
[0084] Based on the risk score calculated in the previous steps, automatically trigger the corresponding response mechanism, aiming to be able to respond to potential security threats in a timely and effective manner, control the adverse effects brought by security risks within the minimum range, and ensure the safety, stability, and reliability of the power grid operation. Specifically, it includes the following sub-steps:
[0085] Step S41: Divide the risk score into different levels to correspond to the requirements of different response mechanisms;
[0086] To determine the grading criteria for the risk score, analyze a large amount of historical data from the past, observe the actual situations and consequences under different risk scores. At the same time, experts in the power field can be invited to jointly determine reasonable grading criteria based on their experience. For example, if the risk score is divided into three levels: low, medium, and high, factors such as the numerical range of the risk score, the tolerance of the power system, and the possible impact degree can be considered. Suppose it is determined to divide by a specific numerical interval. For example, a risk score in the range of [0, 30] is considered a low-risk level, which may mean that the current abnormal situation is relatively mild and has little impact on the power system. After determining the grading criteria, a program can be written to automatically judge which level the input risk score belongs to, so as to trigger the corresponding response mechanism in the follow-up.
[0087] Step S42: Trigger corresponding alarm and handling measures according to different risk levels of low, medium, and high.
[0088] For the low-risk level, the alarm method is relatively mild. For example, display a less conspicuous prompt message on the interface of the power data analysis system, or send a normal email to notify the relevant operation and maintenance personnel. In terms of handling measures, mainly continuously monitor the abnormal situation and record relevant data changes for subsequent analysis. Regular inspection and maintenance work can be arranged to ensure that the problem will not deteriorate further. For example, check the relevant equipment at regular intervals to see if there are potential problems.
[0089] For the medium-risk level, the alarm method is more obvious. Flashing indicator lights, sound alarms, etc. can be issued to attract the attention of relevant personnel. At the same time, send an emergency notice to relevant personnel to ensure that they can pay attention to the problem in a timely manner. In terms of handling measures, immediately start the emergency plan, organize professional personnel to conduct in-depth analysis of the abnormal situation to determine the root cause of the problem. For example, conduct detailed detection and diagnosis of relevant equipment to find possible fault points. At the same time, isolate the affected area to prevent the problem from spreading. Some relevant equipment can be shut down, or the operating parameters of the system can be adjusted to reduce the risk.
[0090] For the high-risk level, the alarm method is extremely strong. Issue continuous alarm sounds, red flashing indicator lights, etc., and at the same time send an emergency notice to all relevant personnel, including management and technical experts. In terms of handling measures, immediately stop the operation of the affected equipment or system to prevent accidents from occurring. Organize a professional emergency rescue team to conduct on-site disposal and take emergency measures to reduce the risk. For example, conduct emergency repairs, replace faulty equipment, etc. At the same time, start the accident investigation procedure, find out the cause of the problem, and formulate corresponding rectification measures to prevent similar problems from occurring again.
[0091] Step S50: Continuously optimize the flight path of the aircraft according to the adjusted control instructions, specifically including the following sub-steps:
[0092] Step S51: Collect relevant data of the processed events, quantitatively evaluate the event processing effect according to the evaluation indicators, and analyze the reasons for the poor effect;
[0093] The system needs to comprehensively collect relevant data of the processed events. This includes information such as the power data characteristics at the time of the event occurrence, the taken processing measures, and the state of the power system after processing. For example, for an event that triggers a security threat due to abnormal voltage, the collected information can include key parameters such as the voltage value, current value, power value, etc. when the abnormality occurs, as well as the specific measures taken to adjust the voltage and whether the voltage returns to normal after processing. Then, quantitatively evaluate the event processing effect according to the pre-set evaluation indicators. The evaluation indicators can include the recovery time, the degree of improvement in system stability, etc. For example, calculate the time taken from the event occurrence to the system returning to normal operation. The shorter the time, the better the processing effect. At the same time, analyze the reasons for the poor processing effect. This may involve multiple aspects, such as inaccurate data collection leading to incorrect model judgment, inappropriate processing measures, limitations of the model itself, etc. If it is a data collection problem, it may be necessary to check the accuracy and stability of the sensors; if it is an inappropriate processing measure, it may be necessary to re-examine the rationality of the response mechanism; if it is a problem with the model itself, it may be necessary to further analyze whether the architecture and parameter settings of the model are reasonable.
[0094] Step S52: Adjust the model training parameters or optimize the response mechanism according to the analysis results;
[0095] Based on the analysis results of the reasons for the poor effect, take corresponding adjustment measures. If it is found that the model training parameters are unreasonable, such as too high or too low learning rate, inappropriate batch size, etc., these parameters can be adjusted. For example, if the high learning rate leads to unstable model training, the learning rate can be appropriately reduced. If there is a problem with the response mechanism, such as untimely alarm, ineffective processing measures, etc., the response mechanism can be optimized. For example, increase the priority of the alarm to ensure that important events can be promptly attended to; or improve the processing measures to make them more targeted and efficient. At the same time, after the adjustment, the new model and response mechanism need to be tested and verified to ensure that they can better handle similar events and improve the security and stability of the power system.
[0096] Embodiment 2
[0097] As Figure 2 shown, Embodiment 2 of the present application provides a power data analysis system based on an AI model, including:
[0098] Data Acquisition and Preprocessing Module 21: The power data analysis system collects and preprocesses data from grid monitoring points, and calculates the data stability baseline value, including the following sub-modules:
[0099] Data Acquisition Sub-module 211: Collect power operation data from grid monitoring points and store it in the database; Data acquisition is the basic part of the entire data analysis system. It is responsible for obtaining raw power operation data from various data sources and transmitting this data to the database for storage. For example, sensors at grid monitoring points will obtain power parameters such as voltage, current, and power in real-time or at regular intervals. The data acquisition module is like a data porter, storing the data collected by these sensors in the database according to certain rules and formats, laying the foundation for subsequent data processing.
[0100] Data Preprocessing Sub-module 212: Clean and standardize the collected data; In the actual data collection process, data may have various problems, such as missing data, data errors, inconsistent data formats, etc. The data preprocessing module is to solve these problems. Cleaning data can remove invalid data records, such as deleting data rows containing error values or too many missing values. Standardizing data is to make data from different monitoring points or different types comparable. For example, unifying voltage data to a reasonable range or performing conversions according to certain standards to ensure that subsequent data analysis and calculations can be carried out on the basis of higher-quality data.
[0101] Data Analysis Sub-module 213: Determine the stability baseline value by calculating the mean and standard deviation of historical data; After data acquisition and preprocessing, relatively clean and standard data can be analyzed. Calculating the mean and standard deviation of historical data to determine the stability baseline value is a common data analysis method.
[0102] Extract historical data from the database. These historical data have a certain representativeness and time span. Through statistical analysis of the historical data, the mean and standard deviation are calculated. The mean represents the central tendency of the data, and the mean is calculated by the following formula:
[0103]
[0104] Where, represents the mean; a is a correction factor with a value range between 0 and 1. It is used to adjust the proportion of the weighted average in the final result, reflecting the relative emphasis on weighted average and geometric average; n is the number of data points, representing the total number of data points in the data set participating in the calculation; ω i is the weight of the i-th data point, usually satisfying and 0 ≤ ω i ≤ 1; x iis the i-th data point, representing each data value in the original dataset involved in the calculation; is the n-th root of the geometric mean, reflecting the average growth or proportional relationship of the data.
[0105] The standard deviation reflects the degree of dispersion of the data. The standard deviation is calculated using the following formula:
[0106]
[0107] where σ is the standard deviation; k represents the number of data groups; i is an index variable used to represent the group involved in the current calculation; f i represents the frequency of the i-th group of data, that is, the number of times the i-th group of data appears; x i is the midpoint of the i-th group of data. In grouped data, each group is usually represented by a representative value, and this representative value is the midpoint; represents the weighted average, which is the average calculated by considering the frequencies of each group of data.
[0108] The stability baseline value can be determined based on the mean and standard deviation. The upper and lower limits of stability can be set by adding and subtracting a certain multiple of the standard deviation from the mean. In this way, a reasonable range can be determined. When the newly collected data exceeds this range, it is considered that the power grid operation has an abnormal situation and further analysis and processing are required.
[0109] Suppose n key features x1, x2,..., x n have been determined, and each feature has its corresponding weight ω1, ω2,..., ω n , these weights reflect the importance of each feature for the stability of the power system. In addition, a set of eigenvalue c1, c2,..., c n under the ideal stable state is defined as the reference point. The data stability baseline value S can be calculated by the following formula:
[0110]
[0111] where S is the data stability baseline value, used to measure the gap between the current data point and the ideal stable state; n is the number of features, representing the key features used to calculate the stability score; i represents the i-th feature; ω i represents the weight of the i-th feature. The greater the weight, the greater the impact of the feature on the stability score; x i represents the actual observed value of the i-th feature, which is the standardized feature value collected from the power grid monitoring point and preprocessed; c iRepresents the ideal stable state value of the i-th feature, which is the value that the feature should reach in the ideal stable state; α represents the weight coefficient of the cross-term, used to adjust the influence degree of the cross-term on the stability score; ω ij Represents the interaction weight between the i-th feature and the j-th feature, reflecting the interaction strength between these two features. If ω ij Is larger, it indicates that the interaction between these two features has a greater impact on the stability score; x j Represents the actual observed value of the j-th feature; c j Represents the ideal stable state value of the j-th feature.
[0112] Based on these values, a stability baseline range is set to determine whether the subsequently real-time collected data is stable. This is a specific operation in the data analysis module for determining the data stability judgment criterion.
[0113] Model Training and Behavior Recognition Module 22: Based on the data stability baseline values, the power data analysis system trains an AI model to identify behaviors deviating from the normal range and judge potential security threats, including the following sub-modules:
[0114] Data Partitioning and Feature Extraction Sub-module 221: Divide the preprocessed historical data into a training set, a validation set, and a test set, and extract relevant features to form a feature vector; before using the data for AI model training, it is necessary to first reasonably partition the preprocessed data. Partitioning the data sets for different purposes is an important prerequisite for ensuring the scientific and reasonable development of model training. The training set is used to let the model learn the rules and patterns in the data, the validation set is used to adjust the model's parameters during training to prevent overfitting and other situations, and the test set is used to objectively evaluate the final performance of the model after the model training is completed. At the same time, extracting relevant features to form a feature vector is also a key task of this module because the original data often contains a large amount of redundant information or not all information is helpful for the model's judgment. Through feature extraction, it is possible to focus on those feature dimensions that are truly valuable for judging behaviors deviating from the normal range and identifying potential security threats, integrate them into a feature vector and input it into the subsequent model to facilitate the model to learn and judge more efficiently.
[0115] AI Model Training and Evaluation Sub-module 222: Select a suitable AI model architecture, train it using the training set, adjust the parameters through the validation set to prevent overfitting, and finally evaluate the model performance using the test set; this sub-module focuses on the key aspects of the entire life cycle of the AI model. First, a suitable model architecture should be selected based on specific tasks and data characteristics factors. Different types of model architectures such as neural networks and support vector machines can be chosen. After selecting the architecture, the prepared training set is used to let the model learn the features and patterns in the data. During the training process, the parameters of the model are continuously adjusted and optimized with the help of the validation set to avoid the situation where the model performs well on the training set but poorly on new data. Finally, an independent test set is used to evaluate the performance of the trained model to see if it can accurately identify behaviors that deviate from the normal range and effectively judge potential security threats, thereby determining whether the model meets the expected practical standards.
[0116] Deep Analysis and Risk Assessment Module 23: For the data marked with security threats by the AI model, the power data analysis system conducts in-depth analysis and calculates the risk scores that comprehensively consider the degree of abnormality, duration, and scope of influence, including the following sub-modules:
[0117] Degree of Abnormality Quantification Sub-module 231: Extract key power parameters from the data marked with security threats by the AI model, calculate the degree of deviation, and obtain the quantified score of the degree of abnormality according to the weights; the key focus is on deeply analyzing the data that has been marked as having security threats by the AI model, extracting key power parameters such as voltage, current, power, etc. from these data, and then accurately calculating their degree of deviation from the normal range. And different weights are assigned according to the importance of each parameter for the overall security threat judgment, and finally a score that can quantitatively reflect the degree of abnormality is obtained comprehensively.
[0118] First of all, determining the key power parameters is an important start of this step. The key parameters can be selected by analyzing the parameters that have a greater impact on the power system security in historical data or combining expert experience. Then, for each key parameter, calculate its degree of deviation from the normal range. Suppose there are n key parameters, which are x1, x2,..., x n , and the corresponding data stability baseline values are S1, S2,..., S n , and the weights are ω1, ω2,..., ω n .
[0119] First, calculate the degree of deviation of each parameter:
[0120]
[0121] Among them, d i represents the quantified value of the degree of deviation of the i-th key power parameter; x irepresents the actual measured value of the i-th key power parameter; S i represents the baseline value of the i-th key power parameter.
[0122] Then, calculate the quantification score of the anomaly degree:
[0123]
[0124] where, E i represents the quantification score of the anomaly degree; n represents the number of key power parameters; i is an index variable used to traverse all key power parameters, from 1 to n; ω i represents the weight of the i-th key power parameter; S i represents the baseline value of the i-th key power parameter.
[0125] Time and scope analysis sub-module 232: Determine the duration of the abnormal data and divide it into intervals for scoring, analyze the influence scope to obtain a score, and synthesize it to provide a basis for subsequent risk scoring; This module focuses on two important dimensions. One is the duration of the abnormal data. By analyzing relevant information such as the data time series, it is clear how long the abnormal situation lasts, and then scores are assigned according to the pre-set time interval standard. For example, if the abnormal duration is short, the score may be low, and if the duration is long, the score may be high. The other is the analysis of the influence scope of the abnormal data, considering whether the anomaly only affects local power grid monitoring points or involves a large-scale power grid area, etc., and giving corresponding scores accordingly. Through the analysis and scoring of these two dimensions, judgment bases are further accumulated for the overall risk scoring from different perspectives, comprehensively reflecting the impact of security threats in terms of time and space scope.
[0126] Score sorting sub-module 233: Sort and summarize the quantification scores of the anomaly degree, duration, and influence scope respectively; This module plays a connecting role from top to bottom. It mainly sorts and summarizes the anomaly degree score, duration score, and influence scope score calculated separately before, integrates the scattered scores of each dimension together, facilitates subsequent unified comprehensive calculations, ensures the accuracy and integrity of each score data, provides a clear data basis for operations such as final weighted summation, and avoids problems such as data chaos or omission affecting the accuracy of risk scoring.
[0127] Risk scoring calculation sub-module 234: Obtain the final risk score by weighted summation of the scores of the anomaly degree, duration, and influence scope by setting weights; Based on the quantification scores of the anomaly degree, duration, and influence scope obtained by each previous sub-module, and according to the scientifically and reasonably set weights, the scores of each dimension are integrated, and finally a risk score that can comprehensively and synthetically reflect the security threat situation represented by the data is calculated.
[0128] The following formula is used to calculate the risk score:
[0129]
[0130] where R is the risk score; a is the overall coefficient, which is used to adjust the magnitude of the calculation result of the entire formula. According to the operating conditions of the actual power system and the experience of data analysis, it can be set to an appropriate value to ensure that the risk score is within a reasonable range; T is the duration of abnormal data; γ is the exponent of the duration T in the trigonometric function part, and different γ values will affect the contribution of this part to the risk score; θ is the angle parameter in the trigonometric function sin(θT), which is combined with the duration T to adjust the period and phase of the trigonometric function, thereby affecting the fluctuation characteristics of this part in the risk score calculation; δ is the exponent of the duration T in the logarithmic function part; b is the base of the logarithmic function log b (T). Different bases will change the growth characteristics of the logarithmic function, thereby affecting the role of this part in the risk score; S is the quantification index of the influence range of abnormal data; ε is the exponent of the influence range S in the trigonometric function part; φ is the angle parameter in the trigonometric function cos(φS); is the exponent of the influence range S in the square root function part.
[0131] Response mechanism trigger module 24: Continuously monitor the actual trajectory of the aircraft, calculate the trajectory deviation value by comparing the actual trajectory with the planned trajectory, and adjust the control instructions, including the following sub-modules:
[0132] Risk level division sub-module 241: Divide the risk score into different levels to correspond to different response mechanism requirements; by setting clear boundary values, divide the risk score into different levels such as low risk, medium risk, high risk, etc. This division process needs to comprehensively consider factors such as the tolerance of the power system, safety thresholds, and past response experience. For example, a risk score of 0 - 30 is the low risk level, 31 - 70 is the medium risk level, and 71 - 100 is the high risk level.
[0133] Response trigger sub-module 242: Trigger corresponding alarm and processing measures according to different risk levels of low, medium, and high; for the low risk level, it may only trigger a simple alarm message, display a yellow warning sign on the system interface, and record relevant data for subsequent manual viewing; for the medium risk level, in addition to the alarm, some preliminary processing measures will be automatically started, adjust the operating parameters of some power equipment, and notify relevant operation and maintenance personnel to pay close attention; for the high risk level, a strong alarm signal will be triggered, a red alarm sound will be issued, an emergency notice will be pushed to operation and maintenance and management personnel, and a series of emergency processing measures will be automatically executed, such as cutting off some lines and starting standby equipment, to minimize the harm of potential safety threats to the power system.
[0134] Model optimization and learning strategy update module 25: Continuously optimize the aircraft trajectory according to the adjusted control instructions, including the following sub-modules:
[0135] Event effect evaluation sub-module 251: Collect relevant data of processed events, quantitatively evaluate the event processing effect according to evaluation indicators, and analyze the reasons for poor effects; First, collect data closely related to these processed events, which cover various aspects such as the original power operation data, risk score situation, triggered response measures, and the final actual results when the events occur. Then, quantitatively evaluate the event processing effect according to the pre-set evaluation indicators to obtain a more objective and accurate evaluation result. And further deeply analyze in those events with poor processing effects whether it is due to inaccurate risk judgment in the early stage, ineffective execution of the response mechanism, or other factors, so as to provide a clear direction for subsequent improvement.
[0136] Model and mechanism optimization sub-module 252: Adjust model training parameters or optimize the response mechanism according to the analysis results; This sub-module is based on the event effect evaluation module and targets the optimization of key parts in the system according to the results analyzed by it. If it is found through analysis that there are biases in the AI model's risk judgment, resulting in poor processing effects, relevant parameters of model training will be adjusted, such as changing the weights of data features, optimizing the hyperparameters of the model, etc., so that the model can more accurately identify security threats in subsequent learning and application processes. If it is found that there are unreasonable aspects in the response mechanism itself, such as untimely alarms, low execution efficiency of processing measures, etc., the response mechanism will be optimized, and the alarm methods under different risk levels will be re-set, and the specific operation process of the processing measures will be improved.
[0137] Corresponding to the above embodiments, an embodiment of the present invention provides a computer storage medium, including: at least one memory and at least one processor;
[0138] The memory is used to store one or more program instructions;
[0139] The processor is used to run one or more program instructions to execute the power data analysis method and system based on the AI model;
[0140] Corresponding to the above embodiments, an embodiment of the present invention provides a computer-readable storage medium, and the computer storage medium contains one or more program instructions, and the one or more program instructions are used to be executed by the processor to execute the power data analysis method and system based on the AI model.
[0141] The embodiments disclosed by the present invention provide a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions run on a computer, the computer is enabled to execute the above-mentioned power data analysis method and system based on an AI model.
[0142] In the embodiments of the present invention, the processor may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0143] It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, a mature storage medium in the art. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.
[0144] The storage medium may be a memory, for example, it may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories.
[0145] Among them, the non-volatile memory may be a read-only memory (ROM for short), a programmable read-only memory (PROM for short), an erasable programmable read-only memory (EPROM for short), an electrically erasable programmable read-only memory (EEPROM for short), or a flash memory.
[0146] The volatile memory may be a Random Access Memory (RAM) which serves as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM).
[0147] The storage media described in the embodiments of the present invention are intended to include but not limited to these and any other suitable types of memories.
[0148] Those skilled in the art should be aware that, in one or more of the above examples, the functions described in the present invention can be implemented by a combination of hardware and software. When applying software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium that can be accessed by a general or special purpose computer.
[0149] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only for the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made on the basis of the technical solutions of the present invention should be included in the protection scope of the present invention.
Claims
1. The power data analysis method based on AI model is characterized by: include: S10, the power data analysis system collects and pre-processes the data from the power grid monitoring points and calculates the data stability baseline value; S20. Based on the data stability baseline value, the power data analysis system trains the AI model to identify behaviors that deviate from the normal range and determine potential security threats; S30. For data marked as security threats by the AI model, the power data analysis system conducts in-depth analysis and calculates a risk score based on the degree of abnormality, duration, and scope of impact; S40. Based on the risk score, the system automatically triggers the corresponding response mechanism; S50. The power data analysis system continuously updates the learning strategy of the AI model by evaluating the effects of processed events.
2. The power data analysis method based on the AI model according to claim 1 is characterized in that: Based on the data stability baseline value, the power data analysis system trains the AI model, including the following sub-steps: Collect power operation data from power grid monitoring points and store them in the database; Clean and standardize the collected data; The stability baseline value is determined by calculating the mean and standard deviation of historical data.
3. The power data analysis method based on the AI model according to claim 1 is characterized in that: Based on the data stability baseline value, the power data analysis system trains the AI model to identify behaviors that deviate from the normal range and judge potential security threats, including the following sub-steps: Divide the preprocessed historical data into training set, validation set and test set, and extract relevant features to form feature vectors; Select a suitable AI model architecture, use the training set for training, adjust parameters through the validation set to prevent overfitting, and finally use the test set to evaluate model performance.
4. The power data analysis method based on the AI model according to claim 1 is characterized in that: For data that the AI model flags as a security threat, the power data analysis system conducts in-depth analysis, including the following sub-steps: Extract key power parameters from data marked as security threats by the AI model, calculate the degree of deviation, and derive a quantitative score of the degree of abnormality based on the weight; Determine the duration of abnormal data and divide it into intervals and assign scores, analyze the impact range to derive scores, and provide a comprehensive basis for subsequent risk scoring.
5. The power data analysis method based on the AI model according to claim 4 is characterized in that: Calculating the risk score based on the degree of abnormality, duration, and impact range includes the following sub-steps: Summarize the quantitative scores of the degree of abnormality, duration, and scope of impact; The final risk score is obtained by setting weights and taking the weighted sum of the scores of the abnormality degree, duration and scope of impact.
6. The power data analysis method based on the AI model according to claim 1 is characterized in that: Based on the risk score, the system automatically triggers the corresponding response mechanism, including the following sub-steps: Divide risk scores into different levels to correspond to different response mechanism requirements; Trigger corresponding alarms and handling measures according to different risk levels: low, medium and high.
7. The power data analysis method based on the AI model according to claim 1 is characterized in that: The power data analysis system continuously updates the learning strategy of the AI model by evaluating the effects of processed events, including the following sub-steps: Collect relevant data of processed events, quantify and evaluate the event processing effects based on evaluation indicators, and analyze the reasons for poor results; Based on the analysis results, adjust model training parameters or optimize response mechanisms.
8. The power data analysis system based on AI model is characterized by: include: Data collection and preprocessing module: The power data analysis system collects and preprocesses data from power grid monitoring points and calculates the data stability baseline value; Model training and behavior recognition module: Based on the data stability baseline value, the power data analysis system trains the AI model to identify behaviors that deviate from the normal range and determine potential security threats; In-depth analysis and risk assessment module: For data marked as security threats by the AI model, the power data analysis system conducts in-depth analysis and calculates risk scores based on the degree of abnormality, duration, and scope of impact; Response mechanism trigger module: Based on the risk score, the system automatically triggers the corresponding response mechanism; Model optimization and learning strategy update module: The power data analysis system continuously updates the learning strategy of the AI model by evaluating the effects of processed events.
9. A computer storage medium, characterized in that: include: at least one memory and at least one processor; A memory for storing one or more program instructions; A processor, used to run one or more program instructions to execute the power data analysis method based on the AI model as described in any one of claims 1-7.
Citation Information
Cited By
Electric power data safety monitoring method and system
CN120524165A