An artificial intelligence driven based power grid data hybrid security encryption method

By employing an AI-driven hybrid security encryption method based on SVM, which dynamically selects encryption algorithms and key strength, the problems of low efficiency and insufficient adaptability in smart meter data transmission are solved, achieving efficient and secure data transmission.

CN120750649BActive Publication Date: 2025-11-18INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511204273.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-18
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing smart meter data encryption methods are ill-suited to the diverse data traffic characteristics and dynamically changing network environments, resulting in low efficiency, insufficient adaptability, and excessive computational overhead, thus failing to effectively protect user privacy.

Method used

We adopt an AI-driven hybrid security encryption method based on SVM. Through a scenario classification model and a key strength decision model, we dynamically select encryption algorithms and key strengths to achieve hybrid encryption of symmetric and asymmetric encryption. We also combine traffic characteristics and encryption algorithm selection tags to optimize the encryption strategy.

Benefits of technology

It improves the adaptability of high-frequency, multi-device data transmission in smart meters, balances data security and encryption overhead, reduces the probability of data being acquired or forged, enhances anti-tampering capabilities during transmission, and solves the problems of complex key management and difficulty in balancing encryption speed and strength in traditional encryption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750649B_ABST
    Figure CN120750649B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of power grid data encryption, and particularly relates to a power grid data hybrid security encryption method based on artificial intelligence driving, which comprises: power grid data flow characteristics generate control flow labels, metering flow labels and encryption algorithm selection labels through a scene classification model, wherein the scene classification model is constructed based on a multi-output classification SVM framework; power grid data flow characteristics are combined with the encryption algorithm selection labels to input a key strength decision model to generate key recommended strength and flow abnormal values, wherein the key strength decision model is constructed based on a hybrid random forest framework; the control flow labels, the metering flow labels, the encryption algorithm selection labels and the key recommended strength are used to calculate comprehensive predicted encryption overhead, and an adaptive hybrid security encryption algorithm for this communication is set. The present application balances data security and encryption overhead through the SVM-based artificial intelligence driving selection algorithm, and provides reliable protection for power grid data security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid data encryption technology, and in particular to a hybrid security encryption method for power grid data based on artificial intelligence. Background Technology

[0002] As a crucial component of modern power systems, smart grids enable intelligent management and efficient operation of the power system through smart meters, communication networks, and data analytics. The widespread application of smart meters allows power grid companies to obtain real-time electricity consumption data from users, thereby achieving precise energy management and distribution optimization.

[0003] However, in this process, with the increasing frequency and scale of data uploaded by smart meters, the issue of protecting user electricity data has become increasingly prominent. This data not only includes users' electricity consumption habits but may also leak sensitive information such as their household routines and financial situation. Therefore, how to achieve efficient data transmission and processing while ensuring data privacy has become a critical issue that urgently needs to be addressed in the smart grid field.

[0004] Currently, various solutions exist for the security of data uploaded by smart meters, primarily including data encryption, privacy protection mechanisms, and security authentication technologies. However, some shortcomings remain. On the one hand, traditional encryption methods typically employ fixed encryption algorithms and key strengths, making them ill-suited to the diverse data traffic characteristics and dynamically changing network environment of smart grids. On the other hand, existing privacy protection schemes often overlook the computational and communication overhead of encryption algorithms during data processing, leading to inefficiencies and resource waste in practical applications.

[0005] Currently, existing solutions to the increasingly frequent and rapidly changing classification and encryption algorithm selection for smart meter uploaded data rely entirely on manual rules or relatively simple machine learning models. For example, port-based traffic classification methods struggle to effectively distinguish between different applications and services when faced with encrypted traffic, while traditional random forest or support vector machine (SVM) models lack the ability to jointly optimize cross-task features when handling multi-task classification problems, resulting in insufficient classification accuracy. Therefore, current methods for encrypting and protecting the privacy of smart meter uploaded data still suffer from inefficiency, insufficient adaptability, and excessive computational overhead, urgently requiring more intelligent, efficient, and flexible solutions.

[0006] Therefore, how to dynamically select the optimal encryption algorithm based on data traffic characteristics to balance data security and encryption overhead is a technical problem that needs to be solved. Summary of the Invention

[0007] To address this, the present invention provides an AI-driven hybrid security encryption method for power grid data. By employing an AI-driven selection algorithm based on SVM, the method intelligently selects the current encryption algorithm and key strength based on the characteristics of the data transmission volume of smart meters. This achieves hybrid encryption of symmetric and asymmetric encryption, improving the adaptability of hybrid encryption to the high-frequency, multi-device data transmission requirements of smart meters, balancing data security and encryption overhead, and providing reliable protection for power grid data security.

[0008] To achieve the above objectives, this invention proposes an artificial intelligence-driven hybrid security encryption method for power grid data, comprising:

[0009] The power grid data flow characteristics reported by smart meters are used to generate control flow labels, metering flow labels, and encryption algorithm selection labels through a scenario classification model. The scenario classification model is built on a multi-output classification SVM framework and trained through a joint loss function. The joint loss function is constructed by weighting the control flow labels based on the dynamic fault tolerance parameters of the control flow and the metering flow labels based on the dynamic fault tolerance parameters of the metering flow.

[0010] The power grid data flow characteristics are combined with the encryption algorithm to select labels and input them into the key strength decision model to generate key recommended strength and flow anomaly values, wherein the key strength decision model is constructed based on a hybrid random forest framework.

[0011] Based on the control flow label, metering flow label, encryption algorithm selection label, and key recommendation strength, a comprehensive prediction encryption cost is calculated. Based on the comprehensive prediction encryption cost, an adaptive hybrid security encryption algorithm is set for the smart meter and cloud server to perform this communication.

[0012] Furthermore, the process by which the scenario classification model generates control traffic labels, metering traffic labels, and encryption algorithm-selected labels includes:

[0013] The joint loss function is constructed based on the norm sum of squares of the control flow classification parameters, metering flow classification parameters, and encryption algorithm selection parameters; the first product of the control flow dynamic fault tolerance parameters, control flow accuracy weights, and overall control flow classification sample slack variables; the second product of the metering flow dynamic fault tolerance parameters, metering flow accuracy weights, and overall metering flow classification sample slack variables; and the third product of the selection weights and overall encryption algorithm selection sample slack variables.

[0014] Based on the binary classification interval constraint framework, control flow classification constraints and metering flow classification constraints are constructed respectively, and an encryption algorithm selection decision function is constructed based on the multi-classification strategy. The flow classification constraints, metering flow classification constraints, and encryption algorithm selection decision function all contain the power grid data flow characteristics.

[0015] During the training process of the scene classification model, the joint loss function is solved based on the traffic classification constraints, the metering traffic classification constraints, and the encryption algorithm selection decision function to generate control traffic labels, metering traffic labels, and encryption algorithm selection labels. Furthermore, during the solution process of the scene classification model, the control traffic dynamic fault tolerance parameters and the metering traffic dynamic fault tolerance parameters are adaptively and dynamically optimized based on the Bayesian sub-model.

[0016] Furthermore, the process of adaptively and dynamically optimizing the control flow dynamic fault-tolerant parameters and the metering flow dynamic fault-tolerant parameters based on the Bayesian sub-model includes:

[0017] The Bayesian sub-model searches the control flow parameter search space and the metering flow parameter search space based on hierarchical cross-validation, and outputs the control flow dynamic fault tolerance parameter and the metering flow dynamic fault tolerance parameter as the optimal parameter combination, wherein the numerical range of the control flow parameter search space is larger than the numerical range of the metering flow parameter search space.

[0018] Furthermore, the process of adaptively and dynamically optimizing the control flow dynamic fault tolerance parameters and the metering flow dynamic fault tolerance parameters based on the Bayesian sub-model also includes:

[0019] If the accuracy of the current scene classification model is less than the target accuracy threshold, then the optimal parameter combination is multiplied by a first coefficient greater than 1.

[0020] If the recall rate of the current scene classification model is less than the target recall rate threshold, then the optimal parameter combination is divided by a second coefficient greater than 1.

[0021] The first coefficient is greater than the second coefficient.

[0022] Furthermore, the calculation process for the control flow accuracy weight, the metering flow accuracy weight, and the selection weight includes:

[0023] During the training and optimization process of the scene classification model, the temporary weight of the control flow or the temporary weight of the metering flow is obtained based on the ratio of the exponential function of the prediction accuracy of the control flow or the metering flow to the exponential function of the overall prediction accuracy. The temporary weight of the metering flow is then self-adjusted to obtain the control flow accuracy weight, and the temporary weight of the metering flow is then self-adjusted to obtain the metering flow accuracy weight.

[0024] The selection weight is determined based on the control flow accuracy weight and the metering flow accuracy weight.

[0025] Furthermore, the process of generating control flow labels, metering flow labels, and encryption algorithm-selected labels using the scenario classification model includes:

[0026] A comprehensive kernel function is constructed by weighting the linear kernel of the control flow label, the radial basis kernel of the metering flow label, and the encrypted selection linear kernel.

[0027] The above solution balances the different security requirements of controlling and measuring traffic, effectively reduces the probability of data being acquired or forged, improves the anti-tampering capability during transmission, and dynamically adjusts the encryption strategy by combining traffic characteristics with encryption algorithm label selection. This solves the problems of complex key management and difficulty in balancing encryption speed and strength in traditional encryption. Furthermore, through dynamic label classification and fault tolerance parameter optimization, it significantly improves the adaptability of the encryption strategy to actual business.

[0028] Furthermore, the process of generating key recommendation strength and traffic outliers includes:

[0029] The power grid data flow characteristics are combined with the encryption algorithm to select tags and generate subtree key strength coefficients by selecting subtrees based on key strength calculated by minimizing mean square error regression.

[0030] The power grid data flow characteristics are used to generate the flow anomaly values ​​through an anomaly detection subtree constructed based on an unsupervised isolation tree architecture;

[0031] The subtree key strength coefficient is adjusted based on the power grid data flow characteristics and the flow anomalies to generate the recommended key strength.

[0032] The key strength decision model includes the key strength selection subtree and the anomaly detection subtree.

[0033] Furthermore, the process of generating the traffic anomaly value through the anomaly detection subtree includes:

[0034] Based on the power grid data flow characteristics, an unsupervised random forest is split to generate an anomaly detection subtree for quantifying the degree of flow anomaly.

[0035] The average path length of the isolated tree and the standard path length are calculated using the anomaly detection subtree;

[0036] The abnormal traffic value is calculated based on the ratio of the average path length of the isolated tree to the standard path length.

[0037] Furthermore, the process of adjusting the subtree key strength coefficients to generate the recommended key strength based on the power grid data flow characteristics and the flow anomalies includes:

[0038] Based on the comparison result between the traffic anomaly value and the first anomaly threshold, the subtree key strength coefficient is increased by a set level to generate the key recommended strength;

[0039] Based on the comparison result between the traffic anomaly value and the second anomaly threshold, the subtree key strength coefficient is increased to the maximum value corresponding to the device capability characteristics to generate the key recommendation strength;

[0040] Wherein, the first abnormal threshold is less than the second abnormal threshold, and the power grid data flow characteristics include equipment capability characteristics.

[0041] Furthermore, the process of calculating the comprehensive predicted encryption overhead and setting the adaptive hybrid security encryption algorithm based on the comprehensive predicted encryption overhead includes:

[0042] The encryption computational overhead is determined based on the selected label using the encryption algorithm, the recommended key strength, and the device performance factor.

[0043] The data inflation rate of the selected tag is based on the encryption algorithm, and the communication overhead is calculated based on the data packet size and latency requirements of the control flow tag or metering flow tag.

[0044] Calculate the security overhead based on the key recommendation strength and traffic anomalies;

[0045] The weighted summation of the encryption computation overhead, communication overhead, and security overhead generates a comprehensive predicted encryption overhead.

[0046] When the predicted encryption overhead exceeds the comprehensive threshold, the encryption algorithm selection label corresponding to the control traffic label is downgraded, the key recommendation strength corresponding to the metering traffic label is downgraded, and an adaptive hybrid security encryption algorithm is generated.

[0047] In the above scheme, the key strength decision model is based on a hybrid random forest framework. It dynamically adjusts the key strength by combining anomaly detection subtrees, which solves the contradiction between the difficulty of key management in traditional symmetric encryption and the slow speed of asymmetric encryption. It dynamically enhances the security level through anomaly detection, while avoiding the energy consumption and latency problems caused by over-encryption, which meets the requirements of smart grids for low energy consumption and high efficiency.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] 1. By using an AI-driven selection algorithm based on SVM, the current encryption algorithm and key strength are intelligently selected based on the characteristics of the data transmission volume of smart meters. This achieves hybrid encryption of symmetric and asymmetric encryption, improving the adaptability of hybrid encryption to the high-frequency, multi-device data transmission requirements of smart meters, balancing data security and encryption overhead, and providing reliable protection for power grid data security.

[0050] 2. It balances the different security requirements of controlling and measuring traffic, effectively reduces the probability of data being acquired or forged, improves the anti-tampering capability during transmission, and combines traffic characteristics with encryption algorithm to select tags, which can dynamically adjust the encryption strategy. It solves the problems of complex key management and difficulty in balancing encryption speed and strength in traditional encryption. Through dynamic tag classification and fault tolerance parameter optimization, it significantly improves the adaptability of encryption strategy to actual business.

[0051] 3. The key strength decision model is based on a hybrid random forest framework and dynamically adjusts the key strength by combining anomaly detection subtrees. This solves the contradiction between the difficulty of key management in traditional symmetric encryption and the slow speed of asymmetric encryption. It dynamically enhances the security level through anomaly detection, while avoiding the energy consumption and latency problems caused by over-encryption, which meets the requirements of smart grids for low energy consumption and high efficiency. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the AI-driven hybrid security encryption method for power grid data according to an embodiment of the present invention.

[0053] Figure 2 This is a flowchart illustrating the SVM scenario classification model of the AI-driven hybrid security encryption method for power grid data, as described in an embodiment of the present invention.

[0054] Figure 3 This is a flowchart illustrating the key strength decision model of the AI-driven hybrid security encryption method for power grid data according to an embodiment of the present invention.

[0055] Figure 4 This is a schematic diagram illustrating the process of determining the encryption algorithm and key strength in the AI-driven hybrid security encryption method for power grid data according to an embodiment of the present invention. Detailed Implementation

[0056] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0057] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0058] It should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0059] like Figures 1 to 4 As shown, this invention provides a hybrid security encryption method for power grid data driven by artificial intelligence. By using an AI-driven selection algorithm based on SVM, the method intelligently selects the current encryption algorithm and key strength based on the characteristics of the data transmission volume of smart meters, thereby achieving hybrid encryption of symmetric and asymmetric encryption. This improves the adaptability of hybrid encryption to the high-frequency, multi-device data transmission requirements of smart meters, balances data security and encryption overhead, and provides a reliable guarantee for power grid data security.

[0060] like Figures 1 to 4 As shown, this embodiment of the invention proposes a hybrid security encryption method for power grid data based on artificial intelligence, including:

[0061] The power grid data flow characteristics reported by smart meters are used to generate control flow labels, metering flow labels, and encryption algorithm selection labels through a scenario classification model. The scenario classification model is built on a multi-output classification SVM framework and trained through a joint loss function. The joint loss function is constructed by weighting the control flow labels based on the dynamic fault tolerance parameters of the control flow and the metering flow labels based on the dynamic fault tolerance parameters of the metering flow.

[0062] The power grid data flow characteristics are combined with the encryption algorithm to select labels and input them into the key strength decision model to generate key recommended strength and flow anomaly values, wherein the key strength decision model is constructed based on a hybrid random forest framework.

[0063] Based on the control flow label, metering flow label, encryption algorithm selection label, and key recommendation strength, a comprehensive prediction encryption cost is calculated. Based on the comprehensive prediction encryption cost, an adaptive hybrid security encryption algorithm is set for the smart meter and cloud server to perform this communication.

[0064] It is understood that this embodiment uses an SVM artificial intelligence-driven model to classify traffic characteristics in real time, and its training process uses smart meter communication data that conforms to the power safety standard IEC 62351.

[0065] Specifically, the power grid data flow characteristics reported by smart meters include tripping commands, frequency regulation signals, priority commands, user electricity consumption data, voltage standard deviation, current standard deviation, and equipment capability characteristics. These equipment capability characteristics include CPU and encryption hardware support, neither of which requires encrypted reporting to the cloud server; they are only used to assist in selecting encryption algorithms and key strength.

[0066] The data belonging to control commands includes trip commands, frequency modulation signals, and priority commands. Control commands are characterized by well-defined communication protocols, low latency, small data packets, and strong linear separability. Since they all involve power grid dispatching, real-time performance is crucial, requiring highly sensitive, low-latency (less than 10ms delay) encryption algorithms, such as AES-256 and ChaCha20-Poly. Specifically, trip commands, being emergency commands, require AES-256 encryption, while frequency modulation signals and priority commands, being regular commands, require ChaCha20-Poly encryption.

[0067] The data belonging to the metering instructions include user electricity consumption data, voltage standard deviation and current standard deviation. The metering instructions are characterized by large data packets and may contain long-term fluctuations and nonlinearity. Due to the sensitivity of the data, highly secure AES-128 and SPECK encryption algorithms are required.

[0068] like Figure 2 As shown, the process by which the scene classification model generates control traffic labels, metering traffic labels, and encryption algorithm-selected labels includes:

[0069] The joint loss function is constructed based on the norm sum of squares of the control flow classification parameters, metering flow classification parameters, and encryption algorithm selection parameters; the first product of the control flow dynamic fault tolerance parameters, control flow accuracy weights, and overall control flow classification sample slack variables; the second product of the metering flow dynamic fault tolerance parameters, metering flow accuracy weights, and overall metering flow classification sample slack variables; and the third product of the selection weights and overall encryption algorithm selection sample slack variables.

[0070] Based on the binary classification interval constraint framework, control flow classification constraints and metering flow classification constraints are constructed respectively, and an encryption algorithm selection decision function is constructed based on the multi-classification strategy. The flow classification constraints, metering flow classification constraints, and encryption algorithm selection decision function all contain the power grid data flow characteristics.

[0071] During the training process of the scene classification model, the joint loss function is solved based on the traffic classification constraints, the metering traffic classification constraints, and the encryption algorithm selection decision function to generate control traffic labels, metering traffic labels, and encryption algorithm selection labels. Furthermore, during the solution process of the scene classification model, the control traffic dynamic fault tolerance parameters and the metering traffic dynamic fault tolerance parameters are adaptively and dynamically optimized based on the Bayesian sub-model.

[0072] It is understandable that the constraints, decision functions, and loss functions of the scene classification model are used during the training phase to optimize the model parameters, and then only the optimized decision functions are used for prediction when using the scene classification model to determine the labels.

[0073] Specifically, the joint loss function is:

[0074]

[0075] In the formula, This indicates the classification parameters for control flow. Flow rate classification parameters Encryption algorithm selection parameters Flow bias control Flow metering bias item And the choice of bias term for encryption algorithm Minimizing is the objective of the solution, where the control flow bias term Flow metering bias item And the choice of bias term for encryption algorithm Located in the constraint / decision function, Indicates the control flow classification parameters Flow rate classification parameters Encryption algorithm selection parameters The norm sum of squares, These represent the dynamic fault tolerance parameters for control flow and the dynamic fault tolerance parameters for metering flow, respectively, to allow for classification errors. These represent the weights for controlling traffic accuracy, measuring traffic accuracy, and selection, respectively, to dynamically adjust the task weights in the loss function and control the severity of the penalty for classification errors. These represent the overall control flow classification sample slack variables, the overall metering flow classification sample slack variables, and the encryption algorithm selection sample slack variables, respectively, where for example... Let N represent the slack variables for the first task on the i-th sample, and N represent the total number of slack variables.

[0076] In order to achieve higher classification accuracy for high-priority task power grid control flow classification, the calculated flow accuracy weights are determined accordingly. Greater than the weight of flow rate accuracy This increases the penalty for classification errors, forcing the scene classification model to prioritize reducing errors in this task, while preventing the scene classification model from overfitting to secondary tasks and maintaining overall generalization ability, based on traffic accuracy weights. And the weight of flow measurement accuracy Calculate and determine the selection weights .

[0077] Specifically, the control flow classification constraints are as follows:

[0078]

[0079] In the formula, This represents the control flow label for the i-th sample in the first task. This indicates the parameters for controlling traffic classification. This indicates the characteristics of power grid data flow. The comprehensive kernel function, This indicates the flow bias term. Let N represent the slack variables for the i-th sample in the first task control traffic classification, and N represent the total number of slack variables.

[0080] Specifically, the constraints for classifying metered flow rates are as follows:

[0081]

[0082] In the formula, This represents the metering flow label for the i-th sample in the second task. Indicates the classification parameters of metered flow rate. This indicates the characteristics of power grid data flow. The comprehensive kernel function, This indicates the metering flow bias term. Let N represent the slack variables for the i-th sample in the second task's flow classification, and let N represent the total number of slack variables.

[0083] Specifically, the encryption algorithm selection decision function based on the multi-classification strategy is as follows:

[0084]

[0085] In the formula, This represents the encryption algorithm selection parameters for the k-th encryption algorithm. This indicates the characteristics of power grid data flow. The comprehensive kernel function, This represents the bias term for selecting the encryption algorithm for the k-th encryption algorithm. This indicates that the encryption algorithm for the third task selects a slack variable for the i-th sample. Let represent the encryption algorithm label selection for the i-th sample in the third task, and N represent the total number of slack variables. , These represent the encryption algorithms AES-256, ChaCha20-Poly, AES-128, and SPECK, respectively. Specifically, the preferred multi-classification strategy is the one-vs-rest (OvR) strategy.

[0086] Among them, control flow label Flow metering labels And encryption algorithm selection label They are respectively:

[0087]

[0088] It is understandable that the decision function is the core of SVM for classification prediction. It outputs a category label based on the input data features. In this embodiment, the multi-class problem is decomposed into multiple binary classification problems through the decision function of the One-vs-Rest strategy. The output of the binary classification problem (controlling traffic labels and metering traffic labels) controls the output of the multi-class problem (the encryption algorithm selects the label).

[0089] Please continue reading. Figure 2 The process of adaptive dynamic optimization of the control flow dynamic fault tolerance parameter and the metering flow dynamic fault tolerance parameter based on the Bayesian sub-model includes:

[0090] The Bayesian sub-model searches the control flow parameter search space and the metering flow parameter search space based on hierarchical cross-validation, and uses the control flow dynamic fault tolerance parameter and the metering flow dynamic fault tolerance parameter as the optimal parameter combination, wherein the numerical range of the control flow parameter search space is larger than the numerical range of the metering flow parameter search space.

[0091] Specifically, the Bayesian sub-model loss function is Hinge, with an adaptive learning rate of 0.1 initially, a maximum number of iterations of 1000, and a tolerance of 1e-3. In the dynamic iterative loop, batches of data and their corresponding labels are acquired from the data stream. Predictions are made on the current batch of data, calculating the accuracy and recall of the current model. Based on the preset accuracy and recall, the parameter combination is dynamically adjusted: if the accuracy is lower than the target, the parameter combination is increased to enhance the model's fit to the data, potentially improving accuracy. If the recall is lower than the target, the parameter combination is decreased to reduce the model's constraints, potentially increasing recall. Incremental training is performed on the model, updating the model parameters to adapt to new data batches. BayesSearchCV is imported for Bayesian optimization to search for the optimal parameter combination. Define the search space for control flow parameters and the search space for metering flow parameters. The search space for metering flow parameters is logarithmically uniformly distributed between 0.1 and 30, and the search space for control flow parameters is logarithmically uniformly distributed between 40 and 80. Initialize the Bayesian optimizer opt, the SVC model, the search space, and the rbf kernel function. Perform 50 iterations and use the F1 score as the evaluation metric. Fit the input data x and output data x of the sample to the data scene classification model. Perform Bayesian optimization to search for and output the optimal parameter combination. .

[0092] It is understandable that the scikit-learn tool, which uses online SVM algorithms, dynamically updates the parameter combinations. This eliminates the need for full retraining. The Bayesian sub-model dynamically updates its parameter combinations by monitoring the model's accuracy and recall. It can improve the adaptability of scene classification models in dynamic environments. Accuracy is the proportion of correctly classified samples, recall is the proportion of correctly identified positive samples, and F1 score is the harmonic average of accuracy and recall. It is suitable for imbalanced data.

[0093] It is understood that the hierarchical cross-validation is manifested as hierarchical cross-validation of accuracy and recall with F1 score. The numerical range of the control flow parameter search space being larger than that of the metering flow parameter search space ensures that the dynamic fault tolerance parameter of the control flow is greater than the dynamic fault tolerance parameter of the metering flow corresponding to the metering flow label. This adapts to the higher classification accuracy required for high-priority tasks such as power grid control flow classification, while allowing a certain degree of error for low-priority metering flow classification tasks.

[0094] like Figure 2 As shown, further, the process of adaptively and dynamically optimizing the control flow dynamic fault tolerance parameters and the metering flow dynamic fault tolerance parameters based on the Bayesian sub-model also includes:

[0095] If the accuracy of the current scene classification model is less than the target accuracy threshold, then the optimal parameter combination is multiplied by a first coefficient greater than 1.

[0096] If the recall rate of the current scene classification model is less than the target recall rate threshold, then the optimal parameter combination is divided by a second coefficient greater than 1.

[0097] The first coefficient is greater than the second coefficient.

[0098] Specifically, the process of adjusting the optimal parameter combination is as follows:

[0099]

[0100] In the formula, The adjusted optimal parameter combination includes dynamic fault-tolerant parameters for metered flow. Dynamic fault tolerance parameters for metered flow , This is the optimal parameter combination before adjustment. The preferred coefficient is 1.2. The second coefficient is preferably 1.1. The accuracy of the current scene classification model. The target accuracy threshold, The recall rate of the current scene classification model. This represents the target recall threshold. It is understandable that the process of adjusting the optimal parameter combination described above can be repeated to achieve gradual adjustment of the parameter combination.

[0101] like Figure 2 As shown, the calculation process for the control flow accuracy weight, the metering flow accuracy weight, and the selection weight further includes:

[0102] During the training and optimization process of the scene classification model, the temporary weight of the control flow or the temporary weight of the metering flow is obtained based on the ratio of the exponential function of the prediction accuracy of the control flow or the metering flow to the exponential function of the overall prediction accuracy. The temporary weight of the metering flow is then self-adjusted to obtain the control flow accuracy weight, and the temporary weight of the metering flow is then self-adjusted to obtain the metering flow accuracy weight.

[0103] The selection weight is determined based on the control flow accuracy weight and the metering flow accuracy weight.

[0104] Specifically, the calculation process for the temporary weight of control flow or the temporary weight of metering flow is as follows:

[0105]

[0106] In the formula, This represents the temporary weights for the t-th task in the k-th evaluation period, namely the temporary weights for control flow and metering flow. This represents the accuracy of the t-th task in the k-th evaluation period. This represents the smoothness parameter of the weight distribution.

[0107] Specifically, the process of self-adjusting to derive the control flow accuracy weight / metering flow accuracy weight is as follows:

[0108]

[0109] In the formula, This represents the accuracy weight for the t-th task, namely the control flow accuracy weight and the metering flow accuracy weight. This represents the temporary weight of the t-th task in the k-th evaluation period. The weight distribution smoothness parameter is 0.7 in the task of controlling the flow accuracy weight and 1.1 in the task of measuring flow accuracy weight, so that the flow accuracy weight is greater than the measuring flow accuracy weight.

[0110] Specifically, the process of calculating and determining the selection weights is as follows:

[0111]

[0112] In the formula, To select weights, To control the weight of traffic accuracy, Weighted for the accuracy of flow measurement.

[0113] Understandably, by monitoring the real-time accuracy of each task during the training optimization process, dynamically adjusting the task weights in the loss function, adaptive focusing is achieved to increase the weight of poorly performing tasks, resource optimization is performed to avoid over-focusing on tasks that have already met the target, and balanced convergence is achieved to prevent a single task from dominating the training. In the actual training and optimization process of the scene classification model, the weights continuously iterate and evolve. Since the difference between the prediction accuracy (F1) of the control flow and the prediction accuracy (F1) of the metered flow is small in the initial stage (the first 5 iterations), the weights for both flow accuracy and metered flow accuracy, determined by the above formula, are 0.33, with a selected weight of 0.25. In the mid-term (around the 5th to 10th iteration), the prediction accuracy (F1) of the control flow is greater than that of the metered flow. At this point, the weights for both flow accuracy and metered flow accuracy, determined by the above formula, are 0.28 and 0.38 respectively, with a selected weight of 0.25, achieving dynamic weight adjustment to improve the metered flow accuracy in the mid-term and achieve phased focus. In the later stage (after the 10th iteration), the encryption selection accuracy (F1) is lower than both the prediction accuracy (F1) of the control flow and the metered flow. The control flow weight remains around 0.25 to balance encryption security. The smart meter encryption system can automatically increase the control flow weight to above 0.6 in fault scenarios and reduce the metering flow weight to below 0.2 during peak hours to release resources.

[0114] Furthermore, the process of generating control flow labels, metering flow labels, and encryption algorithm-selected labels using the scenario classification model includes:

[0115] A comprehensive kernel function is constructed based on a weighted average of the linear kernel for control flow labels, the radial basis kernel for metering flow labels, and the linear kernel for encryption selection.

[0116] The scenario classification model is based on a comprehensive kernel function. Under the constraints of traffic classification, metering traffic classification, and encryption algorithm selection, the joint loss function is solved to generate control traffic labels, metering traffic labels, and encryption algorithm selection labels.

[0117] Understandably, linear kernels are suitable for low-dimensional linearly separable data; therefore, both flow labeling and encryption selection tasks use linear kernels. Radial basis kernels capture time-series nonlinear relationships and are suitable for complex data distributions; therefore, flow labeling tasks use radial basis kernels. Their weighted propensity scores are 10, 5, and 1, respectively.

[0118] The above solution balances the different security requirements of controlling and measuring traffic, effectively reduces the probability of data being acquired or forged, improves the anti-tampering capability during transmission, and dynamically adjusts the encryption strategy by combining traffic characteristics with encryption algorithm label selection. This solves the problems of complex key management and difficulty in balancing encryption speed and strength in traditional encryption. Furthermore, through dynamic label classification and fault tolerance parameter optimization, it significantly improves the adaptability of the encryption strategy to actual business.

[0119] like Figure 3 As shown, the process of generating key recommendation strength and traffic outliers further includes:

[0120] The power grid data flow characteristics are combined with the encryption algorithm to select tags and generate subtree key strength coefficients by selecting subtrees based on key strength calculated by minimizing mean square error regression.

[0121] The power grid data flow characteristics are used to generate the flow anomaly values ​​through an anomaly detection subtree constructed based on an unsupervised isolation tree architecture;

[0122] The subtree key strength coefficient is adjusted based on the power grid data flow characteristics and the flow anomalies to generate the recommended key strength.

[0123] Specifically, the key strength selection subtree uses supervised learning trees to predict key strengths of 128 / 192 / 256 digits, and the anomaly detection subtree uses unsupervised isolation trees (Isolation Forest) to calculate anomaly scores. These subtrees are proportionally distributed to the total number of trees in the hybrid random forest key strength decision model, with the key strength selection subtree accounting for 0.7 and the anomaly detection subtree accounting for 0.3.

[0124] Specifically, the training process of the key strength decision model, including the key strength selection subtree, the generation subtree, and the anomaly detection subtree, includes: labeling a dataset containing features with different key strengths and their corresponding strength labels, and an unlabeled anomaly sample dataset; randomly selecting a subset of samples and features to construct a decision tree. At each node, a feature is randomly selected, and the data is split based on the feature value. Recursive splitting continues until a stopping condition is met: the number of samples in a node is less than a threshold. Multiple decision trees form a random forest, predicting key strength through voting or averaging. Using an unsupervised learning algorithm, anomaly scores for samples are calculated by randomly splitting the data. Randomly selecting a subset of samples and features, an isolation tree is constructed. At each node, a feature is randomly selected, and a value between the minimum and maximum values ​​of that feature is randomly chosen as a split point. Recursive splitting continues until a stopping condition is met. Multiple isolation trees form an unsupervised isolation forest, used to calculate the anomaly score for each sample.

[0125] like Figure 3 As shown, the process of generating the traffic anomaly value through the anomaly detection subtree further includes:

[0126] Based on the power grid data flow characteristics, an unsupervised random forest is split to generate an anomaly detection subtree for quantifying the degree of flow anomaly.

[0127] The traffic anomaly value is calculated based on the ratio of the average path length of the isolated tree to the standard path length.

[0128] Specifically, the input data for the anomaly detection subtree is a one-hot encoded sequence of packet length, protocol type, and traffic entropy value. An isolation forest is constructed using randomly selected features and split points. Each isolation forest recursively selects features and split values ​​randomly until a sample is isolated into a single node or the maximum depth is reached. These features are randomly selected from all dimensions of the input data. There are 10 features, where d is the total number of features.

[0129] Specifically, the calculation process for traffic anomalies is as follows:

[0130]

[0131] In the formula, For input data Traffic anomalies, This represents the average path length of the isolation tree. This is the standard path length.

[0132] like Figure 3 As shown, further, the process of adjusting the subtree key strength coefficient to generate the key recommendation strength based on the power grid data flow characteristics and the flow anomalies includes:

[0133] Based on the comparison result between the traffic anomaly value and the first anomaly threshold, the subtree key strength coefficient is increased by a set level to generate the key recommended strength;

[0134] Based on the comparison result between the traffic anomaly value and the second anomaly threshold, the subtree key strength coefficient is increased to the maximum value corresponding to the device capability characteristics to generate the key recommendation strength;

[0135] Wherein, the first abnormal threshold is less than the second abnormal threshold, and the power grid data flow characteristics include equipment capability characteristics.

[0136] Specifically, the key strength selection subtree adopts the CART splitting criterion. Its input data consists of a one-hot encoded sequence of protocol type, latency, and data sensitivity level, as well as device capabilities such as CPU and encryption hardware support. The output data is an ordered classification of key strength, including 128 / 192 / 256 bits. The final subtree key strength coefficient is determined by majority voting.

[0137] Preferably, the first abnormal threshold is 0.6 and the second abnormal threshold is 0.9.

[0138] like Figure 4 As shown, further, the process of calculating the comprehensive predicted encryption overhead and setting the adaptive hybrid security encryption algorithm based on the comprehensive predicted encryption overhead includes:

[0139] The encryption computational overhead is determined based on the selected label using the encryption algorithm, the recommended key strength, and the device performance factor.

[0140] The data inflation rate of the selected tag is based on the encryption algorithm, and the communication overhead is calculated based on the data packet size and latency requirements of the control flow tag or metering flow tag.

[0141] Calculate the security overhead based on the key recommendation strength and traffic anomalies;

[0142] The weighted summation of the encryption computation overhead, communication overhead, and security overhead generates a comprehensive predicted encryption overhead.

[0143] When the predicted encryption overhead exceeds the comprehensive threshold, the encryption algorithm selection label corresponding to the control traffic label is downgraded, the key recommendation strength corresponding to the metering traffic label is downgraded, and an adaptive hybrid security encryption algorithm is generated.

[0144] Specifically, the encryption computation overhead is the product of the basic overhead of the encryption algorithm selection tag in the algorithm basic overhead table (μs / byte), the factor corresponding to the key recommendation strength, and the device performance factor. The communication overhead is the product of the data inflation rate corresponding to the encryption algorithm selection tag and the data packet size and latency requirement corresponding to the control traffic tag or metering traffic tag. A traffic anomaly value minus 0.6 multiplied by 50 determines a security cost of 50 μs added for every 0.1 anomaly point. The key recommendation strength multiplied by 0.01 determines the security benefit. The security cost is determined by subtracting the security benefit from the security cost greater than 0. The encryption computation overhead, communication overhead, and security overhead are weighted and summed to generate a comprehensive predicted encryption overhead, with corresponding weight values ​​of 0.6, 0.3, and 0.1, respectively.

[0145] When the predicted encryption overhead exceeds the comprehensive threshold, the encryption algorithm selection label corresponding to the control traffic label is downgraded, and the key recommendation strength corresponding to the metering traffic label is downgraded, in order to adapt to the needs of prioritizing low latency for control traffic and balancing security and efficiency for metering traffic.

[0146] In the above scheme, the key strength decision model is based on a hybrid random forest framework. It dynamically adjusts the key strength by combining anomaly detection subtrees, which solves the contradiction between the difficulty of key management in traditional symmetric encryption and the slow speed of asymmetric encryption. It dynamically enhances the security level through anomaly detection, while avoiding the energy consumption and latency problems caused by over-encryption, which meets the requirements of smart grids for low energy consumption and high efficiency.

[0147] In this embodiment, an AI-driven selection algorithm based on SVM intelligently selects the current encryption algorithm and key strength based on the characteristics of the smart meter's uploaded traffic. This achieves hybrid encryption of symmetric and asymmetric encryption, improving the adaptability of hybrid encryption to the high-frequency, multi-device data transmission requirements of smart meters, balancing data security and encryption overhead, and providing reliable protection for power grid data security. It balances the different security requirements of control traffic and metering traffic, effectively reducing the probability of data acquisition or forgery, improving anti-tampering capabilities during transmission, and dynamically adjusting encryption strategies by combining traffic characteristics with encryption algorithm selection tags. This solves the problems of complex key management and difficulty in balancing encryption speed and strength in traditional encryption. Furthermore, dynamic tag classification and fault tolerance parameter optimization significantly improve the adaptability of encryption strategies to actual business operations. The key strength decision model is based on a hybrid random forest framework, dynamically adjusting key strength by combining anomaly detection subtrees. This resolves the contradiction between the difficulty of key management in traditional symmetric encryption and the slow speed of asymmetric encryption. It dynamically enhances security levels through anomaly detection while avoiding energy consumption and latency issues caused by excessive encryption, meeting the smart grid's requirements for low energy consumption and high efficiency.

[0148] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

[0149] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A hybrid security encryption method for power grid data based on artificial intelligence, characterized in that, include: The power grid data flow characteristics reported by smart meters are used to generate control flow labels, metering flow labels, and encryption algorithm selection labels through a scenario classification model. The scenario classification model is built on a multi-output classification SVM framework and trained using a joint loss function. The joint loss function is constructed by weighting the control flow labels based on the dynamic fault tolerance parameters of the control flow and the metering flow labels based on the dynamic fault tolerance parameters of the metering flow. The power grid data flow characteristics are combined with the encryption algorithm to select labels and input them into the key strength decision model to generate key recommended strength and flow anomaly values, wherein the key strength decision model is constructed based on a hybrid random forest framework. Based on the control flow label, metering flow label, encryption algorithm selection label and key recommendation strength, a comprehensive prediction encryption cost is calculated, and an adaptive hybrid security encryption algorithm for single communication between the smart meter and the cloud server is set based on the comprehensive prediction encryption cost. The process of generating key recommendation strength and traffic outliers includes: The power grid data flow characteristics are combined with the encryption algorithm to select tags and generate subtree key strength coefficients by selecting subtrees based on key strength calculated by minimizing mean square error regression. The power grid data flow characteristics are used to generate the flow anomaly values ​​through an anomaly detection subtree constructed based on an unsupervised isolation tree architecture; The subtree key strength coefficient is adjusted based on the power grid data flow characteristics and the flow anomalies to generate the recommended key strength. The key strength decision model includes the key strength selection subtree and the anomaly detection subtree.

2. The AI-driven hybrid security encryption method for power grid data according to claim 1, characterized in that, The process by which the scenario classification model generates control flow labels, metering flow labels, and encryption algorithm-selected labels includes: The joint loss function is constructed based on the norm sum of squares of the control flow classification parameters, metering flow classification parameters, and encryption algorithm selection parameters; the first product of the control flow dynamic fault tolerance parameters, control flow accuracy weights, and overall control flow classification sample slack variables; the second product of the metering flow dynamic fault tolerance parameters, metering flow accuracy weights, and overall metering flow classification sample slack variables; and the third product of the selection weights and overall encryption algorithm selection sample slack variables. Based on the binary classification interval constraint framework, control flow classification constraints and metering flow classification constraints are constructed respectively, and an encryption algorithm selection decision function is constructed based on the multi-classification strategy. The flow classification constraints, metering flow classification constraints, and encryption algorithm selection decision function all contain the power grid data flow characteristics. During the training process of the scene classification model, the joint loss function is solved based on the traffic classification constraints, the metering traffic classification constraints, and the encryption algorithm selection decision function to generate control traffic labels, metering traffic labels, and encryption algorithm selection labels. Furthermore, during the solution process of the scene classification model, the control traffic dynamic fault tolerance parameters and the metering traffic dynamic fault tolerance parameters are adaptively and dynamically optimized based on the Bayesian sub-model.

3. The AI-driven hybrid security encryption method for power grid data according to claim 2, characterized in that, The process of adaptive dynamic optimization of the control flow dynamic fault tolerance parameter and the metering flow dynamic fault tolerance parameter based on the Bayesian sub-model includes: The Bayesian sub-model searches the control flow parameter search space and the metering flow parameter search space based on hierarchical cross-validation, and uses the control flow dynamic fault tolerance parameter and the metering flow dynamic fault tolerance parameter as the optimal parameter combination, wherein the numerical range of the control flow parameter search space is larger than the numerical range of the metering flow parameter search space.

4. The AI-driven hybrid security encryption method for power grid data according to claim 3, characterized in that, The process of adaptive dynamic optimization of the control flow dynamic fault tolerance parameter and the metering flow dynamic fault tolerance parameter based on the Bayesian sub-model also includes: If the accuracy of the current scene classification model is less than the target accuracy threshold, then the optimal parameter combination is multiplied by a first coefficient greater than 1. If the recall rate of the current scene classification model is less than the target recall rate threshold, then the optimal parameter combination is divided by a second coefficient greater than 1. The first coefficient is greater than the second coefficient.

5. The AI-driven hybrid security encryption method for power grid data according to claim 2, characterized in that, The calculation process for the control flow accuracy weight, the metering flow accuracy weight, and the selection weight includes: During the training and optimization process of the scene classification model, the temporary weight of control flow or temporary weight of metered flow is obtained based on the ratio of the exponential function of the prediction accuracy of control flow or metered flow to the exponential function of the overall prediction accuracy. The temporary weight of control flow is then self-adjusted to obtain the control flow accuracy weight, and the temporary weight of metered flow is self-adjusted to obtain the metered flow accuracy weight. The selection weight is determined based on the control flow accuracy weight and the metering flow accuracy weight.

6. The AI-driven hybrid security encryption method for power grid data according to claim 2, characterized in that, The process of generating control flow labels, metering flow labels, and encryption algorithm-selected labels using the scenario classification model includes: A comprehensive kernel function is constructed by weighting the linear kernel of the control flow label, the radial basis kernel of the metering flow label, and the encrypted selection linear kernel.

7. The AI-driven hybrid security encryption method for power grid data according to claim 1, characterized in that, The process of generating the traffic anomaly value through the anomaly detection subtree includes: Based on the power grid data flow characteristics, an unsupervised random forest is split to generate an anomaly detection subtree for quantifying the degree of flow anomaly. The average path length of the isolated tree and the standard path length are calculated using the anomaly detection subtree; The abnormal traffic value is calculated based on the ratio of the average path length of the isolated tree to the standard path length.

8. The AI-driven hybrid security encryption method for power grid data according to any one of claims 1 to 7, characterized in that, The process of adjusting the subtree key strength coefficient and generating the recommended key strength based on the power grid data flow characteristics and the flow anomalies includes: Based on the comparison result between the traffic anomaly value and the first anomaly threshold, the subtree key strength coefficient is increased by a set level to generate the key recommended strength; Based on the comparison result between the traffic anomaly value and the second anomaly threshold, the subtree key strength coefficient is increased to the maximum value corresponding to the device capability characteristics to generate the key recommendation strength; Wherein, the first abnormal threshold is less than the second abnormal threshold, and the power grid data flow characteristics include equipment capability characteristics.

9. The AI-driven hybrid security encryption method for power grid data according to any one of claims 1 to 7, characterized in that, The process of calculating the comprehensive predicted encryption overhead and setting an adaptive hybrid security encryption algorithm based on the comprehensive predicted encryption overhead includes: The encryption computational overhead is determined based on the selected label using the encryption algorithm, the recommended key strength, and the device performance factor. The data inflation rate of the selected tag is based on the encryption algorithm, and the communication overhead is calculated based on the data packet size and latency requirements of the control flow tag or metering flow tag. Calculate the security overhead based on the key recommendation strength and traffic anomalies; The weighted summation of the encryption computation overhead, communication overhead, and security overhead generates a comprehensive predicted encryption overhead. When the predicted encryption overhead exceeds the comprehensive threshold, the encryption algorithm selection label corresponding to the control traffic label is downgraded, the key recommendation strength corresponding to the metering traffic label is downgraded, and an adaptive hybrid security encryption algorithm is generated.

Citation Information

Patent Citations

  • Encryption optimization method for data communication

    CN118944952A

  • Deep reinforcement learning optimization algorithm based on adaptive encryption framework

    CN119513891A