Power data privacy protection method and related equipment based on data classification

By combining feature analysis model and preset privacy algorithm, the problem of inefficiency of traditional power data protection methods is solved, and efficient classification and privacy protection of power data is achieved, which is suitable for power data security in smart grids.

CN118246062BActive Publication Date: 2025-08-19FIBRLINK NETWORKS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410184215.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-08-19
Estimated Expiration
2044-02-19

AI Technical Summary

Technical Problem

Traditional power data protection methods are difficult to cope with changes in complex relationships, are inefficient and difficult to guarantee classification accuracy, and cannot effectively cope with the diversity and complexity of power data in smart grids.

Method used

The privacy level of each attribute feature in the power data is determined through the feature analysis model, and the processing is performed using a preset privacy algorithm, combining the mutual information algorithm and the Transformer model to process nonlinear relationships, and using K-Means clustering and differential privacy algorithms for data hierarchical protection.

Benefits of technology

It realizes effective privacy protection for power data based on data grading, improves classification accuracy and efficiency, and ensures the security of power data and user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118246062B_ABST
    Figure CN118246062B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and related equipment for protecting power data privacy based on data grading. Specifically, the method includes: obtaining a first feature dataset; wherein the first feature dataset includes at least one attribute feature; based on the first feature dataset, using a feature analysis model to determine the privacy level corresponding to each attribute feature; wherein the feature analysis model is a trained neural network model; and based on the first feature dataset and the privacy level corresponding to each attribute feature, using a preset privacy algorithm to perform privacy processing to obtain a second feature dataset. This method provides privacy protection for power data based on data grading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of power data processing, and in particular to a power data privacy protection method based on data classification and related equipment. Background Art

[0002] With the advancement of digitalization in smart grid power data, the volume and frequency of data changes are increasing, making it increasingly difficult to protect the privacy of power data. Traditional rule-based methods struggle to cope with the complex relationships and changes in power data, requiring extensive manual intervention, resulting in low efficiency and difficulty ensuring classification accuracy. Summary of the Invention

[0003] In view of this, the purpose of the present disclosure is to propose a power data privacy protection method based on data classification and related equipment.

[0004] Based on the above objectives, the present disclosure provides a method for protecting power data privacy based on data classification, including:

[0005] Acquire a first feature data set; wherein the first feature data set includes at least one attribute feature;

[0006] Determining the privacy level corresponding to each attribute feature using a feature analysis model based on the first feature data set; wherein the feature analysis model is a trained neural network model;

[0007] According to the first feature data set and the privacy level corresponding to each attribute feature, a preset privacy algorithm is used to perform privacy processing to obtain a second feature data set.

[0008] Based on the same inventive concept, an embodiment of the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the power data privacy protection method as described in any one of the above items is implemented.

[0009] Based on the same inventive concept, an embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the above-mentioned power data privacy protection methods.

[0010] Based on the same inventive concept, an embodiment of the present disclosure further provides a computer program product, including computer program instructions. When the computer program instructions are executed on a computer, the computer executes any of the above-mentioned power data privacy protection methods.

[0011] From the above description, it can be seen that the present disclosure provides a method and related equipment for protecting the privacy of electric power data based on data classification. The method determines the privacy level of each attribute feature in the first feature data set through a feature analysis model, and performs privacy processing using a preset privacy algorithm based on the first feature data set and the privacy level corresponding to each attribute feature to obtain a second feature data set, thereby realizing privacy protection for electric power data based on data classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0013] Figure 1 A schematic diagram illustrating an application scenario of a power data privacy protection method provided by an embodiment of the present disclosure is shown;

[0014] Figure 2 A schematic diagram illustrating a training process of a privacy protection model provided by an embodiment of the present disclosure is shown;

[0015] Figure 3 A schematic diagram illustrating a flow chart of a power data privacy protection method provided by an embodiment of the present disclosure;

[0016] Figure 4 A schematic diagram illustrating a flow chart of another power data privacy protection method provided by an embodiment of the present disclosure;

[0017] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0018] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0019] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.

[0020] Power data covers user electricity usage, energy supply conditions, and system operation information. This information is highly sensitive. Unauthorized access or leakage can pose a serious threat to user privacy and system security.

[0021] The advancement of digitalization of smart grid power data has triggered an urgent need for data security and privacy protection. Traditional power data protection methods typically rely on fixed rules and static models, which perform poorly with the diversity and complexity of power data. Consequently, they suffer from a range of problems, including inaccurate classification and low efficiency.

[0022] In view of this, the embodiments of the present disclosure provide a power data privacy protection method and related equipment based on data classification. Specifically, a feature analysis model is used to determine the privacy level of each attribute feature in a first feature data set. Based on the first feature data set and the privacy level corresponding to each attribute feature, a preset privacy algorithm is used to perform privacy processing to obtain a second feature data set, thereby providing privacy protection for power data based on data classification.

[0023] In order to make the technical solution of the present disclosure clearer and easier to understand, the scenario architecture of the power data privacy protection method provided by the embodiment of the present disclosure is introduced below with reference to the accompanying drawings.

[0024] Figure 1 The power data privacy protection method provided by the embodiment of the present disclosure is provided in a schematic diagram of an application scenario. The power data privacy protection method provided by the embodiment of the present disclosure is not limited to applications such as Figure 1 In the application scenario 100 shown in FIG. Figure 1As shown, the application scenario includes a data storage system 102, a server 104 and devices 106A and 106B. The data storage system 102, the server 104 and the devices 106A and 106B can be connected via a wired or wireless communication network. The data storage system 102 and the server 104 can be independent physical servers, or a server cluster or distributed system composed of multiple physical servers. They can also be cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Devices 106A and 106B include but are not limited to desktop computers, mobile phones, mobile computers, tablet computers, media players, smart wearable devices, personal digital assistants (PDAs), servers or other electronic devices that can achieve the above functions.

[0025] A large amount of training feature data is stored in the data storage system 102. The training feature data includes at least one attribute feature and its corresponding preset privacy level. The server 104 can train the hybrid feature selection model based on the large amount of training feature data and analyze the privacy levels corresponding to the features in the input training attribute feature set and the training candidate feature set. When the loss function of the hybrid feature selection model reaches a minimum value, the server 104 can obtain a first feature data set from the data storage system 102, and then perform feature screening based on the privacy level on the first feature data set based on the feature analysis model in the hybrid feature selection model, so as to facilitate privacy protection of power data of different privacy levels. Finally, the privacy-protected power data corresponding to the permissions of the devices 106A and 106B can be provided. At the same time, the server 104 can also continuously optimize the hybrid feature selection model based on the newly added training feature data.

[0026] It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present disclosure, and the embodiments of the present disclosure are not limited in this respect. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0027] Next, the power data privacy protection method provided by the embodiment of the present disclosure will be described in detail from the perspective of training the hybrid feature selection model and the perspective of privacy protection using the hybrid feature selection model.

[0028] Figure 2A schematic diagram of the training process of a privacy protection model provided by an embodiment of the present disclosure is shown. It should be understood that the Transformer model in the trained privacy protection model can analyze the privacy level of data features, thereby helping to achieve privacy protection and availability of power data.

[0029] See also Figure 2 , obtaining a training feature data set 201. Here, the training feature data set includes at least one attribute feature of the power data and its corresponding privacy level.

[0030] Exemplarily, the power data exists in the form of a time series, recording various parameters and indicators of the power system at different time points. It will be understood by those skilled in the art that the attribute characteristics can be any parameters or indicators that describe the power data, and the present disclosure does not limit this. For example, attribute characteristics may include: timestamp (generally, the timestamp itself may not contain sensitive information), frequency (the frequency information of the power system may be related to the operating mode of certain activities or equipment), voltage deviation (voltage deviation may be related to the stability of the power system and the status of the equipment), power, electric energy, power load (power load information directly reflects the power demand of the system and may reveal the user's life pattern and activities).

[0031] The preset privacy levels include four privacy levels, namely Level 1, Level 2, Level 3, and Level 4. The privacy level classifications corresponding to the attribute features of the training feature dataset 201 can be: Privacy Level 1 - Timestamp, Privacy Level 2 - Frequency, Privacy Level 3 - Voltage Deviation, Privacy Level 4 - Power, Electric Energy, Electric Load.

[0032] Next, data segmentation 202 is performed to determine a training attribute feature set T and its corresponding training candidate feature set W.

[0033] For example, the voltage deviation data of privacy level 3 is selected as the training attribute feature set T, expressed as: T = {voltage deviation: v1}; then the remaining attribute feature data (including but not limited to frequency, electric energy, etc.) constitute the training candidate feature set W, expressed as: W = {timestamp: t1, frequency: f1, power: p1, electric energy: e1, power load: l1}.

[0034] It should be understood that when other privacy levels and other attribute features are selected, corresponding training attribute feature sets T and training candidate feature sets W can be obtained. In other words, based on the differences in privacy levels and attribute features, the training feature dataset can be divided into multiple sets of training attribute feature sets T and training candidate feature sets, and each set of training attribute feature sets T and training candidate feature sets W can be input into the hybrid feature selection model 203 for training.

[0035] The hybrid feature selection model 203 includes a mutual information algorithm and a Transformer model. Among them, mutual information (MI) is used as an indicator to measure the correlation between random variables. In the attribute feature selection of the present disclosure, mutual information is used to quantify the relationship between each attribute feature and the privacy level. The Transformer model is a neural network based on the attention mechanism, which is very suitable for complex sequence modeling tasks, including processing nonlinear relationships. The present disclosure combines the Transformer model with the mutual information algorithm to overcome the limitations of mutual information in processing nonlinear relationship data, taking into account both linear and nonlinear relationships.

[0036] In the hybrid feature selection model 203 , feature selection is first performed based on the mutual information algorithm, and then combined with the Transformer model to make up for the deficiency of the mutual information algorithm in performing poorly in processing nonlinear relationship data.

[0037] Exemplarily, the mutual information algorithm is as follows:

[0038]

[0039] Among them, MI(X i ; Y i ) represents the mutual information between feature X and text category Y, and p(x, y) represents X=x i and Y = y i The probability of simultaneous occurrence, that is, the joint probability distribution, p(x) and p(y) represent X=x i and Y = y i Using the above probability distribution, the mutual information value I(X; Y) between the feature and the target variable is calculated according to the mutual information formula. For each feature, the mutual information value between it and the target variable is calculated to evaluate the correlation between them.

[0040] The mutual information algorithm is used to calculate the correlation between the training attribute feature set T = {v1} and the privacy level P. The following formula can be used:

[0041]

[0042] The joint probability distribution p(v1, P) and the marginal probability distribution p(v1) can be determined based on the training attribute feature set T = {v1} and the training candidate feature set W = {t1, f1, p1, e1, l1}. The selection of the training attribute feature set T and the training candidate feature set W reflects the correlation between the attribute features (e.g., voltage deviation) and the privacy level. The role of the training candidate feature set W is to ensure that the attributes in the training candidate training set W are considered when calculating the mutual information between the training attribute feature set T and the target variable P. This can integrate information from different features to obtain more accurate information entropy, thereby improving the representativeness and accuracy of the training feature subset.

[0043] In some embodiments, the mutual information value between each attribute feature and the target variable (e.g., privacy level) at each privacy level is calculated; the mutual information values of the attribute features at each privacy level are arranged in descending order, and the attribute features are screened using a threshold A to obtain a final training feature subset. Optionally, attribute features greater than the threshold A are selected to enter the training feature subset. Here, the threshold A can be pre-set, and the present disclosure does not limit the specific value.

[0044] Setting the feature selection threshold A helps ensure that more nonlinear attribute features are retained, balancing the number of selected attribute features and their complexity. Using the mutual information algorithm to preliminarily screen the features of the training feature dataset 201 (i.e., the original power data) to form a training feature subset helps narrow the range of features and improve the efficiency of subsequent model training.

[0045] The power system has complex dynamic behaviors and interdependencies, and nonlinear relationships often exist between attribute features in the data. Therefore, based on mutual information feature selection, the Transformer model is introduced to better handle nonlinear relationships.

[0046] To handle the nonlinear relationships in power data, this paper embeds a mutual information loss function into the training of the Transformer model. The mutual information loss function is defined by maximizing the average mutual information across samples. The overall loss function includes the reconstruction loss function of the Transformer model and the mutual information loss function.

[0047] Based on the mutual information algorithm, the mutual information value between the attribute features and the privacy level in the power data is obtained. In some embodiments, the loss function defining the mutual information algorithm is determined based on the average mutual information of each attribute feature in the training feature subset. Exemplarily, the calculation formula of the loss function is as follows:

[0048]

[0049] Where N is the number of attribute features. For example, if the attribute features of the training feature subset are frequency, voltage deviation, power, energy, and power load, then N is 5. MI(xi,yi) represents the mutual information between the attribute features and their corresponding class labels. After training, the average mutual information of each attribute feature is maximized.

[0050] During model training, the sum of the reconstruction loss function of the Transformer model and the loss function of the mutual information algorithm, i.e., the overall loss function, is optimized to simultaneously learn effective attribute feature representation and attribute feature selection. The overall loss function can be defined as the following formula:

[0051] Total Loss=Loss Transformer +α*Loss MI

[0052] Among them, Loss Transformer is the reconstruction loss function of the Transformer model, Loss MI It is the loss function of the mutual information algorithm, and α is a hyperparameter used to balance the importance of the two loss functions.

[0053] Optionally, the reconstruction loss function can be a cross-entropy loss, which is used to measure the difference between the predicted output of the model on the training sample and the actual label.

[0054] Loss MI Mutual information selection loss is used to constrain feature selection within the training feature subset. By maximizing the average mutual information of the training feature subset, the model is guided to select features that are helpful for the classification task. During model training, the training feature subset is used as input for the Transformer model. The Transformer model uses a self-attention mechanism to focus on the relationships between different features during the learning process.

[0055] By combining these two loss functions, Total Loss balances both feature learning and feature selection during model training. During model training, by optimizing the overall loss function—minimizing the reconstruction loss and maximizing the average of the mutual information algorithm—it simultaneously learns effective feature representation and performs feature selection. This allows the model to retain useful feature information while also considering the correlation between features and privacy levels during learning. In other words, by minimizing the overall loss function, the hybrid feature selection model effectively balances feature representation and attribute feature selection.

[0056] It should be noted that the trained Transformer model can analyze the privacy level of attribute features and obtain the privacy level corresponding to the attribute features.

[0057] In order to facilitate understanding of the power data privacy protection method provided by the embodiment of the present disclosure, Figure 3 Provide a detailed introduction.

[0058] Figure 3 A flowchart illustrating a method for protecting power data privacy provided by an embodiment of the present disclosure is shown. First, a first feature dataset 301 is obtained. First feature dataset 301 is similar to the aforementioned training feature dataset, except that at least some of the attribute features in first feature dataset 301 lack corresponding privacy levels. In other words, first feature dataset 301 contains attribute features with unknown privacy levels.

[0059] Next, the privacy level corresponding to each attribute feature is determined using the feature analysis model 302. It should be noted that the feature analysis model 302 is the trained Transformer model mentioned above and will not be described in detail.

[0060] In some embodiments, the privacy level can be identified by a label. For example, data sample 1: {frequency (f1), power (p1), privacy level (category label)}; data sample 2: {frequency (f2), power (p2), privacy level (category label)}... data sample N: {frequency (fn), power (pn), privacy level (category label)}.

[0061] Next, a preset privacy algorithm is used to perform privacy protection. Optionally, privacy protection can be performed with the help of the privacy protection module 303.

[0062] In some embodiments, the privacy protection module 303 includes two parts: a clustering algorithm (eg, a K-Means clustering algorithm) and a differential privacy algorithm. The clustering algorithm and the differential privacy algorithm are introduced below.

[0063] The K-Means clustering algorithm clusters and partitions data based on distance or similarity between them, thereby finding equivalent clusters of central points. The main purpose of this part is to divide the data points in the dataset into K clusters. These clusters can be regarded as collections of data points with similar characteristics.

[0064] The differential privacy algorithm has the following relevant definitions.

[0065] Definition 1: An adjacent dataset refers to a dataset D and a dataset D′ obtained by deleting or modifying a record in D. These two datasets D and D′ are called mutually adjacent datasets.

[0066] Definition 2: ε-differential privacy. Assume that there is a differential privacy mechanism R. When R satisfies the following formula, it means that R can achieve ε-differential privacy protection.

[0067] P r [R(D)∈S R ]≤exp(ε)·P r [R(D′)∈S R ]+δ

[0068] Where, P r represents the probability distribution, P r [R(D)∈S R ] indicates that the output of mechanism R on data set D falls on S R The probability of P r [R(D')∈S R ] indicates that the output of mechanism R on the adjacent dataset D′ falls in S R R(D): represents the output after applying the differential privacy mechanism R on the dataset D. exp(ε): represents e to the power of ε, where ε is the privacy budget. ε is used to control the level of privacy protection, and a smaller ε value indicates stronger privacy protection. S R represents the output space, that is, the set of possible output results, such as the statistical results of a certain range of electricity consumption. D is the original power dataset, which can be the first feature dataset here; D′ represents the near dataset adjacent to D, that is, the dataset with only one data item different. δ is the privacy distortion boundary, which is used to measure the uncertainty of privacy leakage. The smaller the value of δ, the smaller the privacy leakage. It should be noted that if the first feature dataset includes attribute features that do not require privacy protection, then the attribute features do not need to be privacy protected.

[0069] Definition 3: Add a noise mechanism. Two commonly used noise mechanisms are the Laplace mechanism and the exponential mechanism. The Laplace mechanism is suitable for protecting numerical data. The Laplace mechanism implements ε-differential privacy by adding random noise that follows a Laplace distribution to the exact query results. Let Lap(b) be the Laplace distribution with location parameter 0 and scale parameter b. The probability density function of the Laplace distribution is shown below:

[0070]

[0071] In some embodiments, privacy protection begins with the selection of the initial clustering point. In the K-means clustering algorithm, the selection of the initial point can affect the clustering results. Alternatively, the K-means++ algorithm is used to select the initial point.

[0072] For example, the steps of selecting the initial point using the K-means++ algorithm are as follows:

[0073] Step 1: Randomly select a data point from the dataset D as the first cluster center C1.

[0074] Step 2: For each data point x∈D, calculate its closest distance D(x) to the currently selected cluster center set C. Specifically, for each data point x, calculate D(x)=min(distance(x,c)), where distance(x,c) is the distance from the data point x to the cluster center c.

[0075] Step 3: Select the next cluster center based on the probability distribution. For each data point x∈D, calculate the probability of it being selected as the next cluster center, that is, P(x)=D(x) 2 / ∑(D(x) 2 ). In this way, points that are farther away are more likely to be selected as the next cluster center.

[0076] Step 4: Repeat steps 2 and 3 until K cluster centers are selected.

[0077] Step 5: Use the selected K cluster centers as the K initial cluster centers, and then continue the iterative process of the K-means algorithm.

[0078] Next, K-Means clustering is performed. For example, for a data set containing n data points, represented as D = {x1, x2, ..., x n}, where each data point x i It is a d-dimensional vector that represents the characteristics of power data. The K-means clustering algorithm is used to cluster the power data. The specific steps are as follows:

[0079] Step 1: Select K cluster centers according to the previous steps, that is, the number of clusters into which the power data should be divided;

[0080] Step 2: Denote the selected cluster centers as C = {c1, c2, ..., c K}, where c i is a d-dimensional vector.

[0081] Step 3: Calculate the distance between the data point and the cluster center. For each data point x i , calculate its relationship with each cluster center c i The distance is calculated using the Euclidean distance formula as follows:

[0082]

[0083] Among them, x i,k Represents the data point x i The kth feature, c i,k Represents the cluster center c i The kth feature of .

[0084] In some embodiments, by adding Laplace noise when calculating the distance between a data point and a cluster center, data privacy is protected using a differential privacy algorithm.

[0085] Specifically, assuming x i is a data point, c i is the cluster center, and d(x i , c i ), namely Distance(xi,cj), represents the data point x i and cluster center c i Without privacy protection, the distance is directly calculated by the above formula to obtain d(x i , c i When applying differential privacy, it is necessary to perform privacy perturbation on the distance, that is, to add Laplace noise to obtain the perturbed distance d′(x i , c i ), the specific principle is shown in the following formula:

[0086] d′(x i , c i )=d(x i , c i )+LaP(0,S / ε)

[0087] LaP(0, S / ε) represents Laplace noise with mean 0 and scale parameter S / ε. ε is the privacy budget for differential privacy, and S is the sensitivity of the cluster center coordinates, indicating the maximum range of variation of the cluster center coordinates. The privacy budget is determined based on the privacy level.

[0088] By adding Laplace noise, the perturbation distance obtained will directly affect the cluster center. For the i-th cluster, the perturbed cluster center will overlap with the perturbed cluster centers of other clusters, thereby blurring the cluster boundaries and increasing the privacy protection effect.

[0089] Step 4: Assign data points to the nearest cluster center. For each data point x i , assign it to the cluster closest to the cluster center. The specific formula is as follows:

[0090] Cluster(x i )=arg min j Distance(x i , c j )

[0091] Step 5: Update the cluster center. For each cluster, calculate the average value of all data points in the cluster and use it as the new cluster center. The specific formula is as follows:

[0092]

[0093] Step 6: Repeat the iterations. Repeat steps 3 and 4 until the stopping condition is met. The stopping condition can be reaching a specified number of iterations or when the change in cluster center is less than a certain threshold, the algorithm is considered to have converged.

[0094] Step 7: Obtain the final clustering results. The final clustering results are K clusters, each containing a group of similar power data points. These clusters can represent sets of data points with similar power usage patterns or characteristics.

[0095] Finally, the privacy-protected data of the power data is output, for example, the noisy clustering results are published to trusted data users without publishing the original data, and D′ is output as a privacy-protected anonymous dataset.

[0096] The technical solution of the disclosed embodiment introduces an innovative hybrid feature selection method that combines mutual information and the Transformer model. The method uses mutual information to measure the correlation between features and target variables, and uses the Transformer model to better handle the nonlinear relationship of power data, thereby achieving more accurate analysis of the correspondence between attribute features and privacy levels. The trained Transformer model has good robustness, especially when power data has complexity, multiple privacy levels, and multiple attribute features, and can robustly handle diverse data.

[0097] In addition, a privacy protection method provided by an embodiment of the present disclosure integrates hybrid feature selection, K-Means clustering and differential privacy technology to design a comprehensive model to process complex power data. It can comprehensively respond to various characteristics of power data, deeply integrate different technical modules, and provide comprehensive privacy protection for the entire power system.

[0098] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.

[0099] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0100] Figure 4 FIG. 1 is a flow chart showing another method for protecting the privacy of electric power data provided by an embodiment of the present disclosure. Figure 4 As shown in FIG, the power data privacy protection method based on data classification includes:

[0101] S402: Acquire a first feature data set; wherein the first feature data set D includes at least one attribute feature, such as timestamp, power, voltage deviation, etc.;

[0102] S404: Determine the privacy level corresponding to each attribute feature using a feature analysis model based on the first feature data set; wherein the feature analysis model is a trained neural network model; optionally, the feature analysis model is constructed based on a Transformer model;

[0103] S406: Based on the first feature data set and the privacy level corresponding to each attribute feature, a preset privacy algorithm is used to perform privacy processing to obtain a second feature data set. Here, the second feature data set may be D'.

[0104] In some embodiments, S406: the preset privacy algorithm includes a clustering algorithm and a differential privacy algorithm; wherein, the clustering algorithm is used to perform clustering processing on at least part of the attribute features of the first feature data set; and the differential privacy algorithm is used to add noise to at least one parameter of the clustering processing.

[0105] Optionally, the parameter is the Euclidean distance.

[0106] In some embodiments, the differential privacy algorithm includes a differential privacy budget; wherein the differential privacy budget is determined based on the privacy level.

[0107] In some embodiments, the clustering algorithm is a K-means clustering algorithm; the K-means clustering algorithm includes selecting a clustering initial point; wherein the clustering initial point is determined by a K-means++ algorithm.

[0108] In some embodiments, the feature analysis model is configured to:

[0109] The training feature subset is obtained by training a pre-built Transformer model; wherein, the training feature subset is obtained by filtering the acquired training feature data set through a mutual information algorithm; wherein, the training feature data set includes at least one attribute feature and its corresponding privacy level.

[0110] In some embodiments, the overall loss function of the feature analysis model includes a loss function of a mutual information algorithm and a loss function of the Transformer model;

[0111] The feature analysis model is further configured to: in response to determining that the overall loss function is minimized, terminate the Transformer model training.

[0112] In some embodiments, the loss function of the mutual information algorithm is determined based on the average value of the mutual information of each attribute feature in the training feature subset.

[0113] In some embodiments, the training feature data set is configured to be divided into a training attribute feature set T of at least one attribute feature and its corresponding training candidate feature set W according to the at least one attribute feature and its corresponding privacy level; wherein,

[0114] The training attribute feature set is configured as a feature variable in the mutual information algorithm; the training candidate feature set is configured to associate the joint probability distribution and the marginal probability distribution in the mutual information algorithm; and the preset privacy level is configured as a target variable in the mutual information algorithm.

[0115] The method of the above embodiment is used to implement the corresponding power data privacy protection method in any of the above embodiments, and has the beneficial effects of the corresponding power data privacy protection method embodiment, which will not be repeated here.

[0116] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the power data privacy protection method described in any of the above embodiments is implemented.

[0117] Figure 5 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.

[0118] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0119] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0120] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0121] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).

[0122] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).

[0123] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0124] The electronic device of the above embodiment is used to implement the corresponding power data privacy protection method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0125] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the power data privacy protection method described in any of the above embodiments.

[0126] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0127] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the power data privacy protection method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0128] Based on the same inventive concept, corresponding to the power data privacy protection method described in any of the above embodiments, the present disclosure also provides a computer program product comprising computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to perform the color correction method. For each step in each embodiment of the color correction method, the processor executing the step can be a member of the corresponding execution entity.

[0129] The computer program product of the above embodiment is used to enable the computer and / or the processor to execute the power data privacy protection method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0130] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0131] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0132] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0133] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A power data privacy protection method based on data classification, characterized in that: include: Acquire a first feature data set; wherein the first feature data set includes at least one attribute feature; Determining the privacy level corresponding to each attribute feature using a feature analysis model based on the first feature data set; wherein the feature analysis model is a trained neural network model; According to the first feature data set and the privacy level corresponding to each attribute feature, a preset privacy algorithm is used to perform privacy processing to obtain a second feature data set; wherein, The feature analysis model is configured to be obtained by training a pre-built Transformer model using a training feature subset; wherein the training feature subset is obtained by filtering an acquired training feature dataset using a mutual information algorithm; wherein the training feature dataset includes at least one attribute feature and its corresponding privacy level; The training feature data set is configured to be divided into a training attribute feature set of at least one attribute feature and its corresponding training candidate feature set based on the at least one attribute feature and its corresponding privacy level; wherein the training attribute feature set corresponds to an attribute feature.

2. The power data privacy protection method according to claim 1, characterized in that: The feature analysis model is built based on the Transformer model.

3. The power data privacy protection method according to claim 1, characterized in that: The preset privacy algorithm includes a clustering algorithm and a differential privacy algorithm; The clustering algorithm is used to perform clustering processing on at least part of the attribute features of the first feature data set; and the differential privacy algorithm is used to add noise to at least one parameter of the clustering processing.

4. The power data privacy protection method according to claim 3, characterized in that: The differential privacy algorithm includes a differential privacy budget; wherein the differential privacy budget is determined based on the privacy level; and / or The clustering algorithm is a K-means clustering algorithm; the K-means clustering algorithm includes selecting a clustering initial point; wherein the clustering initial point is determined by a K-means++ algorithm.

5. The power data privacy protection method according to claim 1, characterized in that: The overall loss function of the feature analysis model includes the loss function of the mutual information algorithm and the loss function of the Transformer model; The feature analysis model is further configured to: in response to determining that the overall loss function is minimized, terminate the Transformer model training.

6. The power data privacy protection method according to claim 5, characterized in that: The loss function of the mutual information algorithm is determined based on the average value of the mutual information of each attribute feature in the training feature subset.

7. The power data privacy protection method according to claim 1, characterized in that: The training attribute feature set is configured as a feature variable in the mutual information algorithm; the training candidate feature set is configured to associate the joint probability distribution and the marginal probability distribution in the mutual information algorithm; and the privacy level is configured as a target variable in the mutual information algorithm.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein: When executing the computer program, the processor implements the power data privacy protection method according to any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the power data privacy protection method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Differential privacy data publishing method meeting personalized privacy budget allocation

    CN114491644A

  • Energy big data privacy protection method based on privacy level

    CN117216796A

  • Claim settlement data information desensitization acquisition method, device and system

    CN117454426A