A method for protecting the privacy of power customer data based on Fourier low-energy coefficients
By applying Fourier low-energy coefficient transformation and covariance matrix distribution perturbation to power customer data, the problems of insufficient privacy protection and poor curve change rate in traditional methods are solved, thus achieving secure protection of power customer data and availability for data mining.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-05
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for protecting the privacy of electricity customer data are ineffective in maintaining the rate of change of the curve, and traditional methods are not sufficiently privacy-protecting, making it difficult to effectively protect the privacy of electricity customer data and maintain its mining availability in untrusted environments.
A Fourier transform based on low-energy coefficients is used to perform Fourier transformation on power customer data, select low-energy coefficients for small-amplitude perturbation, and generate a perturbed dataset by random perturbation of the covariance matrix distribution. This prevents third parties from reconstructing the real dataset while maintaining the availability of the data change rate within the dataset.
This approach achieves the goal of protecting the privacy of electricity customer data while maintaining the availability of the data change rate for each electricity customer data record within the dataset, thereby improving the privacy and security of electricity customer data and preventing third parties from reconstructing the actual dataset.
Smart Images

Figure CN114021189B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power customer data privacy protection technology, and relates to a method for power customer data privacy protection based on Fourier low energy coefficient. Background Technology
[0002] The continuous development and breakthroughs in power technology have greatly promoted the accumulation of power customer data and the diversification of data structures, thus significantly increasing the value of mining large volumes of diverse data. However, with the development of power customer data mining applications, privacy protection issues have arisen. How to protect the privacy of multi-source power customer datasets in untrusted environments while maintaining their usability for data mining has become a crucial problem.
[0003] Electricity customer data largely involves numerical data, and methods for its privacy protection and mining are mostly based on data distortion techniques. Data distortion techniques maintain the usability of certain aspects of the data, such as the rate of change of curves or clustering characteristics, by perturbing the original data, ensuring that sensitive information in the original data is not disclosed. Data privacy strength refers to the difficulty for an attacker with background knowledge to reconstruct the original dataset from the published dataset.
[0004] Analyzing the rate of change in a company's electricity consumption curve is of great reference value for determining whether there are disturbances in the power grid, stabilizing the system state, and rationally setting electricity prices.
[0005] The most common method to maintain the rate of change of the curve is to process the original data by adding noise and swapping. The processed data is the same as the original data in some statistically relevant properties, and can produce sufficiently similar results in subsequent data mining operations.
[0006] To maintain the rate of change of the curve, some existing traditional privacy protection methods based on data distortion employ data exchange or matrix calculation, which have a certain degree of reversibility but are insufficient in protecting data privacy; others have good privacy protection effects but are not ideal in maintaining the rate of change of the curve. Summary of the Invention
[0007] To address the shortcomings of existing technologies, this application provides a method for protecting the privacy of electricity customer data based on Fourier low-energy coefficients. By utilizing the non-reconfigurability caused by perturbations in the forward and inverse Fourier coefficients and the energy concentration characteristics of the Fourier transform, a small perturbation is applied to the low-energy coefficients in the Fourier coefficients. This method maintains the privacy of individual electricity customer data while preserving the availability of the rate of change of each electricity customer data record in the dataset after the perturbation.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] A method for protecting the privacy of electricity customer data based on the Fourier coefficient low energy factor includes the following steps:
[0010] Step 1: For a given power customer dataset D, set the Fourier low energy coefficient percentage α and the number of groups G. The power customer dataset is a multidimensional numerical dataset.
[0011] Step 2: Perform a Fourier transformation on each electricity customer data record in dataset D, and save the transformed Fourier coefficient sequence of all electricity customer data records to set F;
[0012] Step 3: Sort the Fourier coefficient sequence of each power customer data record in set F in descending order, and select the coefficients of the last α part to form the low energy coefficient sequence of the power customer data record;
[0013] Step 4: Sort the set of low energy coefficient sequences of all power customer data records in descending order according to the first value of each sequence, and divide the set of low energy coefficient sequences into G groups evenly according to the sorting order;
[0014] Step 5: Within each group, use the mean vector of each group and the unbiased estimate of the true data to perform random perturbation based on the distribution of the covariance matrix;
[0015] Step 6: Perform an inverse transform on the perturbed Fourier coefficients to generate the perturbed dataset and publish it.
[0016] The present invention further includes the following preferred embodiments:
[0017] Preferably, in step 1, given an electricity customer dataset D, the number of its attributes is m, and there are n electricity customer data records in D;
[0018] Set the percentage of low energy coefficient α, where α ∈ (0, 50%);
[0019] Let the number of groups be G, G∈[2,|n / 2|], where G is a positive integer, and |n / 2| represents the integer part of n / 2.
[0020] Preferably, in step 2, a one-dimensional discrete Fourier transform is performed on the m attribute values of each power customer data record in dataset D to obtain the corresponding m Fourier coefficients;
[0021] Specifically, the i-th electricity customer data record is transformed to obtain the Fourier coefficient sequence f. i1 ,f i2 ,…,f im ;
[0022] Save the transformed Fourier coefficient sequence of all electricity customer data records to set F for filtering low-energy coefficients, i.e., set F contains coefficient f. 11 ,f 12 ,…,f 1m ,…,f n1 ,f n2 ,…f nm .
[0023] Preferably, in step 2, a one-dimensional discrete Fourier transform is performed on the m attribute values of each power customer data record in the dataset D. Specifically, FFT or DCT is used according to the requirements for conversion efficiency and energy concentration to obtain the corresponding m Fourier coefficients.
[0024] Preferably, in step 2, the original positions of each coefficient are also recorded and stored for subsequent reset.
[0025] Preferably, in step 4, the low-energy coefficient sequence set [FL1,…,FL] of all electricity customer data records is used. n The data is then sorted in descending order based on the first low-energy coefficient value of each electricity customer data record, resulting in the sequence set FL = [FL1′, ..., FL1′]. n ′).
[0026] Preferably, if n is divisible by G, then FL is divided into G groups.
[0027] Otherwise, divide the sequences in FL with indices before (n-(n mod G)) into G groups, and then add the remaining sequences to the last group.
[0028] Preferably, in step 4, the low-energy coefficient sequence is sorted again and divided into G groups, wherein the number of power customer data records in the g-th group (1≤g≤G) is:
[0029]
[0030] Preferably, in step 5, assuming that the number of low-energy coefficients selected in each electricity customer data record in step 3 is J, FL(k) is the overall sequence of the k-th (1≤k≤J) attribute in the sequence set FL, that is, FL(k)={f 1k ,f 2k ,…,f nk};
[0031] Calculate the sample covariance matrix S among the overall sequences of each attribute. F ;
[0032] Calculate the mean of each attribute sequence within each group to obtain the mean vector for that group. and the overall mean vector matrix of each attribute
[0033] according to Calculate the within-group sample covariance matrix
[0034] Use the sample covariance matrix S between the overall sequences of each attribute F and within-group sample covariance matrix Calculate the unbiased estimate S of the covariance matrix of the real data. G ;
[0035] Within a group, based on the mean vector of each group The unbiased estimate of the covariance matrix of the real data, S G The distribution generates perturbation values;
[0036] By replacing the original coefficients at the corresponding positions, random perturbations based on the distribution of the covariance matrix are achieved.
[0037] Preferably, the step of basing the mean vector of each group on the mean vector within the group... The unbiased estimate of the covariance matrix of the real data, S G The distribution generates perturbation values, specifically:
[0038] Generate conformance The perturbation values of the multidimensional normal distribution:
[0039] First, generate an arbitrary G-dimensional normal distribution X containing all elements with low energy coefficients, according to S G Calculate the linear transformation matrix C, and finally generate the perturbation value matrix. Where E is the G-dimensional identity matrix.
[0040] Preferably, in step 5, S F The middle element is It is the covariance between FL(j) and FL(k);
[0041] Where j and k represent the sets of low-energy coefficient sequences [FL1,…,FL], respectively. n Any two attributes of ];
[0042] FL(j) and FL(k) are the overall sequences of the j-th (1≤j≤J) and k-th (1≤k≤J) attributes in FL, respectively;
[0043]
[0044] This represents the sequence of within-group means for the k-th attribute value, i.e.
[0045] The middle element is for The covariance of the j-th (1≤j≤J) and k-th (1≤k≤J) attributes;
[0046] S G The middle element is
[0047] Preferably, step 6 specifically includes:
[0048] Step 6.1: According to the original position of the coefficients corresponding to the power customer data records in Step 2, restore the coefficients that have been reordered and replaced in Steps 3-5 to their original positions to obtain the coefficient set F';
[0049] Step 6.2: Perform an inverse discrete Fourier transform on the coefficient set F' to obtain the dataset D*;
[0050] Step 6.3: Submit D* to the third-party server that provides data mining services.
[0051] The beneficial effects achieved by this application are:
[0052] This invention provides privacy protection for power customer numerical multidimensional datasets based on the (α,G)-grouped perturbation data publishing using Fourier low-energy coefficients. It achieves data privacy protection processing and maintains the rate of change characteristics of data curves in a B / S (Browser / Server) architecture. This invention allows users to control the parameters used in the privacy processing. The server employs a method of replacing the low-energy coefficients of the Fourier transform with random perturbation data based on a grouped normal distribution, preventing third parties from reconstructing the actual dataset and improving the security of power customer data privacy. Attached Figure Description
[0053] Figure 1 This is a flowchart of the method of the present invention;
[0054] Figure 2 This is a flowchart illustrating the method of the present invention;
[0055] Figure 3 This is a schematic diagram of a multidimensional numerical dataset according to an embodiment of the present invention;
[0056] Figure 4 This is a schematic diagram of the converted Fourier coefficient set according to an embodiment of the present invention;
[0057] Figure 5 This is a schematic diagram of the low-energy coefficient of an embodiment of the present invention;
[0058] Figure 6 This is a schematic diagram of low-energy coefficient grouping according to an embodiment of the present invention;
[0059] Figure 7This is a schematic diagram of the low-energy coefficient random perturbation results in an embodiment of the present invention;
[0060] Figure 8 This is a schematic diagram of the Fourier coefficient set after replacing the low-energy coefficients in an embodiment of the present invention;
[0061] Figure 9 This is a schematic diagram of the published dataset according to an embodiment of the present invention. Detailed Implementation
[0062] The present application will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and should not be construed as limiting the scope of protection of the present application.
[0063] like Figure 1-2 As shown, the present invention provides a method for protecting the privacy of electricity customer data based on the Fourier coefficient, comprising the following steps:
[0064] Step 1: For a given power customer dataset D, set the Fourier low energy coefficient percentage α and the number of groups G. The power customer dataset is a multidimensional numerical dataset.
[0065] In practice, given a power customer dataset D, the number of attributes is m, and there are n power customer data records in D;
[0066] Set the percentage of low energy coefficient α, where α ∈ (0, 50%);
[0067] Let the number of groups be G, G∈[2,|n / 2|], where G is a positive integer, and |n / 2| represents the integer part of n / 2.
[0068] Furthermore, the electricity customer data record can be any numerical data, such as a record of an enterprise's electricity consumption over m months, where the first to m attributes correspond to the monthly electricity consumption over the consecutive m months.
[0069] Step 2: Perform a Fourier transformation on each electricity customer data record in dataset D, and save the transformed Fourier coefficient sequence of all electricity customer data records to set F;
[0070] In practice, a one-dimensional discrete Fourier transform is performed on the m attribute values of each power customer data record in dataset D. Depending on the requirements for conversion efficiency and energy concentration, any one of the transformation methods, FFT (Fast Fourier Transform) or DCT (Discrete Cosine Transform), can be used to obtain the corresponding m Fourier coefficients.
[0071] For example, transforming the i-th electricity customer data record yields the Fourier coefficient sequence f. i1 ,f i2 ,…,f im ;
[0072] Save the transformed Fourier coefficient sequence of all electricity customer data records to set F for filtering low-energy coefficients, i.e., set F contains coefficient f. 11 ,f 12 ,…,f 1m ,…,f n1 ,f n2 ,…f nm The power customer data records and stores the original positions of each coefficient for use in the reset in step 6;
[0073] One-dimensional discrete cosine transform (DCT) is a type of Fourier transform and a classic method in the field. DCT can be applied to real number sequences f(0), f(2), ..., f(n-1) to generate real coefficient sequences F(0), F(1), ..., F(n-1). The transformation process is represented by the following equation:
[0074]
[0075] Step 3: Sort the Fourier coefficient sequence of each power customer data record in set F in descending order, and select the coefficients of the α (α is a percentage) part to form the low energy coefficient sequence of the power customer data record;
[0076] For the Fourier coefficients of the same electricity customer data record in set F (i.e., coefficients with the same first index), sort them in descending order and select the coefficients that are α (where α is a percentage) after the total number of coefficients to form the low-energy coefficient sequence FL of that electricity customer data record. i (1≤i≤n, where n is the total number of electricity customer data records).
[0077] Introducing the concept of low energy coefficient, in step 3, the coefficient of the last α part in the Fourier coefficient sequence of each sorted power customer data record is defined as the low energy coefficient.
[0078] The trend of Fourier transform is to concentrate energy in a few high-energy coefficients, and after making small perturbations to the low-energy coefficients, the difference between the inverse Fourier transform power customer data record and the original power customer data record can be kept within a certain range.
[0079] Step 4: Sort the set of low energy coefficient sequences of all power customer data records in descending order according to the first value of each sequence, and divide the set of low energy coefficient sequences into G groups evenly according to the sorting order;
[0080] In practice, the low-energy coefficient sequence set [FL1,…,FL] of all electricity customer data records will be used. n The data is then sorted in descending order based on the first low-energy coefficient value of each electricity customer data record to obtain the sequence set.
[0081] FL=[FL1′,…,FL n ′];
[0082] If n is divisible by G, then FL is divided into G groups.
[0083] Otherwise, divide the sequences in FL with indices before (n-(n mod G)) into G groups, and then add the remaining sequences to the last group;
[0084] In step 4, the low-energy coefficient sequence is sorted again and divided into G groups. The number of power customer data records in the g-th group (1≤g≤G) is:
[0085]
[0086] Step 5: Within each group, use the mean vector of each group and the unbiased estimate of the true data to perform random perturbation based on the distribution of the covariance matrix;
[0087] In practical implementation, assuming that the number of low-energy coefficients selected in each power customer data record in step 3 is J, and FL(k) is the overall sequence of the k-th (1≤k≤J) attribute in the sequence set FL, that is, FL(k)={f 1k ,f 2k ,…,f nk};
[0088] Calculate the sample covariance matrix S among the overall sequences of each attribute. F FL(j) and FL(k) are the overall sequences of the j-th (1≤j≤J) and k-th (1≤k≤J) attributes in FL, respectively. F Middle elements FL(k)) is the covariance between FL(j) and FL(k);
[0089] Calculating the mean of each attribute sequence within each group yields the mean vector for that group. Specifically, the mean vector for the g-th group (1 ≤ g ≤ G) can be represented as:
[0090] Another result is the overall mean vector matrix of each attribute. This represents the sequence of within-group means for the k-th attribute value, i.e.
[0091] according to Calculate the within-group sample covariance matrix Middle elements for The covariance of the j-th (1≤j≤J) and k-th (1≤k≤J) attributes;
[0092] Using the overall sample covariance matrix S F and within-group sample covariance matrix Calculate the unbiased estimate S of the covariance matrix of the real data. G S G Middle elements Where j and k represent the sets of low-energy coefficient sequences [FL1,…,FL], respectively. n Any two attributes of ]; S is the sample covariance matrix among the overall sequences of each attribute. F The corresponding element in; Within-group sample covariance matrix The corresponding element;
[0093] Within a group, based on the average vector of each group. The unbiased estimate of the covariance matrix S of the real data G The generated perturbation value of the distribution is used to replace the original low-energy coefficient at the corresponding position.
[0094] For example: generating a mean vector conforms to The covariance matrix conforms to S G The perturbation matrix of a multidimensional normal distribution.
[0095] By the known theorem: For an n-dimensional random variable X that follows a normal distribution N(μ,B), if an m-dimensional random variable Y is a linear transformation of X, i.e., Y = XC, where C is an n×m matrix, then Y follows a normal distribution N(μC,C). T BC).
[0096] According to the above theorem, an n-dimensional normal sample with a covariance matrix E can be transformed into a C-dimensional normal sample by a linear transformation. T C is an n-dimensional normal sample. For an n-dimensional normal distribution matrix with covariance R, where R is a positive definite matrix, the Cholesky decomposition R = C is possible. T C.
[0097] The present invention can generate a perturbation value matrix according to the following steps:
[0098] Generate an arbitrary G-dimensional normal distribution matrix X containing the total number of low-energy coefficient elements;
[0099] According to S G Calculate the linear transformation matrix C; for a given covariance S G The multidimensional normal distribution matrix, S GS is a positive definite matrix, and can be decomposed into Cholesky matrix. G =C T C.
[0100] Generate perturbation value matrix Where E is the G-dimensional identity matrix.
[0101] Step 6: Perform an inverse transform on the perturbed Fourier coefficients to generate the perturbed dataset and publish it. This includes:
[0102] Step 6.1: According to the original position of the coefficients recorded in Step 2, restore the coefficients that have been reordered and replaced in Steps 3-5 to their original positions to obtain the coefficient set F';
[0103] Step 6.2: Perform an inverse discrete Fourier transform (such as IDCT) on the coefficient set F' to obtain the dataset D*;
[0104] IDCT can be applied to real coefficient sequences F(0), F(1), ..., F(n-1) to generate real number sequences:
[0105] f(0), f(2), ..., f(n-1);
[0106] The transformation process is performed according to the following formula:
[0107]
[0108] Step 6.3: Submit D* to the third-party server that provides data mining services.
[0109] Example
[0110] The present invention provides a method for protecting the privacy of electricity customer data based on the Fourier coefficient, comprising the following steps:
[0111] Step 1: As Figure 3 As shown, a multidimensional numerical dataset D0 of a certain electricity customer consists of n=5 electricity customer data records, and each electricity customer data record contains 6 attribute values.
[0112] It can be any data with numerical attributes, such as 6 attributes that record the company's electricity consumption over 6 consecutive months.
[0113] Suppose a user sends a privacy protection request to the browser, with privacy protection constraints G=2 and α=70%.
[0114] Numerical multidimensional datasets can typically be represented using two-dimensional matrices.
[0115] Steps 2-4: Server privacy protection processing steps are as follows:
[0116] (1) The server will process dataset D0 according to the formula
[0117]
[0118] Perform a one-dimensional DCT transform, and represent the obtained Fourier coefficients using a two-dimensional matrix F, such as... Figure 4 As shown;
[0119] (2) Sort the coefficients in F. Based on α = 30%, the low-energy coefficients are located at the th position in the sorted coefficient sequence.
[0120] Following the column, the selected low-energy coefficient FL is as follows: Figure 5 As shown, the original locations of each low-energy coefficient are stored;
[0121] (3) Sort the values of the low-energy coefficients in each row in descending order according to the maximum value of each row, and group them according to G=2.
[0122] (n-(n mod G)) / G=2
[0123] n mod G = 5 mod 2 = 1 to obtain the following: Figure 6 The groups shown are:
[0124] Step 5: From Figure 6 calculate:
[0125] Mean vector matrix
[0126] population sample covariance matrix
[0127] Within-group sample covariance matrix
[0128] Unbiased estimate of the covariance matrix of real data
[0129] Within a group, the average vector based on each group The unbiased estimate of the covariance matrix S of the real data G The perturbation values generated by the multivariate normal distribution are as follows Figure 7 As shown;
[0130] Use according to the original position of each stored low-energy coefficient. Figure 7 Random perturbation data replacement Figure 4 Some of the original coefficients, the results are as follows Figure 8 ;
[0131] Step 6: [Regarding...] Figure 8 The Fourier coefficients are calculated according to the formula Performing the inverse DCT transform, we obtain the perturbed published dataset as follows: Figure 9 As shown, it is difficult for third parties to reconstruct the real dataset from the published dataset, and the rate of change of the curve of the published dataset remains basically unchanged compared with the real dataset.
[0132] As can be seen, this invention utilizes the non-reconfigurability caused by perturbations in the forward and inverse Fourier coefficients, as well as the energy concentration characteristics of the Fourier transform, to maintain the availability of the data change rate of each power customer data record after perturbation by making small perturbations to the low-energy coefficients in the Fourier coefficients, while maintaining the privacy of individual power customer data.
[0133] This invention addresses privacy protection applications for numerical multidimensional datasets from electricity customers, implementing data privacy protection processing and maintaining the rate of change characteristics of data curves in a B / S (Browser / Server) architecture. The invention allows users to control the parameters used in the privacy processing. The server employs a method of replacing low-energy coefficients of the Fourier transform with randomly perturbed data based on a grouped normal distribution, preventing third parties from reconstructing the actual dataset and improving the security of electricity customer data privacy.
[0134] The applicant of this invention has provided a detailed description of the embodiments of the invention in conjunction with the accompanying drawings. However, those skilled in the art should understand that the above embodiments are merely preferred embodiments of the invention. The detailed description is only intended to help readers better understand the spirit of the invention and is not intended to limit the scope of protection of the invention. On the contrary, any improvements or modifications made based on the inventive spirit of the invention should fall within the scope of protection of the invention.
Claims
1. A method for protecting the privacy of electricity customer data based on the Fourier coefficient, characterized in that: The method includes the following steps: Step 1: Given a power customer dataset D with m number of attributes and n number of power customer data records in D; for the given power customer dataset D, set the Fourier low-energy coefficient percentage G, the power customer dataset is a multi-dimensional numerical dataset; Step 2: Perform a Fourier transform on each electricity customer data record in dataset D, and save the transformed Fourier coefficient sequence of all electricity customer data records to set F; specifically, perform a one-dimensional discrete Fourier transform on the m attribute values of each electricity customer data record in dataset D to obtain the corresponding m Fourier coefficients; among them, for the th Transform the data records of each electricity customer to obtain the Fourier coefficient sequence. Save the transformed Fourier coefficient sequence of all electricity customer data records to set F for filtering low-energy coefficients; that is, set F contains coefficients. ; Step 3: Sort the Fourier coefficient sequence of each electricity customer data record in set F in descending order, and then select... The coefficients in this part constitute the low-energy coefficient sequence of this electricity customer data record; Step 4: Sort the low-energy coefficient sequence set of all electricity customer data records in descending order according to the first value of each sequence, and divide the low-energy coefficient sequence set evenly into G groups according to the sorting order; specifically, this involves the low-energy coefficient sequence set of all electricity customer data records. Sort the data in descending order based on the first low-energy coefficient value of each electricity customer data record to obtain the sequence set. ; Step 5: Within each group, use the mean vector of each group and the unbiased estimate of the true data to perform random perturbation based on the distribution of the covariance matrix; Assume that in step 3, the number of low-energy coefficients selected in each electricity customer data record is J, and FL(k) is the overall sequence of the k-th (1≤k≤J) attribute in the sequence set FL, i.e., FL(k) = { }; Calculate the sample covariance matrix among the overall sequences of each attribute. Calculate the mean of each attribute sequence within each group to obtain the mean vector for that group. (1≤g≤G) and the overall mean vector matrix of each attribute ;according to Calculate the within-group sample covariance matrix Use the sample covariance matrix between the overall sequences of each attribute. and within-group sample covariance matrix Calculate the unbiased estimate of the covariance matrix of the real data Within a group, based on the mean vector of each group. (1≤g≤G) and the unbiased estimate matrix of the covariance matrix of the real data The distribution generates perturbation values, which replace the original coefficients at the corresponding positions, thus realizing random perturbation based on the distribution of the covariance matrix; Step 6: Perform an inverse transform on the perturbed Fourier coefficients to generate the perturbed dataset and publish it.
2. The method for protecting the privacy of electricity customer data based on the Fourier low energy coefficient according to claim 1, characterized in that: In step 1, set the percentage of low energy coefficient. , ∈(0, 50%); Let the number of groups be G, G∈[2,|n / 2|], where G is a positive integer, and |n / 2| represents the integer part of n / 2.
3. The method for protecting the privacy of electricity customer data based on the Fourier low energy coefficient according to claim 1, characterized in that: In step 2, a one-dimensional discrete Fourier transform is performed on the m attribute values of each power customer data record in dataset D. Specifically, FFT or DCT is used according to the requirements for conversion efficiency and energy concentration to obtain the corresponding m Fourier coefficients.
4. The method for protecting the privacy of electricity customer data based on the Fourier low energy coefficient according to claim 1, characterized in that: In step 2, the original positions of each coefficient are also recorded and stored for subsequent reset.
5. The method for protecting the privacy of electricity customer data based on the Fourier low energy coefficient according to claim 1, characterized in that: If n is divisible by G, then FL They were all divided into Group G; Otherwise, for FL The sequences with indices before (n-(n mod G)) are divided into G groups, and the remaining sequences are added to the last group.
6. A method for protecting the privacy of electricity customer data based on the Fourier low energy coefficient according to claim 1 or 5, characterized in that: In step 4, the low-energy coefficient sequence is sorted again and divided into G groups, where the first group is the lowest energy coefficient. The number of electricity customer data records in the group is: num= .
7. The method for protecting the privacy of electricity customer data based on the Fourier low energy coefficient according to claim 1, characterized in that: The mean vector within each group (1≤g≤G) and the unbiased estimate matrix of the covariance matrix of the real data The distribution generates perturbation values, specifically: Generate a conforming ( , The perturbation value of the multidimensional normal distribution: First, generate an arbitrary G-dimensional normal distribution X containing all elements with low energy coefficients, according to... Calculate the linear transformation matrix C, and finally generate the perturbation value matrix. =X*C+E* ,in Let E be the dataset, and E be the G-dimensional identity matrix.
8. The method for protecting the privacy of electricity customer data based on the Fourier low energy coefficient according to claim 1, characterized in that: In step 5, The middle element is =s(FL(j), FL(k) is the covariance between FL(j) and FL(k); Where j and k represent the sets of low-energy coefficient sequences, respectively. Any two attributes; FL(j) and FL(k) are the overall sequences of the j-th (1≤j≤J) and k-th (1≤k≤J) attributes in FL, respectively; ={ }; = , (1≤k≤J) represents the sequence of within-group mean values for the k-th attribute value, i.e. ={ }; The middle element is ,for The covariance of the j-th (1≤j≤J) and k-th (1≤k≤J) attributes; The middle element is .
9. A method for protecting the privacy of electricity customer data based on the Fourier low energy coefficient according to claim 1, characterized in that: Step 6 specifically includes: Step 6.1: According to the original position of the coefficients corresponding to the electricity customer data records in Step 2, restore the coefficients that have been reordered and replaced in Steps 3-5 to their original positions to obtain the coefficient set. ; Step 6.2: For the coefficient set Perform an inverse discrete Fourier transform to obtain the dataset D*; Step 6.3: Submit D* to the third-party server that provides data mining services.
Citation Information
Patent Citations
Remote meter reading method and device based on power line
CN101751766A
Method and device for performing multi-party joint dimension reduction processing on user privacy data
CN110889139A