Differential privacy data de-identification method based on time-series random mapping and related device
Patent Information
- Application Number
- CN202211734104.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2042-12-30
AI Technical Summary
随着电力现货市场建设推进,市场主体多元化趋势愈加明显,各类市场主体对信息披露要求越来越高,然而目前电力市场中对数据安全防护的相关研究较少,针对现有数据没有建立其合理的数据脱敏方法,无法对电力交易中心的时序数据进行合理防护,这将限制未来电力市场的进一步发展,市场主体数据也将面临泄露风险,进而扰乱市场运行
[0012]1)采用基于DTW的k-medoids聚类,将庞大的时序数据集分为几个类簇,从而更好的区分数据特征,并基于此选用不同的隐私保护参数进行差异化的数据脱敏处理;2)采用基于差分隐私的时序随机化映射的方法,考虑从时序特征的角度保护时序数据的隐私安全,一方面可以通过差分隐私的框架可以实现隐私性与数据可用性的权衡,另一方面扰乱时序的方法更适合于交易数据的脱敏发布;3)时序随机化映射方法不会给用电的时序数据带来独立噪声,减少了由于引入数字噪声带来的尖峰波动,不会让扰乱后的电力市场时序数据失去本身的数据价值。
Smart Images

Figure CN116011011B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a differential privacy data desensitization method, apparatus, electronic device, and computer-readable storage medium based on time-series random mapping. Background Technology
[0002] In recent years, the power industry has attached great importance to data security. Business departments have implemented a series of data protection measures using technologies such as data access control, data anonymization, watermarking, and encryption. However, risks remain regarding data exchange and external transmission, including incomplete auditing and uncontrollable data security. With the development of the power market, key initiatives such as power market pilot programs have achieved significant breakthroughs in several provinces (regions). Various provinces have successively issued local market rules, and market-based transactions have been intensively carried out, forming a multi-cycle, multi-product trading system. The power market information flowing within the market is a crucial foundation supporting its development, but this data is currently subject to serious privacy leaks. On the one hand, data exported from the power operation system during data anonymization lacks security protection measures such as obfuscation and watermarking, increasing the possibility of exposure to sensitive customer information. Once data is leaked, it is difficult to quickly locate and trace the source. On the other hand, the power operation system contains a large amount of important data, including highly sensitive personal information of electricity users, covering a wide range of areas. This data is shared and distributed to various professional departments, grassroots units, and external partners within the power grid company, providing business applications such as big data analysis and decision support. Once a data leak occurs, there are problems with effectively locating the source of the leak and tracing security responsibility. The operations department needs to take appropriate security control measures to prevent security risks in the data sharing process.
[0003] Currently, electricity market information disclosure is strictly carried out through the electricity market information disclosure platform in accordance with the requirements of documents issued by energy regulatory agencies. Due to the large variety and volume of data involved, the disclosure must ensure data validity and privacy. With the advancement of the electricity spot market, the diversification of market participants is becoming increasingly apparent, and various market participants have higher and higher requirements for information disclosure. However, there is currently limited research on data security protection in the electricity market, and no reasonable data anonymization methods have been established for existing data. This makes it impossible to adequately protect the time-series data of the electricity trading center, which will limit the further development of the electricity market in the future. Market participant data will also face the risk of leakage, thereby disrupting market operations. Summary of the Invention
[0004] The purpose of this application is to provide a differential privacy data desensitization method, apparatus, electronic device, and computer-readable storage medium based on time-series random mapping.
[0005] To achieve the above objectives, this application provides a differential privacy data anonymization method based on temporal random mapping in its first aspect. The method includes: acquiring various target time-series data that need to be anonymized by a power trading center, and clustering the target time-series data using a k-medoids clustering algorithm based on dynamic time curvature distance metric to obtain time-series datasets of different clusters; calculating the window size corresponding to different target time-series data within the same cluster based on the maximum compression ratio of the target time-series data, determining the average compression ratio of all target time-series data, and using the average compression ratio as the target window size set during anonymization; determining the privacy budget required for differential privacy based on the predetermined security protection requirement level of the target time-series data, and processing the target time-series data according to the target window size and the differential privacy anonymization method of temporal random mapping to obtain time-anonymized power market entity data.
[0006] To achieve the above objectives, this application provides a differential privacy data desensitization device based on time-series random mapping in a second aspect. The device includes: a clustering processing unit, used to acquire various target time-series data that need to be desensitized by the power trading center, and to cluster the target time-series data using a k-medoids clustering algorithm based on dynamic time curvature distance metric to obtain time-series datasets of different clusters; a target window size determination unit, used to calculate the window size corresponding to different target time-series data in the same cluster based on the maximum compression ratio of the target time-series data, and to determine the average compression ratio of all target time-series data, using the average compression ratio as the target window size set during desensitization; and a desensitization processing unit, used to determine the privacy budget required for differential privacy based on the predetermined security protection requirement level of the target time-series data, and to process the target time-series data according to the target window size and the differential privacy desensitization method of time-series random mapping to obtain time-series desensitized power market entity data.
[0007] To achieve the above objectives, this application provides an electronic device in a third aspect, the electronic device comprising:
[0008] Memory, used to store computer programs;
[0009] A processor, configured to implement the steps of the differential privacy data desensitization method based on temporal random mapping as described in the first aspect above, when executing a computer program stored in memory.
[0010] To achieve the above objectives, this application provides a computer-readable storage medium in a fourth aspect, on which a computer program is stored, and when executed by a processor, the computer program implements the steps of differential privacy data desensitization based on temporal random mapping as described in the first aspect above.
[0011] The differential privacy data anonymization scheme based on temporal random mapping provided in this application has the following beneficial effects:
[0012] 1) DTW-based k-medoids clustering is used to divide the massive time-series dataset into several clusters, thereby better distinguishing data features and selecting different privacy protection parameters for differentiated data anonymization processing; 2) A time-series randomization mapping method based on differential privacy is adopted to protect the privacy and security of time-series data from the perspective of time-series features. On the one hand, the differential privacy framework can achieve a trade-off between privacy and data availability; on the other hand, the time-series randomization method is more suitable for the anonymized release of transaction data; 3) The time-series randomization mapping method will not introduce independent noise into the electricity consumption time-series data, reducing spike fluctuations caused by the introduction of digital noise, and ensuring that the disturbed electricity market time-series data does not lose its inherent data value.
[0013] The differential privacy data anonymization method based on temporal random mapping provided in this application takes into account the temporal characteristics of power trading data and is more suitable for privacy anonymization of time-series data in power trading centers.
[0014] This application also provides a differential privacy data desensitization device, electronic device, and computer-readable storage medium based on time-series random mapping, which have the above-mentioned beneficial effects, and will not be elaborated here. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0016] Figure 1 A flowchart illustrating a differential privacy data desensitization method based on temporal random mapping provided in this application embodiment;
[0017] Figure 2 A flowchart illustrating another differential privacy data desensitization method based on temporal random mapping provided in this application embodiment;
[0018] Figure 3 This is a schematic diagram illustrating the effect of differential privacy data desensitization based on temporal random mapping, provided in an embodiment of this application.
[0019] Figure 4 This is a schematic diagram showing the difference between and without anonymization of a certain time-series data after undergoing time-series random mapping for privacy data.
[0020] Figure 5This is a structural block diagram of a differential privacy data desensitization device based on temporal random mapping, provided in an embodiment of this application. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] Please see Figure 1 , Figure 1 A flowchart for differential privacy data desensitization based on temporal random mapping provided in this application includes the following steps:
[0023] Step 101: Obtain various target time series data that need to be de-identified by the power trading center, and use the k-medoids clustering algorithm based on dynamic time curvature distance metric to cluster the target time series data to obtain time series datasets of different clusters;
[0024] This step aims to have an entity suitable for executing the differential privacy data anonymization method based on temporal random mapping provided in this application (e.g., a local server or cloud server for data processing and analysis) collect various types of time-series data from the power trading center that require anonymization. Then, a k-medoids clustering algorithm based on DTW (Dynamic Time Warping) distance metric is used to cluster the required time-series data, resulting in time-series datasets of different clusters. Furthermore, the distance similarity of these data varies considerably; therefore, data within the same cluster can be anonymized using the same parameters.
[0025] Specifically, the following steps can be followed:
[0026] 1) Collect various types of time-series data circulating in the power trading center, and perform k-medoids clustering on the time-series data to form several clusters. The clustering uses the cumulative distance of dynamic time warping to measure data similarity. The specific calculation method is as follows:
[0027] a) Assume the lengths of the two sequences are respectively n , m , respectively , First, construct an n The matrix grid of m, matrix elements ( i , j )express and Distance between two points The calculation method uses Euclidean distance:
[0028] .
[0029] b) Find the minimum distance path from the starting point of each sequence to its ending point (i.e., from one end of the diagonal to the other) within a distance matrix. This process requires satisfying three conditions:
[0030] ① Boundary condition: The selected path must start from the bottom left corner and end at the bottom right corner;
[0031] ② Continuity: When selecting a path, it can only align with points adjacent to itself;
[0032] ③ Monotonicity: When selecting a path, the dashed lines cannot intersect.
[0033] c) Find the optimal path that satisfies the above constraints, and use the cumulative distance of this path as the sequence similarity measure, as shown in the following formula:
[0034] ;
[0035] In the formula, Wk represents the path points of the optimal path, K represents the number of path points of the optimal path, and the cumulative calculation means that the cumulative path distance is calculated from the optimal path. Wk serves as the value in the distance matrix, but the minimum value is taken as the final DTW distance.
[0036] 2) Based on the defined sequence similarity, several clusters are formed using the k-medoids clustering method. The steps of the k-medoids clustering method are as follows:
[0037] a) Randomly select k objects as the initial representative objects. The remaining objects are assigned to the nearest cluster based on their distance from the representative object;
[0038] b) Then, non-representative objects are repeatedly used to replace representative objects, and a cost function is calculated to evaluate the average dissimilarity between the object and its reference object. The final clustering is achieved by continuously evaluating the dissimilarity. The calculation formula is shown below:
[0039] ;
[0040] in, Indicated by Clusters centered on a central point p This refers to non-representative objects. k This refers to the initial number of representative objects. The subscript j ranges from 1 to k.
[0041] Step 102: Based on the maximum compression ratio of the target time series data, calculate the window size corresponding to different target time series data in the same cluster, determine the average compression ratio of all target time series data, and use the average compression ratio as the target window size set during desensitization;
[0042] Based on step 101, this step aims to have the aforementioned executing entity calculate the required window size for different time-series data in the same cluster based on the maximum compression ratio of the time-series data, then obtain the average compression ratio of all time-series data, and use the finally obtained average compression ratio as the window size set during desensitization, so that the random perturbation of time series can achieve better results.
[0043] Specifically, clusters of the input time series can be obtained through clustering. Based on the maximum compression ratio of the time series data, dimensionality reduction and compression are performed on the time series data to determine the maximum value of the window during time series segmentation, and thus determine the specific random mapping window size. The window size parameter used for data within the same cluster remains consistent. The calculation method is as follows:
[0044] , ;
[0045] in, F For the frequency of change in time series data, n The number of data points contained in a sequence. Represents the first in the sequence Data points, At the maximum compression ratio, It is a standard constant. This is the allowable error. Generally, it is set as follows: , .
[0046] Step 103: Based on the predetermined security protection requirement level of the target time series data, determine the privacy budget required for differential privacy, and process the target time series data according to the differential privacy desensitization method of target window size and time series random mapping to obtain the electricity market entity data after time series desensitization.
[0047] Building upon step 102, this step aims to have the aforementioned implementing entity determine the privacy budget ε required for differential privacy based on the system-defined time-series data security protection requirement level. Simultaneously, using the window size obtained from clustering calculations, the original time-series data is processed using a time-series random mapping differential privacy desensitization method to obtain time-series desensitized electricity market entity data. This enables the time-series data of the market entity to interact more securely with other market entities.
[0048] Specifically, the following steps can be followed:
[0049] a) Assuming a privacy budget for a user input The maximum privacy budget achievable by the random mapping mechanism is ,and The formula for calculating probability is as follows:
[0050] ;
[0051] Where k represents the window value, and m is the sum of the window value and the threshold. The difference is g(k, 1) = 2 / k.
[0052] b) Initialize the interval for the bisection method as follows: ,in , ;Pick , and by and Get the maximum privacy budget ;if Then take Otherwise proceed to step c).
[0053] c) by and Calculate the maximum privacy budget ,if Then take Otherwise take ;
[0054] d) Judgment Is it true? If not, then proceed by... and Get the maximum privacy budget ,pass and Get the maximum privacy budget Proceed to step (e) to determine the threshold. ;
[0055] e) If Then the threshold ;if Then the threshold The process ends here. Otherwise, proceed to step f).
[0056] f) Let , ,when Then proceed to step g).
[0057] g) Calculation and through new and get . judge If true, then take Otherwise take ; and then judge If the condition is met, then determine the threshold. Otherwise, return to step f).
[0058] Furthermore, based on the data type, the perturbation threshold and window value k are determined. Using a differential privacy data desensitization method based on temporal random mapping, the original time series is randomly mapped forward within a fixed window to form a new time series in the perturbed sequence. The steps are as follows:
[0059] a) Initialize the mapping sequence to empty, select a time series as the original sequence, and set its size to [value missing]. k The window is placed at the beginning of both the original sequence and the mapped sequence;
[0060] b) Calculate the number of bits in the mapped sequence window that are empty at position 1, and denot them as . ;
[0061] c) Comparing numerical values With threshold The size, if The position at the beginning of the original sequence window will with The probability is mapped to any point where the mapping sequence window is empty;
[0062] d) If This is not valid. If the beginning of the mapping sequence is empty at this point, then mapping takes precedence; otherwise, then mapping takes precedence. The probability is mapped to any point where the mapping sequence window is empty.
[0063] For details, please refer to the following: Figure 2 The processing flowchart is shown.
[0064] By applying the differential privacy data anonymization method based on temporal random mapping provided in this embodiment, the following technical effects can be achieved: 1) Using DTW-based k-medoids clustering, the massive time-series dataset is divided into several clusters, thereby better distinguishing data features and selecting different privacy protection parameters for differentiated data anonymization processing; 2) Employing a time-series randomization mapping method based on differential privacy, the privacy and security of time-series data are protected from the perspective of time-series features. On the one hand, the differential privacy framework can achieve a trade-off between privacy and data availability; on the other hand, the time-series disruption method is more suitable for anonymizing and publishing transaction data; 3) The time-series randomization mapping method does not introduce independent noise into the electricity consumption time-series data, reducing spike fluctuations caused by the introduction of digital noise and preventing the disrupted electricity market time-series data from losing its inherent data value. In other words, the differential privacy data anonymization method based on temporal random mapping provided in this application considers the temporal characteristics of electricity transaction data and is more suitable for privacy anonymization of time-series data in electricity trading centers.
[0065] To enhance understanding of the solution provided in this embodiment, an example and appendix are also provided below. Figures 3-4 Further explanation:
[0066] Step 1: The power trading center collects a large amount of time-series data that needs to be anonymized. First, based on the data type, the number of clusters k to be formed is determined. Then, k objects are randomly selected from the collected data as initial cluster centers. The remaining objects are assigned to the nearest cluster based on their DTW distance from the representative object. Next, the total cost of the current cluster centers is calculated using the cost function. Simultaneously, a new cluster center is selected, and its total cost is calculated. If the cost decreases, it indicates that the new center has a better clustering effect, and this new center is selected as the cluster center. Otherwise, if the total cost of replacement is consistently greater than the original, an effective replacement has not been achieved, and the algorithm converges.
[0067] Step 2: Clustering yields k clusters. Since the time series data within each cluster have high similarity, we consider using the same window size for perturbation. We then determine the cluster window size using the maximum compression ratio method. Because each sequence has a corresponding maximum compression ratio, we use an averaging method to calculate the maximum compression ratio of all sequences within a cluster and then average them to obtain the required window size k for the entire cluster. For simplicity, let's assume a time series requiring desensitization is within a cluster with 96 sites, and the calculated in-cluster k value is 10. The sequence is:
[0068] [0.2915 0.2315 0.2307 0.3 0.4765 0.6368 2.5715 1.454 1.4049 1.3902 1.3611 1.2618 1.2393 1.2477 1.3879 1.1453 2.1676 2.199 1.8181 2.0468 2.9576 1.2284 1.2011 1.2161 1.526 0.4397 1.5298 3.3394 2.4967 1.2255 0.3543 0.475 0.3065 0.2979 0.3027 0.3032] 0.7968 1.1543 1.0282 0.3001 0.3532 0.3715 0.36 0.2821 0.285 0.3643 0.3068 0.3927 0.3389 0.3105 0.3682 0.3104 0.4047 1.2753 1.0465 1.8433 0.3161 0.2434 0.2419 0.3674 0.2455 0.2453 0.2517 0.2483 0.315 0.237 0.3668 0.1649 0.2216 1.0217 0.983 1.921 0.1428 0.2134 0.1331 0.133 0.1913 0.1928 0.2065 0.1362 0.1331 0.1332 0.2562 0.2081 0.1328 0.1327 0.1327 0.2525 0.2001 0.133 0.1557 0.1591 0.2173 0.2257 0.1507 0.244).
[0069] Step 3: Determine the privacy budget ε based on the data security protection level, and then perturb the sequence using a time-series data differential privacy desensitization method based on time-series random mapping in the power trading center. Assume that the privacy budget ε = 1 is determined for this sequence through data security protection, and the final threshold is calculated using the aforementioned bisection method. If the value is 3, then according to the random mapping method, a scrambled random sequence can be obtained. However, since the sequence is always mapped forward, there may be empty points in the scrambled sequence during the mapping process. In this case, the average value of the values to the left and right of the empty point is calculated and used as the value of the empty point. Finally, the scrambled sequence can be obtained as follows:
[0070] [0.1526 0.1447 0.2174 0.1628 0.1537 0.2307 0.1558 0.2107 0.2915 0.2315 0.4765 0.3 2.5715 1.4049 0.6368 1.454 1.3611 1.3902 1.1453 1.2618 1.3879 1.2393 1.2477 2.1676 1.8181 1.2011 2.199 2.9576 2.0468 1.526 1.2284 1.2161 1.5298 0.3543 0.4397 0.3065] 3.3394 2.4967 1.2255 0.475 0.7968 0.3027 0.2979 1.1543 0.3032 1.0282 0.3001 0.2821 0.3715 0.3532 0.36 0.3105 0.3643 0.285 0.3068 0.3927 0.3389 0.3682 0.3104 0.4047 1.2753 1.8433 1.0465 0.2483 0.2419 0.3161 0.2434 0.3674 0.2455 0.2517 0.2453 0.3668 0.237 0.315 0.2216 0.1649 1.0217 0.2134 0.1428 0.983 1.921 0.1331 0.1928 0.133 0.2065 0.1913 0.1362 0.1327 0.1331 0.2562 0.1332 0.2081 0.1328 0.2525 0.133 0.1327].
[0071] See another example. Figure 3 The diagram shown is shown in the image.
[0072] Step 4: The final disturbed sequence is the de-identified sequence obtained through the differential privacy de-identification method for time-series data in the power trading center based on time-series random mapping. The actual situation of the two sequences is as follows: Figure 4 As shown. The trading center can directly publish this scrambled sequence as the original sequence after privacy protection, or it can use this method to scramble the data when certain untrustworthy entities inquire about market-related time-series information, preventing them from obtaining accurate market operation information.
[0073] The data anonymization methods provided in any of the above embodiments are based on the concept of localized differential privacy considering time series. They protect the privacy of numerical time-series information flowing in the electricity market according to the assessed security requirement level, including unit bidding information, load information, and generation information. The importance of data anonymization in electricity trading centers is reflected in several aspects: the data of the trading center is directly related to the entire electricity trading process, and some data needs to be effectively disclosed to achieve market transparency; in addition, for some dishonest entities in the electricity trading process, it is necessary to anonymize data during data interaction to prevent these entities from infringing on the privacy information of other entities. User load information, generation information, and bidding information have strong time series characteristics; the information they contain is not only reflected in numerical values but also in their temporal order. When protecting this type of data, it is necessary to pay attention not only to the numerical information of the sequence but also to the information contained in the temporal sequence.
[0074] Accordingly, this application proposes a differential privacy data desensitization method based on temporal random mapping. By disrupting the temporal information positions, it hides sensitive information in time-series data, providing an effective method for desensitization processing before information disclosure by power trading centers. At the same time, it can serve as a penalty for dishonest parties to obtain data, regulate the behavior of market participants in the trading process, and act as a safeguard for the data security of trading centers.
[0075] Due to the complexity of the situation, it is impossible to list and elaborate on them all. Those skilled in the art should realize that there are many examples based on the basic method principles provided in this application and in combination with actual situations. Without sufficient creative effort, they should all be within the protection scope of this application.
[0076] Please see below. Figure 5 , Figure 5 This application provides a structural block diagram of a differential privacy data desensitization device 500 based on temporal random mapping. This embodiment is a device embodiment corresponding to the above method embodiment. The differential privacy data desensitization device 500 based on temporal random mapping may include:
[0077] Clustering processing unit 501 is used to acquire various target time series data that need to be desensitized by the power trading center, and to cluster the target time series data using the k-medoids clustering algorithm based on dynamic time curvature distance metric to obtain time series datasets of different clusters.
[0078] The target window size determination unit 502 is used to calculate the window size corresponding to different target time series data in the same cluster based on the maximum compression ratio of the target time series data, and determine the average compression ratio of all target time series data, and use the average compression ratio as the target window size set during desensitization.
[0079] The desensitization processing unit 503 is used to determine the privacy budget required for differential privacy based on the predetermined security protection requirement level of the target time series data, and to process the target time series data according to the differential privacy desensitization method of target window size and time series random mapping, so as to obtain the power market entity data after time series desensitization.
[0080] In some other embodiments of this example, the clustering processing unit 501 may be further configured to:
[0081] Randomly select k objects as the initial cluster centers;
[0082] Non-cluster centroids are repeatedly used to replace the initial cluster centroids for clustering the target time series data;
[0083] Calculate the cost function for each cluster center, and determine the final cluster center based on the dissimilarity determined by each cost function.
[0084] In some other embodiments of this example, the target window size determination unit 502 may be further configured to:
[0085] Calculate using the following formula:
[0086] , ;
[0087] in, F For the frequency of change in time series data, n The number of data points contained in a sequence. Represents the first in the sequence Data points, At the maximum compression ratio, It is a standard constant. Allowed error. Default. , .
[0088] In some other embodiments of this example, the desensitization processing unit 503 may be further configured to:
[0089] Initialize the perturbation sequence to empty, and select a time series as the original sequence, with a size of [missing information]. k The window is placed at the beginning of the original sequence and the beginning of the scrambled sequence;
[0090] Calculate the number of bits in the mapped sequence window that are empty at position 1, and denote them as . ;
[0091] Compare values With threshold The size, if The position at the beginning of the original sequence window will be The probability is mapped to any point in the scrambled sequence window where the value is empty;
[0092] like If the condition is not met and the beginning of the window that disrupts the sequence is empty, then mapping is prioritized; if... If this condition is not met and the beginning of the window that disrupts the sequence is not empty, then... The probability is mapped to any site where the perturbation sequence window is empty.
[0093] This embodiment exists as a device embodiment corresponding to the above method embodiment.
[0094] By applying the differential privacy data anonymization device based on temporal random mapping provided in this embodiment, the following technical effects can be achieved: 1) Using DTW-based k-medoids clustering, the massive time-series dataset is divided into several clusters, thereby better distinguishing data features and selecting different privacy protection parameters for differentiated data anonymization processing; 2) Employing a time-series randomization mapping method based on differential privacy, considering the protection of time-series data privacy from the perspective of time-series features, on the one hand, the differential privacy framework can achieve a trade-off between privacy and data availability, and on the other hand, the time-series disruption method is more suitable for anonymizing and publishing transaction data; 3) The time-series randomization mapping method does not introduce independent noise into the electricity consumption time-series data, reducing spike fluctuations caused by the introduction of digital noise, and preventing the disrupted electricity market time-series data from losing its inherent data value. In other words, the differential privacy data anonymization device based on temporal random mapping provided in this application considers the time-series characteristics of electricity transaction data and is more suitable for privacy anonymization of time-series data in electricity trading centers.
[0095] Based on the above embodiments, this application also provides an electronic device, which may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, it can implement the steps provided in the above embodiments. Of course, the electronic device may also include various necessary network interfaces, a power supply, and other components.
[0096] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by an execution terminal or processor, can perform the steps provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0097] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0098] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0099] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. For those skilled in the art, various improvements and modifications can be made to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0100] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.
Claims
1. A differential privacy data anonymization method based on temporal random mapping, characterized in that, include: Obtain various target time series data that need to be de-identified by the power trading center, and use the k-medoids clustering algorithm based on dynamic time curvature distance metric to cluster the target time series data to obtain time series datasets of different clusters; Based on the maximum compression ratio of the target time series data, the window size corresponding to different target time series data in the same cluster is calculated, and the average compression ratio of all the target time series data is determined. The average compression ratio is then used as the target window size set during desensitization. Based on the predetermined security protection requirement level of the target time series data, the privacy budget required for differential privacy is determined, and the target time series data is processed according to the target window size and the differential privacy desensitization method of time series random mapping to obtain the electricity market entity data after time series desensitization. The step of calculating the window size corresponding to different target time series data in the same cluster based on the maximum compression ratio of the target time series data includes: Calculate using the following formula: , ; in, F For the frequency of change in time series data, n The number of data points contained in a sequence. Represents the first in the sequence Data points, At the maximum compression ratio, For standard constants, To allow for error, the default is... , ; The step of processing the target time-series data using a differential privacy desensitization method based on the target window size and temporal random mapping includes: Initialize the perturbation sequence to empty, and select a time series as the original sequence, with a size of [missing information]. k The window is placed at the beginning of the original sequence and the beginning of the scrambled sequence; Calculate the number of bits in the mapped sequence window that are empty at position 1, and denote them as . ; Compare values With threshold The size, if The position at the beginning of the original sequence window will with The probability is mapped to any point in the perturbation sequence window where the value is empty; like If this condition is not met and the beginning of the window of the scrambling sequence is empty, then mapping is prioritized; if... If this is not true and the beginning of the window of the scrambling sequence is not empty, then... The probability is mapped to any site where the perturbation sequence window is empty.
2. The method according to claim 1, characterized in that, The clustering of the target time-series data using the k-medoids clustering algorithm based on dynamic time curvature distance metric includes: Randomly select k objects as the initial cluster centers; Non-clustering centroids are repeatedly used to replace the initial clustering centroids for clustering the target time series data; Calculate the cost function for each cluster center, and determine the final cluster center based on the dissimilarity determined according to each cost function.
3. A differential privacy data desensitization device based on temporal random mapping, characterized in that, include: The clustering processing unit is used to acquire various target time series data that need to be de-identified by the power trading center, and to cluster the target time series data using the k-medoids clustering algorithm based on dynamic time curvature distance metric to obtain time series datasets of different clusters. The target window size determination unit is used to calculate the window size corresponding to different target time series data in the same cluster based on the maximum compression ratio of the target time series data, and to determine the average compression ratio of all the target time series data, and to use the average compression ratio as the target window size set during desensitization. The desensitization processing unit is used to determine the privacy budget required for differential privacy based on the predetermined security protection requirement level of the target time series data, and to process the target time series data according to the target window size and the differential privacy desensitization method of time series random mapping to obtain the electricity market entity data after time series desensitization. The target window size determination unit is further configured to: Calculate using the following formula: , ; in, F For the frequency of change in time series data, n The number of data points contained in a sequence. Represents the first in the sequence Data points, At the maximum compression ratio, It is a standard constant. To allow for error, the default is... , ; The desensitization processing unit is further configured to: Initialize the perturbation sequence to empty, and select a time series as the original sequence, with a size of [missing information]. k The window is placed at the beginning of the original sequence and the beginning of the scrambled sequence; Calculate the number of bits in the mapped sequence window that are empty at position 1, and denote them as . ; Compare values With threshold The size, if The position at the beginning of the original sequence window will with The probability is mapped to any point in the perturbation sequence window where the value is empty; like If this condition is not met and the beginning of the window of the scrambling sequence is empty, then mapping is prioritized; if... If this is not true and the beginning of the window of the scrambling sequence is not empty, then... The probability is mapped to any site where the perturbation sequence window is empty.
4. The apparatus according to claim 3, characterized in that, The clustering processing unit is further configured to: Randomly select k objects as the initial cluster centers; Non-clustering centroids are repeatedly used to replace the initial clustering centroids for clustering the target time series data; Calculate the cost function for each cluster center, and determine the final cluster center based on the dissimilarity determined according to each cost function.
5. An electronic device, characterized in that, include: Memory, used for computer programs; A processor configured to implement, when executing a computer program stored in the memory, the steps of the differential privacy data desensitization method based on temporal random mapping as described in any one of claims 1 to 2.
6. A readable storage medium, characterized in that, The readable storage medium stores a computer program, which, when executed by a processor, can implement the steps of the differential privacy data desensitization method based on temporal random mapping as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Efficient intelligent sensing and collecting method for micro-seismic signal detection
CN111046737A
Differential privacy dynamic data publishing method based on mutual information correlation technology
CN112131605A