Security protection strategy optimization method and system for energy big data
By performing timing analysis and clustering of power data of energy data entities, combined with RSA encryption algorithm, adaptive encryption processing of energy big data is realized, solving the shortcomings of traditional encryption algorithms in key distribution management and waste of computing resources, and improving the security protection and sharing efficiency of energy big data.
Patent Information
- Application Number
- CN202410151014.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-02-02
AI Technical Summary
Energy big data faces the security risks of frequent occurrence of multiple points in the process of cross-domain interaction, and traditional encryption algorithms have shortcomings in key distribution management and waste of computing resources.
By conducting timing analysis of the power data of energy data entities, a peak extreme sequence and significant power oscillation factor are constructed, combined with DPC clustering algorithm and RSA encryption algorithm, adaptive encryption processing is realized and security protection strategies are optimized.
It effectively improves the security protection sharing efficiency of energy big data, avoids the use of high-complexity encryption for data with small differences but high protection value, and saves computing resources.
Smart Images

Figure CN117857019B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data encryption, and in particular to a security protection strategy optimization method and system for energy big data. Background Art
[0002] In the external interaction model of energy big data, there are many energy data entities across management domains. With the development of the times, cross-domain interactive business has risen sharply, the number of energy data entities has increased, and energy data security risks have shown a multi-point frequent trend. Among them, the operation planning and resource structure optimization of energy data entities will lead to changes in energy data. The change characteristics of this energy data belong to the high protection value data of the energy data entity, and a more secure protection strategy should be adopted to protect the change characteristics of the energy data entity.
[0003] The security protection strategy of energy big data usually adopts encryption algorithm to protect energy data. Traditional algorithms such as AES (Advanced Encryption Standard) algorithm are symmetric encryption algorithms, which are simple and easy to implement. However, energy big data requires frequent key exchange, and AES symmetric keys are difficult to distribute and manage. RSA (Rivest-Shamir-Adleman) is an asymmetric encryption algorithm, which has significant advantages in key distribution and management in energy big data. However, a large amount of data in energy big data is of low protection value. If all energy data are encrypted with the same complexity, a lot of computing resources will be wasted, which is not conducive to the security protection and sharing of energy data. Summary of the invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a security protection strategy optimization method and system for energy big data. The technical solutions adopted are as follows:
[0005] In a first aspect, an embodiment of the present invention provides a security protection strategy optimization method for energy big data, the method comprising the following steps:
[0006] For energy data, each power energy consumption demander is regarded as a power data entity, and the power consumption of the power data entity at each collection time is collected to form an entity power data sequence. The entity power data sequence is evenly divided into each time series interval, and the power consumption of each time series interval constitutes each entity interval power sequence;
[0007] For each entity interval power sequence, the peak extreme value sequence of each power period is constructed according to each power data in the entity interval power sequence and the corresponding time. The secondary power sequence of each moment is obtained according to each power data and neighboring power data in the entity interval power sequence, and the power oscillation significance factor of each power period is constructed according to the power data of each power period in the entity interval power sequence and the time proportion of the extreme point in the peak extreme value sequence. The local power peak intensity of each moment is obtained according to the power data at each moment and each power data in the secondary power sequence. The comprehensive power peak probability of each moment is obtained according to the local power peak intensity at each moment and the power oscillation significance factor of the power period at each moment. The normalized time proportion value, local power peak intensity and comprehensive power peak probability of each moment are used to form the moment peak vector of each moment, and the DPC algorithm is used to cluster the moment peak vector, and the cluster with the largest mean value of the comprehensive power peak probability is taken as the representative peak cluster of the entity interval.
[0008] According to the local density and time proportion of each data point in the entity interval representative peak cluster, the true coefficient of the power peak interval of the entity interval representative peak cluster is obtained; according to the time proportion of each data in the entity interval power sequence and the true coefficient of the power peak interval, the peak moment similarity extension range of each time series interval is constructed, and the power data is adaptively encrypted according to the peak moment similarity extension range combined with the RSA encryption algorithm.
[0009] Preferably, the step of calculating the time proportion value at the corresponding time according to each power data in the physical interval power sequence to construct the peak extreme value sequence of each power period includes:
[0010] The ratio of each moment in the physical interval power sequence to the maximum moment in the physical interval power sequence is taken as the moment proportion value of each moment;
[0011] The entity interval power sequence is used as the input of the Otsu method, and the power segmentation threshold is output. The moment corresponding to the power data in the entity interval power sequence that is greater than or equal to the power segmentation threshold is recorded as the high power moment, and the high power moment is used as the initial growth point of the regional growing algorithm. The growth criterion is that the adjacent moment of growth must be a high power moment, and each growth area is obtained. The growth area with a high power moment greater than 4 is recorded as the power period;
[0012] The least square method is used to perform nonlinear curve fitting on the data of the power period to obtain the extreme points of the fitting curve, and the extreme points are sorted in order from small to large according to the time proportion value at the corresponding time to form a peak extreme value sequence.
[0013] Preferably, the step of obtaining the secondary power sequence at each moment according to each power data and neighboring power data in the physical interval power sequence includes:
[0014] Taking each moment as the center, w adjacent moments are selected from the left and right sides respectively, and the power data of each moment and the adjacent moments of each moment are sorted in chronological order to form a secondary power sequence of each moment, wherein the number of data in the secondary power sequence is 2w+1. When the number of data in the secondary power sequence is less than 2w+1, the mean of all data in the secondary power sequence is used for filling.
[0015] Preferably, the construction of the power oscillation significant factors of each power period includes:
[0016] For each power period, the mean of the power data in the power period is calculated, and the absolute value of the difference between the mean and the power segmentation threshold of the entity interval power sequence is obtained; the absolute value of the difference between the time proportions of the latter extreme point and the previous extreme point in the adjacent extreme point sequence of the peak extreme value is calculated, recorded as the first absolute value of the difference, and the first absolute value of the difference is used as the exponent of an exponential function with a natural constant as the base, and the sum of the calculation results of the exponential function of all the adjacent extreme points is obtained;
[0017] The ratio of the absolute value of the difference to the sum is calculated, and the product of the number of elements in the peak extreme value sequence and the ratio is used as the power oscillation significance factor of the power period.
[0018] Preferably, the local power peak intensity at each moment includes:
[0019] The difference between the power data at each moment and the power data in the secondary power sequence is calculated, and the sum of the differences between the power data at each moment and all the power data in the secondary power sequence is normalized as the local power peak intensity at each moment.
[0020] Preferably, the comprehensive peak probability of power at each moment includes:
[0021] For each moment, if the ath moment belongs to any power period in the physical interval power sequence, the sum of the power oscillation significance factor of the power period where the ath moment is located and the number 1 is calculated, and the product of the local power peak intensity at the ath moment and the sum is taken as the power comprehensive peak probability at the ath moment;
[0022] If the ath moment does not belong to any power period in the physical interval power sequence, the local power peak intensity at the ath moment is taken as the power comprehensive peak probability at the ath moment.
[0023] Preferably, the method of obtaining the real coefficient of the power peak interval of the entity interval representative peak cluster according to the local density and time proportion of each data point in the entity interval representative peak cluster includes:
[0024] Calculate the absolute value of the difference between the time proportions of any two data points in the entity interval representative peak cluster, obtain the calculation result of an exponential function with the natural constant as the base and the opposite number of the absolute value of the difference as the exponent, and use the sum of the calculation results obtained from all any two data points in the entity interval representative peak cluster as the power peak timing clustering coefficient of the entity interval representative peak cluster;
[0025] The product of the local density mean of all data points in the entity interval representative peak cluster and the power peak time series clustering coefficient is used as the power peak interval true coefficient of the entity interval representative peak cluster.
[0026] Preferably, the peak time similarity extension range of each time series interval is constructed according to the time proportion value of each data in the physical interval power sequence and the real coefficient of the power peak interval, including:
[0027] The mean of the time proportion values of all data in the peak cluster represented by the entity interval is taken as the peak time value of the entity interval power sequence, the peak time values of the power sequences of all entity intervals in the nth power data entity are combined into the peak time sequence of the nth power data entity, the peak time sequence is taken as the input of the Bayesian online change point detection algorithm, and the time series interval corresponding to the output Bayesian mutation point is taken as the peak time mutation interval;
[0028] For the mth time series interval of the nth power data entity, the peak moment mutation interval on its left side and with the closest time interval is used as the left mutation interval of the mth time series interval, and the peak moment mutation interval on its right side and with the closest time interval is used as the right mutation interval of the mth time series interval;
[0029] The number of adjacent similar intervals taken to the left and right of the mth time series interval in the nth power data entity is recorded as The expressions are:
[0030]
[0031]
[0032] In the formula, is the power feature vector of the mth time series interval in the nth power data entity, are the power feature vectors corresponding to the left and right mutation intervals of the mth time series interval in the nth power data entity, Cos() is the cosine similarity, is the peak mutation probability corresponding to the left and right mutation intervals of the mth time series interval in the nth power data entity, α is a parameter adjustment factor preset to be greater than zero, wherein the power characteristic vector is composed of the peak time value and the real coefficient of the power peak interval;
[0033] The similar extension range of the peak time of the mth time series interval in the nth power data entity is
[0034] Preferably, the adaptive encryption processing of the power data based on the similar expansion range at the peak time in combination with the RSA encryption algorithm includes:
[0035] For the mth time interval of the nth power data entity, when the peak moment similar extension range of any other time interval includes the mth time interval, the accumulator of the mth time interval is increased by 1, and the accumulator value of the mth time interval is counted;
[0036] The accumulator values of each time series interval are counted using the method for obtaining the accumulator value of the mth time series interval, and the power feature vector corresponding to the time series interval with the largest accumulator value is recorded as the power entity feature vector of the nth power data entity;
[0037] The DBSCAN algorithm is used to cluster the power entity feature vectors of all power data entities, and the entity power data sequence corresponding to the entity cluster with the least number of isolated points and data points is used as the high protection value data of this round, and combined with the ciphertext data obtained in the previous round as the data to be encrypted in this round of encryption. If this round of encryption belongs to the first round of encryption, the high protection value data of this round is used as the original plaintext data of this round of encryption;
[0038] Among them, the encryption algorithm adopts the RSA encryption algorithm.
[0039] In a second aspect, an embodiment of the present invention further provides a security protection strategy optimization system for energy big data, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of any one of the above-mentioned methods when executing the computer program.
[0040] The present invention has at least the following beneficial effects:
[0041] The present invention constructs the real coefficient of the power peak interval of the power data entity time series interval according to the peak characteristics of high power moments between power data in the time series interval, combines the Otsu method and the DPC clustering algorithm, and reflects the confidence of the peak moment of the power data calculated. The beneficial effect is that it comprehensively considers the power data peak time series aggregation characteristics of the power data entity, and avoids the problem that the obtained power data peak cannot represent the entire power data entity time series interval; based on the similarity of the peak moments in the power data entity time series interval, combined with the Bayesian online change point detection algorithm and the DBSCAN clustering algorithm, the data to be encrypted in the multi-round RSA encryption process is obtained. The beneficial effect is that it avoids the problem of high encryption complexity for power data entities with small differences and high protection value, optimizes the security protection strategy of energy big data, and improves the efficiency of energy big data security protection sharing. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0043] Figure 1 A flowchart of the steps of a security protection strategy optimization method for energy big data provided by one embodiment of the present invention;
[0044] Figure 2 Schematic diagram of the clustering effect of power entity feature vectors of power data entities. DETAILED DESCRIPTION
[0045] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the security protection strategy optimization method and system for energy big data proposed by the present invention, its specific implementation method, structure, features and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0046] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0047] The specific scheme of the security protection strategy optimization method and system for energy big data provided by the present invention is described in detail below with reference to the accompanying drawings.
[0048] See also Figure 1 , which shows a flowchart of a method for optimizing a security protection strategy for energy big data provided by an embodiment of the present invention, the method comprising the following steps:
[0049] It should be noted that the present invention aims to optimize the security protection strategy for energy data. Energy data includes various types, including electricity data, coal data (coal consumption), natural gas data (natural gas consumption), oil data (oil consumption), etc. This embodiment takes electricity data as an example to explain in detail the data security protection strategy optimization process. For other energy data, implementers can use the same processing and protection strategy optimization methods as electricity data to optimize the security protection strategies for other energy data.
[0050] Step S001, collecting power data of each power data entity through the database of the energy big data center and performing pre-processing.
[0051] In the energy big data center, each power energy consumption demander is regarded as a power data entity. Through the database of the energy big data center, the power consumption of the power data entity is taken as power data every t sampling time intervals. For ease of understanding, the power data described later in this embodiment is also the power consumption data, and will not be explained one by one later. The power data of each power data entity is collected, and the total collection time is T. The collection time interval is set to 30 minutes in this embodiment, and the total collection time is 30 days. Since the power data of the power data entity may be missing during the transmission process, this embodiment first uses the Lagrangian filling method to fill in the data of all the acquired power data, wherein the Lagrangian filling method is a well-known technology and will not be repeated here.
[0052] The filled power data of each power data entity is sorted in chronological order to construct the entity power data sequence A. The total number of power data entities is N. Secondly, in this embodiment, 1 day is taken as a time series interval, and the above entity power data sequence is divided into 30 time series intervals. The data sequence corresponding to each time series interval is recorded as the entity interval power sequence B, and the total number of time series intervals is recorded as M. Among them, A n Represents the entity power data sequence corresponding to the nth electronic data entity in the energy big data center, It represents the entity interval power sequence of the mth time interval in the nth power data entity in the energy big data center.
[0053] For other energy data, the implementer may obtain the energy data entity, time series interval and corresponding entity interval energy sequence of other energy data according to the above method of this embodiment.
[0054] Step S002, based on the peak characteristics of high power moments between power data in the time series interval, construct the real coefficient of the power peak interval of the power data entity time series interval, and combine the Bayesian online change point detection algorithm and the DBSCAN clustering algorithm to obtain the data to be encrypted in the multi-round RSA encryption process.
[0055] In the energy big data center, different power data entities may have different power data, and the changes in these power data will reflect the operation and development direction of the electronic data entity. For example, a factory belongs to a power data entity. When the factory increases the production of its products, the power data of the factory, that is, the power consumption, will increase accordingly, and continue thereafter. After the attacker obtains these power data, he may determine when the factory increases production, causing the factory to suffer significant losses. Therefore, it is usually necessary to encrypt the power data of these power data entities. At the same time, not all power data entities have more important power data information. If the same level of encryption is used for all power data entity data, it will lead to long encryption and decryption time and waste of computing resources. Therefore, the purpose of this embodiment is to use a more complex encryption method for power data entities with important power data information and a high security protection level, and to use a less complex encryption method for power data entities with relatively unimportant power data information and a low security protection level.
[0056] In the energy big data center, the power data corresponding to the power data entity may have certain peak characteristics. For example, the production activities of a factory usually have peak periods within a day. Most factories work during the day. When the factory equipment and machines are running, the power demand will increase, forming a peak.
[0057] The entity interval power sequence corresponding to the mth time series interval of the nth power data entity in the energy big data center For example, the entity interval power sequence As the input of Otsu's method, the power split threshold is output The physical interval power sequence The moment when the medium power data is greater than or equal to the power segmentation threshold is recorded as the high power moment. The Otsu method is a well-known technology and the specific process will not be repeated. Secondly, the regional growing algorithm is used to obtain the power time period. Specifically, the high power moment is used as the initial growth point of the regional growing algorithm, and the growth criterion is set that the adjacent moments of growth must also be high power moments. The growth is carried out in sequence to obtain each growth area; further, the growth area with a high power moment greater than 4 is recorded as the power time period. Calculate the power sequence of the entity interval The moment proportion value of each moment in the sequence is calculated as follows: the ratio of each moment in the sequence to the maximum moment in the physical interval power sequence is taken as the moment proportion value of each moment. For example, if the maximum moment in the physical interval sequence is 48, then the moment proportion value of the 5th moment in the physical interval sequence is Note it as μ5. Further, based on the entity interval power sequence Taking the rth power period as an example, the moment proportion value corresponding to each high power moment is taken as the horizontal coordinate, and the power data value is taken as the vertical coordinate. The least squares method is used to perform nonlinear curve fitting on the power period data to obtain the fitting curve. The extreme value points of the fitting curve are counted, and the extreme value points are sorted in order from the smallest to the largest according to the moment proportion value of the corresponding moment. The composed sequence is recorded as the peak extreme value sequence.
[0058] Furthermore, based on the physical interval power sequence Take the ath moment as the center, take 4 adjacent moments to the left, and take 4 adjacent moments to the right, and construct a secondary power sequence with a length of 9 with the corresponding power data values in order of the above adjacent moments and the ath moment. If the number of data in the secondary power sequence is less than 9, the mean of the existing power data values in the secondary power sequence will be filled.
[0059] Based on the above analysis, the physical interval power sequence can be obtained: The probability of the comprehensive peak power corresponding to the ath moment ρ a :
[0060]
[0061]
[0062]
[0063] Among them, s r It is the entity interval power sequence The significant factor of power oscillation in the rth power period, It is the entity interval power sequence The mean value of the power data in the rth power period, It is the entity interval power sequence The power split threshold, C q+1 , C q is the time proportion corresponding to the q+1th and qth extreme value points in the peak extreme value sequence, Q r is the total number of data elements in the peak extreme value sequence.
[0064] When the difference between the time proportions corresponding to two adjacent extreme points is smaller, that is, The smaller it is, the denser the extreme points appear during the power period. At the same time, when the number of extreme points is greater, that is, Q r The larger the value is, the more frequent the fluctuation of power data is. The larger the value is, the greater the overall power consumption during the power period is. The more the power period fluctuates within the high power data range, the more significant and frequent the fluctuation characteristics are, and the more likely it is the peak period of the power data entity in a day. The peak period fluctuation significance factor S r The bigger.
[0065] U a It is the entity interval power sequence The local peak power intensity at the ath moment, H a is the power data value corresponding to the ath moment, H a,g is the power data value of the gth data element in the secondary power sequence at the ath moment, G is the sequence length of the secondary power sequence, and the empirical value is 9. Norm[] is the normalization function, which normalizes the local power peak intensity to (0,1) to ensure that the local power peak intensity at each moment is greater than zero.
[0066] The larger the power data value corresponding to the ath moment is compared with the power data value in the secondary power sequence, that is, (H a -H a,g ) is larger, it means that the power sequence in the physical interval The higher the local power intensity of the power data value at the ath moment, the higher the local power peak intensity U a The bigger.
[0067] ρ a It is the entity interval power sequence The probability of the comprehensive peak power corresponding to the ath moment in time is: It is the entity interval power sequence The set of all moments in each power period, Indicates that the ath moment belongs to the entity interval power sequence During any power period, Indicates that the ath moment does not belong to the physical interval power sequence Any power period in the
[0068] When the local power intensity of the power data value at the ath moment is greater, that is, U a The larger the value is, the more likely the power data corresponding to the ath moment is to be a physical interval power sequence. At the same time, the more significant and frequent the oscillation characteristics of the power band corresponding to the a-th moment are, the more likely it is the peak period of the power data entity in a day, that is, Sr The larger the value is, the less likely it is that the power data value at the ath moment is an independent noise point, and the more likely it is that the power sequence of the entire entity interval The comprehensive power peak value, then the power comprehensive peak probability ρ a The bigger.
[0069] So far, the physical interval power sequence can be obtained The moment proportion μ, local power peak intensity U, and power comprehensive peak probability ρ corresponding to each moment in the equation are standardized using the Z-score method due to the different dimensions of the above indicators. Taking the ath moment in the example, the standardized moment proportion value, local power peak intensity, and power comprehensive peak probability are combined to form the moment peak vector V a =[v a,1 ,v a,2 ,v a,3 ], where v a,1 ,v a,2 ,v a,3 Represents the entity interval power sequence The standardized moment proportion value, local power peak intensity, and power comprehensive peak probability of the ath moment in the equation are used as dimensions to construct a three-dimensional feature space. The moment peak vector corresponding to each moment in is projected into the feature space, and the obtained feature space data points are recorded as the moment peak data set. The moment peak data set is used as the input of the density peak clustering DPC (Density peaks Clustering) algorithm. The cross-validation method is used to obtain the truncation distance in the DPC algorithm. The output of the DPC algorithm is the entity interval power sequence The local density and entity interval power series corresponding to each moment in There are K clusters.
[0070] Calculate the mean of all data points in all DPC clusters in the power comprehensive peak probability dimension, and take the cluster with the largest power comprehensive peak probability mean as the entity interval representative peak cluster. Based on the above analysis, the entity interval power sequence can be obtained. The entity interval represents the real coefficient of the peak power interval of the peak cluster
[0071]
[0072]
[0073] in, It is the entity interval power sequence The entity interval represents the power peak timing clustering coefficient of the peak cluster, C f , C f′ The entity interval power series The entity interval represents the time proportion corresponding to the fth and f′th data points in the peak cluster. It is the entity interval power sequence The entity interval represents the number of all data points in the peak cluster. When the entity interval represents the time proportion values corresponding to any two data points in the peak cluster, the closer the value is, The larger the value, the greater the physical interval represents. The time corresponding to the peak cluster is in the physical interval power sequence. The more they appear together, the more the entity interval represents the peak cluster, which conforms to the entity interval power sequence. Time clustering characteristics of the mid-peak interval and the power peak timing clustering coefficient The bigger.
[0074] It is the entity interval power sequence The entity interval represents the real coefficient of the peak power interval of the peak cluster, It is the entity interval power sequence The entity interval represents the local density mean of all data points in the peak cluster. When the entity interval represents the local density mean of the peak cluster, the larger the The larger the value, the better the DPC clustering result is, the greater the credibility of the calculated power peak timing clustering coefficient is, and the more the entity interval represents the peak cluster, the more consistent it is with the entity interval power sequence. The moment aggregation characteristics of the mid-peak interval, that is, The larger the value is, the more the peak cluster of the entity interval can represent the peak interval of the entire entity interval power sequence, and the true coefficient of the power peak interval is The bigger.
[0075] According to the above steps, calculate the physical interval power sequence The entity interval represents the mean of the time proportion of all data in the peak cluster, recorded as the peak time value Secondly, the physical interval power sequence Peak time value and the real coefficient of the peak power interval Composition of physical interval power sequence The power characteristic vector of the time series interval is: Peak time value Characterizing the power sequence of a physical interval The peak time of power data and the real coefficient of power peak interval Used to characterize the confidence level of the peak moments of the obtained power data.
[0076] By repeating the above method of this embodiment, the power feature vector of each time series interval can be obtained, which is used to characterize the peak value of the power data in the time series interval.
[0077] Since the power usage of the power data entity has a certain regularity, for example, the start and stop of the factory equipment usually occurs at a fixed time every day, and the working hours of the factory staff also show regularity, the peak time of the power data of the power data entity is relatively similar every day.
[0078] According to the above steps, the peak time values corresponding to the power sequences of all entity intervals in the nth power data entity in the energy big data center are used to form the peak time sequence of the nth power data entity, and the peak time sequence is used as the input of the Bayesian online change point detection algorithm. The time series interval corresponding to the output Bayesian mutation point is recorded as the peak time mutation interval, and the mutation probability of the output Bayesian mutation point is recorded as the peak mutation probability of the peak time mutation interval. The Bayesian online change point detection algorithm is a well-known technology, and the specific process is not repeated. Taking the mth time series interval as an example, the peak time mutation interval on the left side of the mth time series interval and the closest time interval is used as the left mutation interval of the mth time series interval, and the peak time mutation interval on the right side of the mth time series interval and the closest time interval is used as the right mutation interval of the mth time series interval.
[0079] Based on the above analysis, the similar extension range of the peak moment of the mth time series interval in the nth power data entity can be obtained:
[0080]
[0081]
[0082]
[0083] in, is the number of adjacent similar intervals taken to the left of the mth time interval in the nth power data entity, The power feature vector corresponding to the mth time series interval in the nth power data entity, is the power feature vector corresponding to the left mutation interval of the mth time series interval in the nth power data entity, Cos() is the cosine similarity, It is the peak mutation probability corresponding to the left mutation interval of the mth time series interval in the nth power data entity. α is a parameter adjustment factor preset to be greater than zero. Its function is to prevent the denominator from being 0. In this embodiment, the value of α is 0.01, and round() is a rounding function.
[0084] is the number of adjacent similar intervals taken to the right of the mth time interval in the nth power data entity, is the power feature vector corresponding to the right mutation interval of the mth time series interval in the nth power data entity, It is the peak mutation probability corresponding to the right mutation interval of the mth time series interval in the nth power data entity.
[0085] It is the similar extension range of the peak moment of the mth time series interval in the nth power data entity.
[0086] The greater the cosine similarity of the power feature vector between the mth time interval and the corresponding left mutation interval, that is, The larger it is, the more similar the power data characteristics of the physical interval power sequence of the mth time interval and the left mutation interval are. At the same time, the smaller the mutation probability at the peak moment of the left mutation interval is, that is, The smaller it is, the lower the change degree of the peak time value of the left mutation interval in the peak time sequence, and the more time series intervals on the left side are similar to the mth time series interval. The bigger, the same The larger the range, the higher the similarity of power characteristics of each time series interval in a similar extension range at the same peak moment.
[0087] Based on the above steps, an accumulator with an initial value of 0 is created for each time series interval in the nth power data entity. Taking the mth time series interval as an example, when the peak moment similar extension range corresponding to any other time series interval includes the mth time series interval, the accumulator of the mth time series interval is increased by 1. After traversing the peak moment similar extension range corresponding to all time series intervals of the nth power data entity, the accumulator value of the mth time series interval is counted. Similarly, the accumulator value of each time series interval is obtained. For the time series interval with the largest accumulator value, the time series interval has a higher power characteristic similarity with other time series times, and can better represent the entire power data entity. Therefore, the power characteristic vector Y corresponding to the time series interval with the largest accumulator value is recorded as the power entity characteristic vector of the nth power data entity.
[0088] Since the vector values of each dimension in the power entity feature vector corresponding to each power data entity have different dimensions, Z-score standardization is used to eliminate the dimension of the power entity feature vectors of all power data entities. The standardized power entity feature vectors of all power data entities are used as the input of the DBSCAN clustering algorithm, where the neighborhood radius (ε) is set to an empirical value of 0.1, the minimum number of neighborhood points (MinPts) is set to 5, and the metric distance is Euclidean distance. Z clusters are output, and each of the above DBSCAN clusters is recorded as an entity cluster cluster. Each data point in the entity cluster cluster corresponds to a power data entity and an entity power data sequence. The schematic diagram of the clustering effect of the power entity feature vectors of each power data entity is as follows: Figure 2 As shown, the clustering results are visualized. Figure 2 In , × is an isolated point, and the triangle, square and solid circle represent different clusters. The Z-score standardization and DBSCAN clustering algorithm are both well-known technologies, and the specific process will not be repeated here.
[0089] Count the number of data points in each entity cluster. Since the isolated points in the DBSCAN clustering results are greatly different from the data points of other entity clusters, and the fewer the data points in the entity cluster, the fewer similar entity power data sequences are. These data points are more likely to be power data with greater differences and higher protection value in the energy big data center, and higher security protection strategies should be adopted for these data.
[0090] Therefore, the RSA algorithm is used as the encryption algorithm. In this round of encryption, the entity power data sequence corresponding to the isolated points of the DBSCAN clustering results and the entity clustering cluster with the least number of data points is recorded as the high protection value data of this round, and combined with the ciphertext data obtained in the previous round, it is used as the data to be encrypted in this round of encryption. It should be noted that if this round of encryption belongs to the first round of encryption, the high protection value data of this round is directly used as the original plaintext data of this round of encryption. Each round of encryption algorithm adopts the RSA encryption algorithm to realize the security protection strategy optimization method for energy big data. Among them, the multi-round RSA encryption algorithm is a well-known technology, and the specific process will not be repeated.
[0091] At this point, the security protection strategy for power data has been optimized. The same method as in this embodiment can be used to optimize the security protection strategy for other energy data.
[0092] Based on the same inventive concept as the above method, an embodiment of the present invention also provides a security protection strategy optimization system for energy big data, including a memory, a processor, and a computer program stored in the memory and running on the processor, and when the processor executes the computer program, the steps of any one of the above-mentioned security protection strategy optimization methods for energy big data are implemented.
[0093] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The above is a description of a specific embodiment of this specification. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0094] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A security protection strategy optimization method for energy big data, characterized in that: The method comprises the following steps: For energy data, each power energy consumption demander is regarded as a power data entity, and the power consumption of the power data entity at each collection time is collected to form an entity power data sequence. The entity power data sequence is evenly divided into each time series interval, and the power consumption of each time series interval constitutes each entity interval power sequence; For each entity interval power sequence, the moment proportion is calculated according to each power data in the entity interval power sequence and the corresponding moment, and the peak extreme value sequence of each power time period is constructed; the moment proportion is the ratio of the position of each moment in the entity interval power sequence to the position of the maximum moment, and the moments of the power data that are aggregated in the entity interval power sequence and are greater than or equal to the power segmentation threshold constitute the power time period, and the extreme value points of the power time period are obtained, and the extreme value points are sorted according to the moment proportion to obtain the peak extreme value sequence; the secondary power sequence of each moment is obtained according to each power data and neighboring power data in the entity interval power sequence, and a preset number of power data are selected around each moment and sorted in chronological order as the secondary power sequence of each moment, and the power data of each power time period in the entity interval power sequence and the moment proportion of the extreme value points in the peak extreme value sequence are constructed. Establish the power oscillation significance factor of each power period; obtain the local power peak intensity at each moment according to the power data at each moment and each power data in the secondary power sequence, wherein the local power peak intensity is the normalized value of the sum of the differences between the power data at each moment and each power data in the secondary power sequence; obtain the power comprehensive peak probability at each moment according to the local power peak intensity at each moment and the power oscillation significance factor of the power period at each moment, and obtain the power comprehensive peak probability through the power oscillation significance factor and the local power peak intensity according to the relationship between the moment and the power period; form the moment peak vector of each moment with the normalized moment proportion value, local power peak intensity and power comprehensive peak probability, cluster the moment peak vector using the DPC algorithm, and take the cluster with the largest mean value of the power comprehensive peak probability as the representative peak cluster of the entity interval; The real coefficient of the peak interval of the physical interval representing the peak cluster is obtained according to the local density and time proportion of each data point in the physical interval representing the peak cluster; the peak time similarity extension range of each time series interval is constructed according to the time proportion value of each data in the physical interval power sequence and the real coefficient of the peak interval, and the power data is adaptively encrypted according to the peak time similarity extension range combined with the RSA encryption algorithm; The construction of significant factors of power fluctuations in each power period includes: For each power period, the mean of the power data in the power period is calculated, and the absolute value of the difference between the mean and the power segmentation threshold of the entity interval power sequence is obtained; the absolute value of the difference between the time proportions of the latter extreme point and the previous extreme point in the adjacent extreme point sequence of the peak extreme value is calculated, recorded as the first absolute value of the difference, and the first absolute value of the difference is used as the exponent of an exponential function with a natural constant as the base, and the sum of the calculation results of the exponential function of all the adjacent extreme points is obtained; Calculating the ratio of the absolute value of the difference to the sum, and taking the product of the number of elements in the peak extreme value sequence and the ratio as the power oscillation significance factor of the power period; The method of constructing a peak time similarity extension range of each time series interval according to the time proportion value of each data in the physical interval power sequence and the real coefficient of the power peak interval includes: The mean of the time proportion values of all data in the peak cluster represented by the entity interval is taken as the peak time value of the entity interval power sequence, the peak time values of the power sequences of all entity intervals in the nth power data entity are combined into the peak time sequence of the nth power data entity, the peak time sequence is taken as the input of the Bayesian online change point detection algorithm, and the time series interval corresponding to the output Bayesian mutation point is taken as the peak time mutation interval; For the mth time series interval of the nth power data entity, the peak moment mutation interval on its left side and with the closest time interval is used as the left mutation interval of the mth time series interval, and the peak moment mutation interval on its right side and with the closest time interval is used as the right mutation interval of the mth time series interval, n represents the nth power data entity, and m represents the mth time series interval; The number of adjacent similar intervals taken to the left and right of the mth time series interval in the nth power data entity is recorded as The expressions are: In the formula, is the power feature vector of the mth time series interval in the nth power data entity, are the power feature vectors corresponding to the left and right mutation intervals of the mth time series interval in the nth power data entity, Cos() is the cosine similarity, is the peak mutation probability corresponding to the left and right mutation intervals of the mth time series interval in the nth power data entity, α is a parameter adjustment factor preset to be greater than zero, wherein the power characteristic vector is composed of the peak time value and the real coefficient of the power peak interval; The similar extension range of the peak time of the mth time series interval in the nth power data entity is The adaptive encryption processing of the power data according to the similar expansion range at the peak time combined with the RSA encryption algorithm includes: For the mth time interval of the nth power data entity, when the peak moment similar extension range of any other time interval includes the mth time interval, the accumulator of the mth time interval is increased by 1, and the accumulator value of the mth time interval is counted; The accumulator values of each time series interval are counted using the method for obtaining the accumulator value of the mth time series interval, and the power feature vector corresponding to the time series interval with the largest accumulator value is recorded as the power entity feature vector of the nth power data entity; The DBSCAN algorithm is used to cluster the power entity feature vectors of all power data entities, and the entity power data sequence corresponding to the entity cluster with the least number of isolated points and data points is used as the high protection value data of this round, and combined with the ciphertext data obtained in the previous round as the data to be encrypted in this round of encryption. If this round of encryption belongs to the first round of encryption, the high protection value data of this round is used as the original plaintext data of this round of encryption; Among them, the encryption algorithm adopts the RSA encryption algorithm.
2. The security protection strategy optimization method for energy big data according to claim 1 is characterized in that: The method of calculating the time proportion value according to each power data in the physical interval power sequence and the corresponding time to construct the peak extreme value sequence of each power period includes: The ratio of each moment in the physical interval power sequence to the maximum moment in the physical interval power sequence is taken as the moment proportion value of each moment; The entity interval power sequence is used as the input of the Otsu method, and the power segmentation threshold is output. The moment corresponding to the power data in the entity interval power sequence that is greater than or equal to the power segmentation threshold is recorded as the high power moment, and the high power moment is used as the initial growth point of the regional growing algorithm. The growth criterion is that the adjacent moment of growth is the high power moment, and each growth area is obtained. The growth area with a high power moment greater than 4 is recorded as the power period; The least square method is used to perform nonlinear curve fitting on the data of the power period to obtain the extreme points of the fitting curve, and the extreme points are sorted in order from small to large according to the time proportion value at the corresponding time to form a peak extreme value sequence.
3. The security protection strategy optimization method for energy big data according to claim 2 is characterized in that: The method of obtaining the secondary power sequence at each moment according to each power data and neighboring power data in the physical interval power sequence includes: Taking each moment as the center, w adjacent moments are selected from the left and right sides respectively, and the power data of each moment and the adjacent moments of each moment are sorted in chronological order to form a secondary power sequence of each moment, wherein the number of data in the secondary power sequence is 2w+1. When the number of data in the secondary power sequence is less than 2w+1, the mean value of all data in the secondary power sequence is used for filling, and w is a preset number.
4. The security protection strategy optimization method for energy big data according to claim 1 is characterized in that: The local power peak intensity at each moment includes: The difference between the power data at each moment and the power data in the secondary power sequence is calculated, and the sum of the differences between the power data at each moment and all the power data in the secondary power sequence is normalized as the local power peak intensity at each moment.
5. The security protection strategy optimization method for energy big data according to claim 4 is characterized in that: The comprehensive peak probability of power at each moment includes: For each moment, if the ath moment belongs to any power period in the physical interval power sequence, the sum of the power oscillation significance factor of the power period where the ath moment is located and the number 1 is calculated, and the product of the local power peak intensity at the ath moment and the sum is taken as the power comprehensive peak probability at the ath moment; If the ath moment does not belong to any power period in the physical interval power sequence, the local power peak intensity at the ath moment is taken as the power comprehensive peak probability at the ath moment, and a is each moment.
6. The security protection strategy optimization method for energy big data according to claim 1 is characterized in that: The method of obtaining the real coefficient of the power peak interval of the entity interval representative peak cluster according to the local density and time proportion of each data point in the entity interval representative peak cluster includes: Calculate the absolute value of the difference between the time proportions of any two data points in the entity interval representative peak cluster, obtain the calculation result of an exponential function with a natural constant as the base and the opposite number of the absolute value of the difference as the exponent, and use the sum of the calculation results obtained from all any two data points in the entity interval representative peak cluster as the power peak timing clustering coefficient of the entity interval representative peak cluster; The product of the local density mean of all data points in the entity interval representative peak cluster and the power peak time series clustering coefficient is used as the power peak interval true coefficient of the entity interval representative peak cluster.
7. A security protection strategy optimization system for energy big data, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Power monitoring network security situation awareness method based on improved AHP method
CN110443037A
KR1018666930000B1