Energy big data secure storage method, storage medium and system

By analyzing the mutual influence and regularity of energy data, calculating the encryption strength and compensation coefficient, and adaptively adjusting the key length, solving the problems of low encryption efficiency and insufficient encryption strength in the existing technology, and achieving efficient and secure encrypted storage of energy big data.

CN120046169AActive Publication Date: 2025-05-27HULUDAO POWER SUPPLY COMPANY OF STATE GRID LIAONING ELECTRIC POWER

Patent Information

Application Number
CN202510141269.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-27
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

In the process of energy big data storage, it is difficult for the prior art to improve the encryption strength while ensuring encryption efficiency, resulting in the possible leakage of important data.

Method used

By analyzing the mutual influence between energy data and the regularity of the data itself, calculating the encryption strength coefficient and encryption compensation coefficient, adaptively adjusting the key length, and encrypting the energy data using the AES encryption algorithm.

Benefits of technology

While ensuring encryption efficiency, it improves the encryption strength of important energy data, avoids the problems of low encryption efficiency or low encryption strength caused by fixed key length, and ensures data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046169A_ABST
    Figure CN120046169A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of big data processing, in particular to an energy big data secure storage method, a storage medium and a system.The method comprises the steps that energy data of any type of energy is obtained, the energy data comprises multiple data types, and any group of data in any data type is arranged according to a time sequence to form a data sequence; determining an association value and a data regularity of each type of data; the encryption strength coefficient of each type of data is obtained; determining the local feature approximation degree of each type of data based on the mutation moment difference and the distribution range difference corresponding to the subsequences between each type of data and all other types of data; determining the period approximation degree of each type of data; further obtaining a key length coefficient of each type of data; and when the energy data is stored, encrypting various types of data by adopting an AES encryption algorithm based on the key length coefficient. The invention aims to improve the encryption strength while ensuring the encryption efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of big data processing, and particularly relates to a method, storage medium and system for secure storage of energy big data. Background Art

[0002] The secure storage of energy big data is one of the key factors to ensure the success of the digital transformation of the energy industry. With the advancement of the digital transformation of the energy industry, data security issues have become increasingly prominent. Especially in a new power system dominated by new energy, there are security risks in all aspects of data collection, storage, use, processing, transmission, provision and disclosure.

[0003] Among them, during the process of storing energy big data, in order to avoid data leakage or theft, the stored data is usually encrypted to ensure its security in the static state. Among existing encryption algorithms, the AES encryption algorithm is widely used in big data encryption due to its advantages such as high efficiency, high security and easy implementation.

[0004] The AES encryption algorithm usually encrypts with a fixed-size key. However, the amount of data contained in energy big data is large and the content involved is relatively extensive. Not all types of data need to be encrypted with high intensity. For example, the charging pile location data in electric vehicle charging pile data usually needs to be made public to the public, while the charging user data in electric vehicle charging pile data needs to be highly confidential. If a large key is used to encrypt all types of data, a large amount of encryption time and resources will be consumed, resulting in a significant reduction in encryption efficiency. However, if a short key is used for encryption, the data with a higher degree of importance in energy big data may be easily stolen due to the low encryption intensity, causing losses. Summary of the Invention

[0005] In view of the above, it is necessary to provide a method, storage medium and system for secure storage of energy big data, which can improve the encryption intensity while ensuring the encryption efficiency compared with the traditional method for secure storage of energy big data:

[0006] In a first aspect, an embodiment of the present application provides a method for secure storage of energy big data, and the method includes the following steps:

[0007] Obtain energy data of any type of energy, which contains multiple data categories, and each data category contains multiple groups of data. Arrange any group of data in any data category in time sequence to form a data sequence;

[0008] Based on the difference in the change trend between all corresponding data sequences of each type of data and the remaining all types of data, and the difference in the time at which the extreme values are located within all corresponding data sequences, determine the correlation value of each type of data;

[0009] Determine the data regularity degree of each type of data based on the similarity between any two data sequences in all kinds of data and the chaos degree of all data sequences;

[0010] Based on the correlation value and the data regularity degree, obtain the encryption strength coefficient of each type of data;

[0011] Divide each data sequence into subsequences of each preset length, and determine the local feature approximation degree of each type of data based on the mutation time difference and distribution range difference of the corresponding subsequences between each type of data and all other types of data;

[0012] Determine the period approximation degree of each type of data based on the similarity degree of periodic characteristics between any two subsequences of each data sequence in each type of data and the similarity degree of data fluctuation between them;

[0013] Combine the local feature approximation degree, the period approximation degree and the encryption strength coefficient to obtain the key length coefficient of each type of data;

[0014] When storing energy data, encrypt each type of data using the AES encryption algorithm based on the key length coefficient.

[0015] In one embodiment, the expression of the correlation value is:

[0016] In the formula, A i is the correlation value of the i-th type of data; B i,j is the mean value of the trend strength difference of the j-th data sequence between the i-th type of data and all other types of data; C i,j is the mean value of the differences of the moments where all corresponding extreme values are located in the j-th data sequence between the i-th type of data and all other types of data; J i is the minimum value of the total number of data sequences in all types of data; α is a first preset value greater than 0.

[0017] In one embodiment, the determination process of the data regularity degree is as follows:

[0018] Calculate the approximate entropy mean value of all data sequences in each type of data;

[0019] Denote the mean value of the distances between any two data sequences in each type of data as the distance mean value;

[0020] The data regularity degree is inversely proportional to the approximate entropy mean value and the distance mean value respectively.

[0021] In one embodiment, the determination process of the local feature approximation degree is as follows:

[0022] Denote any data category as i, calculate the difference in the moments where the corresponding mutation points of the x-th subsequence of the j-th data sequence between the i-th category of data and the data of the remaining categories, and denote it as the mutation moment difference;

[0023] Denote the ratio of the range of each subsequence to the sequence length as the first ratio, calculate the difference in the first ratio of the x-th subsequence of the j-th data sequence between the i-th category of data and the data of the remaining categories, and denote it as the range difference;

[0024] The expression for the local feature approximation degree of the i-th category of data is:

[0025] In the formula, G i is the local feature approximation degree of the i-th category of data; H i,j is the mean value of all the mutation moment differences corresponding to the j-th data sequence in the i-th category of data; K i,j is the mean value of all the range differences corresponding to the j-th data sequence in the i-th category of data; J i is the minimum value of the total number of data sequences in all categories of data; γ is a preset third value greater than 0.

[0026] In one of the embodiments, the expression for the period approximation degree is:

[0027] In the formula, M i is the period approximation degree of the i-th category of data; N i,j and P i,j are respectively the mean value of the phase-locking values and the sum value of the differences in the degree of dispersion between all any two subsequences of the j-th data sequence in the i-th category of data; I is the total number of data sequences in the i-th category of data; ε is a preset fourth value greater than 0.

[0028] In one of the embodiments, the key length coefficient is the fusion result of the local feature approximation degree, the period approximation degree, and the encryption strength coefficient.

[0029] In one of the embodiments, the process of encrypting various types of data is as follows:

[0030] Use the threshold segmentation algorithm to obtain the segmentation threshold of all key length coefficients;

[0031] Mark the category data with a key length coefficient greater than or equal to the segmentation threshold as high-security data, otherwise, mark it as ordinary-security data;

[0032] Use the AES encryption algorithm to encrypt the high-security data and the ordinary-security data with different lengths of keys.

[0033] In one of the embodiments, the key length during high-security data encryption is greater than the key length during normal-security data encryption.

[0034] In a second aspect, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which when executed by a processor implements an energy big data secure storage method as described in the first aspect.

[0035] In a third aspect, an embodiment of the present application further provides an energy big data secure storage system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of an energy big data secure storage method as described in any one of the above.

[0036] The present application has at least the following beneficial effects:

[0037] When storing energy data, the present application analyzes the mutual influence situation among energy data and the law situation of the energy data itself to obtain an encryption intensity coefficient, which is used to characterize the importance degree of the stored energy data and the intensity of encryption required. The higher the importance degree of the energy data, the greater the encryption intensity;

[0038] Furthermore, when there is a missing situation in energy data that affects the data characteristics, the present application analyzes the approximation situation of local characteristics among energy data and the periodic characteristics of the energy data itself to obtain an encryption compensation coefficient, which can characterize the correlation degree among energy data and the regularity of the energy data itself in the case of data missing, and avoid misjudgment of the encryption intensity required for data due to data missing; combining the encryption intensity coefficient and the encryption compensation coefficient to obtain a key length coefficient to characterize the encryption intensity required for energy data. Then, when using the AES encryption algorithm to encrypt and store energy data, the key length can be adaptively adjusted, ensuring the encryption efficiency while improving the encryption intensity of important energy data, and solving the problem that when using a fixed key length to encrypt and store energy data, the encryption efficiency is low or the encryption intensity is low, which may lead to the leakage of important energy data. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1The flowchart of steps of a method for securely storing energy big data provided by an embodiment of the present application;

[0041] Figure 2 The schematic diagram of the STL algorithm decomposition of the social electricity consumption data sequence;

[0042] Figure 3 The schematic diagram of the extreme point detection of the social electricity consumption data sequence;

[0043] Figure 4 The schematic diagram of the determination process of the data regularity degree;

[0044] Figure 5 The schematic diagram of the acquisition process of the key length coefficient. Specific embodiments

[0045] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example", etc. are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "or", "for example", etc. is intended to present related concepts in a specific manner.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. It should be understood that unless otherwise stated in this application, " / " means "or".

[0047] In addition, it should be noted that the terms "first" and "second" in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0048] The following specifically describes the specific solutions of a method, a storage medium and a system for securely storing energy big data provided by this application with reference to the accompanying drawings.

[0049] Please refer to Figure 1 , which shows the flowchart of steps of a method for securely storing energy big data provided by an embodiment of the present application. The method includes the following steps:

[0050] Step 1, obtain the energy data of any type of energy, which includes multiple data categories, and each data category includes multiple groups of data. Arrange any group of data in any data category in time series to form a data sequence.

[0051] Obtain the latest energy data stored in the energy big data center. Since there are many energy categories, each energy category contains multiple data categories, and each data category contains multiple groups of data. For example, in electric power energy, the data categories include power consumption, power generation, and power grid operation data. Among the power generation data, there are thermal power generation and wind power generation, etc.

[0052] In this application, taking the electric power energy data as an example, any group of data in any data category in the electric power energy data is arranged in time series to form a data sequence. For example, the power consumption data of any enterprise is arranged in time series to form the power consumption data sequence of the said any enterprise.

[0053] Thus, data sequences of all groups of data in all data categories in the electric power energy data can be obtained.

[0054] Step 2: Determine the correlation value of each type of data based on the difference in the change trend of all corresponding data sequences between each type of data and all other types of data, and the difference in the moments where the extreme values are located within all corresponding data sequences.

[0055] In the electric power energy data, there is a certain mutual influence among various types of data. If a certain type of data has a stronger influence on other types of data, it indicates that the importance of this type of data in the electric power energy data may be higher, and a longer key needs to be given for encryption during the storage of this type of data to protect its security. Taking the i-th type of data as an example, specifically, the data change characteristics between each group of data in the i-th type of data and each group of data in other types of data are similar, that is, the times when the peak and valley values appear are close, and the change amplitudes are similar.

[0056] Based on the above analysis, determine the correlation value of each type of data based on the difference in the change trend of all corresponding data sequences between each type of data and all other types of data, and the difference in the moments where the extreme values are located within all corresponding data sequences. The expression is:

[0057] In the formula, A i is the correlation value of the i-th type of data; B i,j is the average value of the difference in the trend strength between the i-th type of data and the j-th data sequence among all other types of data; C i,j is the average value of the difference in the moments where all corresponding extreme values are located in the j-th data sequence between the i-th type of data and all other types of data; J iis the minimum value of the total number of data sequences in all categories of data; α is a first preset value greater than 0, the purpose is to prevent the denominator from being 0, and the value of α is preset manually, and the implementer can set it by himself. In this embodiment, the value of α is 0.001; among them, the method for obtaining the trend intensity of the data sequence is: first use the STL (Seasonal and Trend decomposition using Loess) algorithm to obtain the trend term and the residual term of the data sequence, and then use the trend intensity calculation formula to obtain the trend intensity of the data sequence. The trend intensity calculation formula is a well-known technology and will not be elaborated in this application. Taking the social electricity consumption data sequence in the electric power energy data as an example, the STL algorithm decomposition schematic diagram of the social electricity consumption data sequence is as Figure 2 shown; the extreme values in the data sequence can be obtained through the extreme point detection algorithm, where the extreme point detection algorithm is a well-known technology and will not be elaborated in this application. Still taking the social electricity consumption data sequence in the electric power energy data as an example, the extreme point detection schematic diagram of the social electricity consumption data sequence is as Figure 3 shown.

[0058] In this embodiment, the difference between the trend intensities is the absolute value of the difference. As other implementation manners, on the basis of being able to measure the difference between the trend intensities, the implementer can adopt other calculation methods to measure the difference between the trend intensities, such as the ratio relationship, the square of the difference, etc. The implementer can limit according to the actual situation, and this application does not make special restrictions.

[0059] In this embodiment, the difference between the moments where the extreme values are located is the absolute value of the difference. As other implementation manners, on the basis of being able to measure the difference between the moments where the extreme values are located, the implementer can adopt other calculation methods to measure the difference between the moments where the extreme values are located, such as the ratio relationship, the square of the difference, etc. The implementer can limit according to the actual situation, and this application does not make special restrictions.

[0060] It should be noted that: when the change trend of the data sequence between the i-th category of data and the data sequences of all the remaining categories of data is more similar, and the moments where the corresponding extreme values in the data sequences are closer, it means that the i-th category of data has a greater influence on other categories of data, and the correlation value is greater.

[0061] Step 3, based on the similarity between all any two data sequences in each category of data and the degree of chaos of all data sequences, determine the data regularity degree of each category of data.

[0062] Since there are certain regularities in various types of data in power energy data, such as residential electricity consumption being usually higher in the evening and lower in the early morning, and enterprise electricity consumption being higher in the peak season and lower in the off-season of the corresponding industries of enterprises. The higher the regularity of the data, the greater the possibility of being cracked and stolen after encryption. Therefore, when the data itself has a high regularity, the length of the encryption key required should also be increased. The regularity of a certain type of data is specifically manifested as that the data sequences in a certain type of data have a high regularity and there is a high similarity between the data sequences.

[0063] Based on the above analysis, based on the similarity between all any two data sequences in various types of data and the degree of chaos of all data sequences, the data regularity degree of various types of data is determined to characterize the regularity characteristics of various types of data. The expression is:

[0064] D i is the data regularity degree of the i-th type of data; E i is the average approximate entropy of all data sequences in the i-th type of data; F i is the average distance between all any two data sequences in the i-th type of data; β is a preset second value greater than 0, the purpose is to prevent the denominator from being 0, the value of β is preset by humans, and the implementer can set it by himself. In this embodiment, the value of β is 0.001; among them, the calculation of approximate entropy is a well-known technology and will not be elaborated in this application.

[0065] In this embodiment, the distance between data sequences is the DTW (Dynamic Time Warping) distance. As other implementation manners, on the basis of being able to measure the distance between two data sequences, the implementer can use other existing technologies to obtain the distance between two data sequences, such as Euclidean distance, Manhattan distance, etc., and this application does not make special restrictions.

[0066] It should be noted that: if the regularity of the data sequences in the i-th type of data is greater and the similarity between the data sequences is stronger, then the average approximate entropy is smaller and the average distance is smaller; that is, the greater the data regularity degree, the stronger the regularity existing in the i-th type of data. The schematic diagram of the determination process of the data regularity degree is as Figure 4 shown.

[0067] Step 4, based on the correlation value and the data regularity degree, obtain the encryption strength coefficient of various types of data.

[0068] The fusion result of the correlation value and the data regularity degree of various types of data is used as the encryption strength coefficient of various types of data.

[0069] It should be noted that: Fusion refers to combining multiple independent variables together in a way that enhances the overall effect, such as an additive relationship, a multiplicative relationship, etc. The implementer can set it according to the actual situation, and this application does not make special restrictions.

[0070] In this embodiment, the sum value of the correlation value and the data regularity degree of various types of data is used as the encryption strength coefficient of various types of data.

[0071] In another embodiment, the product of the correlation value and the data regularity degree of various types of data is used as the encryption strength coefficient of various types of data.

[0072] It should be noted that: Taking the i-th type of data as an example, if the i-th type of data has a stronger influence on the data of the remaining categories among all power energy data, then the correlation degree between the i-th type of data and the data of the remaining categories is greater, and the data regularity contained in the i-th type of data is stronger, indicating that the importance of the i-th type of data in all power energy data is higher, and the possibility of being cracked and stolen after encryption is greater. Therefore, a longer key is required for encryption to improve the encryption effect.

[0073] Step 5: Divide each data sequence into subsequences of each preset length, and determine the local feature approximation degree of each type of data based on the mutation moment difference and distribution range difference of the corresponding subsequences between each type of data and all other types of data.

[0074] Due to the large amount of power energy data, the data collection methods are diverse and the sources are diverse, which may lead to low data quality. Specifically, it is manifested as: there are data missing situations in the data stored in the energy big data center, resulting in null values in the data sequence, thereby affecting the overall data characteristics of some types of data in the power energy data.

[0075] For example, a data sequence with strong original regularity has its regularity decreased due to missing some data; or the correlation between a certain type of data and the data of the remaining categories was originally strong, but due to missing some data, the correlation is weakened, and finally the calculated encryption strength coefficient is low. As a result, a shorter key is used for encrypting the certain type of data, resulting in a poor encryption effect. Although there are data missing situations in this part of the data, there may still be complete data, and these complete data may be relatively important content in the power energy data, such as the addresses and identity information of residential users. If the encryption effect of this part of the data is poor, it will lead to these data being easily cracked and stolen, resulting in privacy leakage. Therefore, only selecting the key length through the encryption strength coefficient is not comprehensive and further analysis is required.

[0076] When there is a missing value in any category of data, some data features will be lost, but the non-missing part will still retain its data features, and the data features between the non-missing part of the data and other category data with which it is correlated are still similar. Taking the data of the i-th category as an example, specifically, the data mutation feature and the data change amplitude feature in a certain data interval of the j-th data sequence in the i-th category of data are still relatively similar to the data features of the data in the remaining categories.

[0077] To characterize the local features of the data sequence, taking the data of the i-th category as an example, each data sequence in the i-th category of data is divided, and every preset number of data is divided into a subsequence, and a mutation point detection algorithm is used to obtain the mutation points in each subsequence. It should be noted that: in this process, if the amount of data is insufficient, no processing is performed.

[0078] In this embodiment, the value of the preset number is 30, and the value of the preset number is preset artificially. The implementer can set it according to the actual situation, and this application does not make special restrictions.

[0079] In this embodiment, the Pettitt mutation point detection algorithm is used to obtain the mutation points in the subsequence. As other implementation manners, on the basis of being able to obtain the mutation points in the subsequence, the implementer can use other existing technologies to obtain the mutation points in the subsequence, such as the Bayesian mutation point detection algorithm, the Mann-Kendall mutation point detection algorithm, etc., and this application does not make special restrictions.

[0080] Based on the mutation time difference and the distribution range difference of the corresponding subsequences between each category of data and all the remaining category data, the local feature approximation degree of each category of data is determined. The specific process is as follows:

[0081] Calculate the difference in the time at which the corresponding mutation points of the x-th subsequence of the j-th data sequence between the data of the i-th category and the data of the remaining categories are located, and record it as the mutation time difference;

[0082] Denote the ratio of the range of each subsequence to the sequence length as the first ratio, and calculate the difference in the first ratio of the x-th subsequence of the j-th data sequence between the data of the i-th category and the data of the remaining categories, and record it as the range difference;

[0083] The expression for the local feature approximation degree of the i-th category of data is:

[0084] In the formula, G i is the local feature approximation degree of the i-th category of data; H i,j is the mean value of all the mutation time differences corresponding to the j-th data sequence in the i-th category of data; K i,j is the mean value of all the range differences corresponding to the j-th data sequence in the i-th category of data; J iis the minimum value of the total number of data sequences in all categories of data; γ is a preset third value greater than 0, the purpose of which is to prevent the denominator from being 0. The value of γ is preset manually, and the implementer can set it by himself. In this embodiment, the value of γ is 0.001.

[0085] In this embodiment, the difference between the moments where the mutation points are located is the absolute value of the difference. On the basis of being able to measure the difference between the moments where the mutation points are located, the implementer can use other calculation methods for measurement, such as ratio relationship, square of the difference, etc. The implementer can make limitations according to the actual situation, and this application does not make special restrictions.

[0086] In this embodiment, the difference between the first ratios is the absolute value of the difference. On the basis of being able to measure the difference between the first ratios, the implementer can use other calculation methods for measurement, such as ratio relationship, square of the difference, etc. The implementer can make limitations according to the actual situation, and this application does not make special restrictions.

[0087] It should be noted that: if the difference between the moments where the corresponding mutation points of the subsequences between the i-th category of data and the data of the remaining categories is smaller, and the difference in the distribution range of the subsequences is smaller, it means that the local characteristics of the data sequences between the i-th category of data and the data of the remaining categories are more similar, then the value of the local feature approximation is larger.

[0088] Step 6, based on the similarity degree of the periodic characteristics between all any two subsequences of each data sequence in each category of data, and the similarity degree of the data fluctuation situation therebetween, determine the period approximation of each category of data.

[0089] If there is a large periodicity in the original data of the i-th category of data, then even if there is a certain data missing situation in the i-th category of data, the periodic characteristics between its subsequences are still relatively similar. Specifically, it is manifested that the periodic characteristics between the subsequences of the data sequences in the i-th category of data are similar, and the degree of data fluctuation is similar.

[0090] Based on the similarity degree of the periodic characteristics between all any two subsequences of each data sequence in each category of data, and the similarity degree of the data fluctuation situation therebetween, determine the period approximation of each category of data. The expression is:

[0091] In the formula, M i is the period approximation of the i-th category of data; N i,j 、P i,jThey are respectively the mean value of the phase locking values between any two subsequences of the j-th data sequence in the i-th type of data and the sum value of the differences in the degree of dispersion therebetween; I is the total number of data sequences in the i-th type of data; ε is a fourth preset value greater than 0, the purpose of which is to prevent the denominator from being 0, and the value of ε is preset manually and the implementer can set it by himself. In this embodiment, the value of ε is 0.001. Among them, the calculation of the phase locking value is a well-known technology and will not be elaborated in this application.

[0092] In this embodiment, the degree of dispersion is the variance. As other implementation manners, on the basis of being able to describe the uneven degree of data distribution in the subsequence, the implementer can use other existing statistical measures for measurement, such as the standard deviation, the coefficient of variation, etc., and this application does not make special restrictions.

[0093] In this embodiment, the difference between the degrees of dispersion is the absolute value of the difference. On the basis of being able to measure the difference between the degrees of dispersion, the implementer can use other calculation methods for measurement, such as the ratio relationship, the square of the difference, etc., and the implementer can make a limit according to the actual situation, and this application does not make special restrictions.

[0094] It should be noted that: the stronger the periodicity of the i-th type of data, the higher the similarity degree of the periodic characteristics between the subsequences in the i-th type of data, the larger the phase locking value between the subsequences and the smaller the difference in the degree of dispersion; that is, the larger the value of the period approximation degree, the more obvious the periodicity of the i-th type of data.

[0095] Step 7, combine the local feature approximation degree, the period approximation degree and the encryption strength coefficient to obtain the key length coefficient of each type of data.

[0096] Based on the local feature approximation degree and the period approximation degree of each type of data, determine the encryption compensation coefficient of each type of data.

[0097] Take the fusion result of the local feature approximation degree and the period approximation degree of each type of data as the encryption compensation coefficient of each type of data.

[0098] In this embodiment, take the sum value of the local feature approximation degree and the period approximation degree of each type of data as the encryption compensation coefficient of each type of data.

[0099] In another embodiment, take the product of the local feature approximation degree and the period approximation degree of each type of data as the encryption compensation coefficient of each type of data.

[0100] It should be noted that: if the local features between the i-th type of data and the data of the remaining categories are more similar, and the period approximation degree of the i-th type of data is larger, then the feature correlation degree between the i-th type of data and the data of the remaining categories is larger, the importance degree of the i-th type of data is larger, and the more regular the i-th type of data is, the larger the encryption compensation coefficient of the i-th type of data is, and the longer the key length required for encryption is.

[0101] Further, based on the encryption strength coefficients and encryption compensation coefficients of various types of data, determine the key length coefficients of various types of data.

[0102] Take the fusion result of the encryption strength coefficients and encryption compensation coefficients of various types of data as the key length coefficients of various types of data.

[0103] In this embodiment, take the sum value of the encryption strength coefficients and encryption compensation coefficients of various types of data as the key length coefficients of various types of data.

[0104] In another embodiment, take the product of the encryption strength coefficients and encryption compensation coefficients of various types of data as the key length coefficients of various types of data.

[0105] It should be noted that: if the influence degree of the i-th type of data on the rest of the types of data is greater and its own regularity degree is higher, it indicates that the importance degree of this type of data in the power energy data is higher, and a longer key is required for encryption to ensure its security during storage. The schematic diagram of the acquisition process of the key length coefficient is as Figure 5 shown.

[0106] Step 8, when storing energy data, encrypt various types of data based on the key length coefficient using the AES encryption algorithm.

[0107] Use the threshold segmentation algorithm to obtain the segmentation threshold of all key length coefficients, mark the category data with key length coefficients greater than or equal to the segmentation threshold as high-confidentiality data, and mark the category data with key length coefficients less than the segmentation threshold as ordinary-confidentiality data.

[0108] In this embodiment, use the Otsu threshold segmentation algorithm to obtain the segmentation threshold. As other implementation manners, on the basis of being able to obtain the segmentation threshold, the implementer can use other existing technologies to obtain the segmentation threshold, such as global threshold segmentation, iterative threshold segmentation, etc., and this application does not make special restrictions.

[0109] When the energy big data center stores energy data, first mark the energy data as high-confidentiality data or ordinary-confidentiality data, then use the AES encryption algorithm to encrypt all the marked energy data to obtain the encrypted ciphertext, and finally store the obtained ciphertext in the database of the energy big data center. Among them, when encrypting the energy data, use a 192-bit key to encrypt the data marked as high-confidentiality data, and use a 128-bit key to encrypt the data marked as ordinary-confidentiality data. Among them, the AES encryption algorithm is a well-known technology, and this application will not elaborate further.

[0110] Based on the same inventive concept as the above method, a computer-readable storage medium is proposed. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method for secure storage of energy big data as described in the first aspect. For its specific functions and the technical effects brought, reference can be specifically made to the method embodiment section, which will not be elaborated here.

[0111] Based on the same inventive concept as the above method, an embodiment of the present application further provides an energy big data secure storage system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods for secure storage of energy big data.

[0112] In summary, when storing energy data, the present application analyzes the mutual influence situation among energy data and the law situation of the energy data itself, obtains an encryption strength coefficient, which is used to characterize the importance degree of the stored energy data and the encryption strength required, and the higher the importance degree of the energy data, the greater the encryption strength.

[0113] Furthermore, when analyzing that the data characteristics are affected due to the missing situation of energy data, the approximation situation of local characteristics among energy data and the periodic characteristics of the energy data itself are analyzed, and an encryption compensation coefficient is obtained, which can, in the case of data missing, characterize the correlation degree among energy data and the regularity of the energy data itself, and avoid misjudgment of the encryption strength required for data due to data missing; combining the encryption strength coefficient and the encryption compensation coefficient to obtain a key length coefficient to characterize the encryption strength required for energy data. Then, when using the AES encryption algorithm to encrypt and store energy data, the key length can be adaptively adjusted, which can improve the encryption strength of important energy data while ensuring the encryption efficiency, and solves the problem that when using a fixed key length to encrypt and store energy data, the encryption efficiency is relatively low or the encryption strength is relatively low, which may lead to the leakage of important energy data.

[0114] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a part thereof, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. In the description corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. Each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0115] For those skilled in the art, it is obvious that the present application is not limited to the details of the above-described exemplary embodiments, and without departing from the basic characteristics of the present application, the present application can be implemented in other specific forms. Therefore, from any point of view, the above-described embodiments of the present application should be regarded as exemplary and non-restrictive.

Claims

1. A method for secure storage of energy big data, characterized in that: The method comprises the following steps: Obtain energy data of any type of energy, which includes multiple data categories, each data category includes multiple groups of data, and arrange any group of data in any data category in time sequence to form a data sequence; Determine the correlation value of each category of data based on the difference in the change trend of all corresponding data sequences between each category of data and all other categories of data, as well as the difference in the time of extreme values ​​in all corresponding data sequences; Based on the similarity between any two data sequences in each type of data and the degree of disorder of all data sequences, the data regularity of each type of data is determined; Based on the association value and the data regularity, obtaining encryption strength coefficients of various types of data; Divide each data sequence into subsequences of preset lengths, and determine the local feature approximation of each type of data based on the difference in mutation time and distribution range between each type of data and all other types of data corresponding to the subsequences; Determine the periodic approximation of each type of data based on the similarity of the periodic characteristics between any two subsequences of each data sequence in each type of data and the similarity of the data fluctuations between them; Combining the local feature approximation, the periodic approximation and the encryption strength coefficient, obtaining a key length coefficient for each type of data; When storing energy data, various types of data are encrypted using the AES encryption algorithm based on the key length coefficient.

2. A method for secure storage of energy big data as claimed in claim 1, characterized in that: The expression of the associated value is: In the formula, A i is the correlation value of the i-th category data; B i,j is the mean of the trend intensity difference between the i-th category data and all other categories of data in the j-th data series; C i,j is the mean of the differences between the i-th category data and all other category data at the time of the corresponding extreme values ​​in the j-th data sequence; J i is the minimum value of the total number of data sequences in all categories of data; α is the first value preset to be greater than 0.

3. A method for secure storage of energy big data as claimed in claim 1, characterized in that: The process of determining the data regularity is as follows: Calculate the approximate entropy mean of all data sequences in each type of data; The mean of the distances between any two data sequences in each type of data is recorded as the mean distance; The data regularity is inversely proportional to the approximate entropy mean and the distance mean, respectively.

4. The method for secure storage of energy big data according to claim 1, characterized in that: The process of determining the local feature approximation is as follows: Let any data category be denoted as i, and calculate the difference in the time of each corresponding mutation point of the x-th subsequence of the j-th data sequence between the i-th data and the rest of the categories of data, and record it as the mutation time difference; The ratio of the range of each subsequence to the sequence length is recorded as the first ratio, and the difference of the first ratio of the x-th subsequence of the j-th data sequence between the i-th category data and the rest of the categories of data is calculated, recorded as the range difference; The expression of the local feature approximation of the i-th type of data is: In the formula, G i is the local feature approximation of the i-th type of data; H i,j K is the mean of all the mutation time differences corresponding to the j-th data sequence in the i-th category of data; i,j is the mean of all the range differences corresponding to the jth data sequence in the i-th category of data; J i It is the minimum value of the total number of data sequences in all categories of data; γ is a third value that is preset to be greater than 0.

5. The method for secure storage of energy big data according to claim 1, characterized in that: The expression of the period approximation is: Where M i is the periodic approximation of the i-th type of data; N i,j , P i,j are respectively the mean of the phase locking values ​​between all arbitrary two subsequences of the jth data sequence in the ith category of data and the sum of the differences in the degree of discreteness between them; I is the total number of data sequences in the ith category of data; ε is a fourth value preset to be greater than 0.

6. A method for secure storage of energy big data as claimed in claim 1, characterized in that: The key length coefficient is a fusion result of the local feature approximation, the periodic approximation and the encryption strength coefficient.

7. The method for secure storage of energy big data according to claim 1, characterized in that: The process of encrypting various types of data is as follows: Using the threshold segmentation algorithm to obtain the segmentation thresholds of all key length coefficients; Mark the category data whose key length coefficient is greater than or equal to the segmentation threshold as high confidentiality data, otherwise, mark it as ordinary confidentiality data; The AES encryption algorithm is used to encrypt high-confidentiality data and ordinary confidentiality data using keys of different lengths.

8. A method for secure storage of energy big data as claimed in claim 7, characterized in that: The key length when encrypting high confidentiality data is greater than the key length when encrypting ordinary confidentiality data.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements a method for secure storage of energy big data as described in any one of claims 1 to 8.

10. An energy big data security storage system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of a method for secure storage of energy big data as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • New energy intelligent settlement method and system based on information intelligent matching

    CN116089777A

  • Electric power big data security sharing method based on block chain

    CN117874556A

  • Multi-subject data cross-domain security interaction system based on encryption algorithm

    CN118094628A

  • Energy sensitive data intelligent identification method based on association mining

    CN118427244A

  • Block chain-based multi-dimensional data transmission method for power equipment

    CN119322962A

Cited By

  • Power information protection method and system based on communication security

    CN122001657A