Power distribution network data privacy protection method

Through hash division technology and Hadoop distributed file system, combined with multi-key homomorphic encryption algorithm and intermediate key-value pair compression mechanism, the problems of low computing efficiency and complex key management in large-scale data processing are solved, and efficient and secure data processing and storage are achieved.

CN120217423APending Publication Date: 2025-06-27HUZHOU ELECTRIC POWER SUPPLY CO OF STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510210146.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional encryption methods are inefficient in processing large-scale data, consume a lot of resources, and are difficult to meet the needs of real-time and high concurrency, and have high complexity in key management and data verification.

Method used

The hash division technology is used to divide the data into multiple data subsets, and the Hadoop distributed file system and MapReduce framework are used for parallel processing. Through the multi-key homomorphic encryption algorithm and the intermediate key-value pair compression mechanism, efficient data encryption and storage are achieved.

Benefits of technology

It significantly improves computing performance and processing capabilities, solves the problems of low computing efficiency, low resource utilization and complex key management in traditional methods, and ensures the security and integrity of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217423A_ABST
    Figure CN120217423A_ABST
Patent Text Reader

Abstract

The power distribution network data privacy protection method comprises the following steps: reading power distribution network data, dividing the power distribution network data into a plurality of data subsets through Hash division, and storing the data subsets to an HDFS (Hadoop Distributed File System); task scheduling is carried out through YARN, a plurality of data subsets are transmitted to corresponding processing units, homomorphic encryption is carried out, intermediate key value pairs are generated, and compression is carried out; and the Reduce reads the compressed intermediate key value pairs, decompresses the intermediate key value pairs, groups the intermediate key value pairs according to the intermediate key values, and verifies each group of intermediate key value pairs. According to the invention, through the optimized multi-key homomorphic encryption algorithm and the intelligent compression and adaptive transmission mechanism, the efficiency and security of data encryption are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network data information security, and particularly relates to a method for protecting distribution network data privacy. Background Art

[0002] With the rapid development of big data and cloud computing technologies, data security and privacy protection have become urgent problems to be solved. Especially in application scenarios such as cross-domain data sharing and federated learning, how to achieve efficient data processing and computing while protecting data privacy has become the focus of research. Traditional data encryption methods have problems such as low computing efficiency and high resource consumption when dealing with large-scale data, and it is difficult to meet the requirements of real-time and high concurrency. Therefore, the present invention aims to provide a fast homomorphic encryption calculation method based on parallel computing to solve these problems and achieve efficient and secure data processing and computing.

[0003] Traditional data encryption methods mainly rely on serial processing methods, such as classic encryption algorithms like AES and RSA. Although these methods can provide high data security, when dealing with large-scale data, they have low computing efficiency, high resource consumption, and it is difficult to meet the requirements of real-time and high concurrency. In addition, although traditional homomorphic encryption algorithms (such as Paillier, BGN, etc.) can perform calculations in the encrypted state, due to their high computational complexity, there are obvious performance bottlenecks in practical applications. When dealing with large-scale data, these traditional methods often take a long time and are difficult to scale to a distributed computing environment.

[0004] In existing technical solutions, some studies have tried to improve the effect of data privacy protection through distributed computing frameworks and homomorphic encryption algorithms. For example, a big data encryption storage method is disclosed on the Chinese Patent Network, with the application number: CN20231131646.1. In this patent, the Hadoop distributed computing framework is used for data storage and processing, and the Paillier homomorphic encryption algorithm is used to encrypt the data. However, these solutions still have some deficiencies. First, existing solutions usually require complex task scheduling and data partitioning mechanisms, increasing the complexity of the system and the development difficulty. Second, existing solutions still face challenges in data transmission and synchronization when dealing with large-scale data, resulting in limited overall performance. In addition, although existing homomorphic encryption algorithms have to some extent solved the problem of data encryption storage, in practical applications, the complexity of key management and data verification is relatively high, affecting the usability and reliability of the system. Summary of the Invention

[0005] The object of the present invention is to solve the problem of difficult verification of data integrity and correctness in traditional encryption.

[0006] Another object of the present invention is to solve the problems of complex traditional encryption task scheduling and low resource utilization rate.

[0007] To achieve the above object, the technical solution of the present invention is as follows.

[0008] Read the distribution network data, divide the distribution network data into several data subsets through hash partitioning, store the data subsets in HDFS; perform task scheduling through YARN, transfer several data subsets to the corresponding processing units, perform homomorphic encryption, generate intermediate key-value pairs, and perform compression; Reduce reads the compressed intermediate key-value pairs, decompresses them, groups them according to the intermediate keys, and verifies each group of intermediate key-value pairs.

[0009] Preferably, the steps of the task scheduling include: transferring the data subsets to the master node RM of YARN, parsing the data content in the data subsets through RM, determining the required parsing resource amount, determining the distribution of the data subsets in HDFS through the task node AM of YARN, and allocating them to the processing units in the slave node NM of YARN for processing.

[0010] Preferably, the steps of the verification include: obtaining the ciphertext sum and plaintext sum of each group of intermediate key-value pairs, encrypting the plaintext sum by the private key of each group of intermediate key-value pairs through Reduce to obtain the decrypted ciphertext sum, comparing the decrypted ciphertext sum with the plaintext sum, and obtaining the comparison result.

[0011] Preferably, if the comparison results are consistent, the integrity verification passes; if the comparison results are inconsistent, the integrity verification fails.

[0012] Preferably, the steps of the compression include: obtaining the coding probability distribution of each group of intermediate key-value pairs, and performing arithmetic coding according to the coding probability distribution to obtain the compression code of each group of intermediate key-value pairs.

[0013] Preferably, the steps of the decompression include: decompressing the arithmetic coding of each group of intermediate key-value pairs according to the coding probability distribution of each group of intermediate key-value pairs.

[0014] Preferably, the steps of the hash partitioning include: obtaining the hash value of each data, constructing a hash ring according to the sizes of the hash values of all data, and partitioning the data according to the hash ring.

[0015] Preferably, the steps of the grouping include: dividing the intermediate key-value pairs with the same identification element into the same group according to the identification element of the intermediate key-value pairs.

[0016] Preferably, for each group of intermediate key-value pairs, perform a quick sort according to the identification of the intermediate key-value pairs to generate a sorted list.

[0017] Preferably, the middle key-value pairs in the sorted list are merged into an encrypted data file, which is split into multiple data blocks of a fixed size through a data chunking mechanism, and a checksum is generated for each data block to ensure the integrity and consistency of the data. Then, the data blocks and their checksums are written into HDFS through the API of Hadoop to achieve encrypted data storage.

[0018] Compared with the prior art, the technical solution of this application has the following technical effects: By adopting the hash partitioning technology and the Hadoop Distributed File System (HDFS) for data partitioning and storage, the present invention solves the problem of low computational efficiency of traditional homomorphic encryption algorithms when dealing with large-scale data; the hash partitioning technology efficiently divides the data set to be encrypted into multiple data subsets and uploads them to HDFS, ensuring the uniform distribution and efficient storage of the data. This not only reduces the burden on a single processing unit but also improves the data reading and writing speeds, thus significantly enhancing the computational performance and processing capacity of the entire system.

[0019] By using the MapReduce framework and YARN of Hadoop for task scheduling, the present invention solves the problems of complex task scheduling and low resource utilization in traditional parallel computing; the Map stage in the MapReduce framework is responsible for reading and processing data subsets, while YARN ensures that the data subsets are executed on the processing unit closest to their storage location through a data locality optimization mechanism, reducing the latency of data transmission. This efficient task scheduling mechanism not only improves the utilization rate of computing resources but also accelerates the data processing process, enabling the system to maintain high performance and high response speed in large-scale data processing.

[0020] By means of the multi-key homomorphic encryption algorithm and the middle key-value pair compression mechanism, the present invention solves the problems of low key management and data transmission efficiency in traditional homomorphic encryption algorithms; the multi-key homomorphic encryption algorithm allows each data subset to generate an independent key pair, solving the complexity of key holding and management and improving the flexibility and security of the system. At the same time, the middle key-value pair compression mechanism reduces the amount of intermediate data transmission and storage space occupancy, reducing the consumption of network bandwidth and storage resources, and further enhancing the overall performance and efficiency of the system.

[0021] The present invention solves the problem of difficult verification of data integrity and correctness in traditional homomorphic encryption algorithms by verifying and sorting intermediate key-value pairs in the Reduce stage; the integrity verification and correct value verification mechanisms ensure that the encryption results of each data subset are accurate and error-free, avoiding data tampering and error propagation. In addition, through the quicksort algorithm and the data chunking mechanism, the verified intermediate key-value pairs are efficiently sorted and merged to generate the final encrypted data file. This not only ensures the order and consistency of the data, but also improves the data access and management efficiency, making the system more reliable and efficient when processing large-scale data.

[0022] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, so as to be implemented in accordance with the content of the specification, and in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following takes the preferred embodiments of the present application and combines the drawings to describe in detail as follows.

[0023] Those skilled in the art will understand the above and other purposes, advantages and features of the present application more clearly according to the following detailed description of the specific embodiments of the present application in conjunction with the drawings. Brief Description of the Drawings

[0024] Figure 1 It is a schematic diagram of the method related to the present invention. Detailed Description of the Specific Embodiments

[0025] Embodiment 1. This embodiment mainly describes a fast homomorphic encryption calculation method based on parallel computing, as Figure 1 shown, including the following steps: S1. Read the data file, divide the data set to be encrypted in the data file into multiple data subsets through the hash partitioning technology, and upload them to be stored in HDFS of the Hadoop distributed file system; Further, the hash partitioning technology includes the following content; Define the data set D to be encrypted, the data set D to be encrypted contains m data items, that is, D = {d1, d2 ··· d m}, where d1 is the first data item; calculate the hash value h(d i ) of each data item d i , where i is the i-th data subset; allocate the data item d i to the h(d i ) + 1-th data subset The formula is: D i = {d j | h(d j ) = i - 1, j ∈ {1, 2,..., m}}, where d jDenote the j-th data item, where i ∈ {1, 2, …, n}, D i Denote the i-th data subset; divide the data set D into n data subsets D1, D2 ··· D n ; By dividing the data set to be encrypted in the data file into multiple data subsets through the hash partitioning technique and uploading them to HDFS, the problem of low computational efficiency in large-scale data processing is solved, and the data reading and writing speeds are improved.

[0026] S2. Use Map in the MapReduce framework of the Hadoop distributed file system to read the data subsets in HDFS, and perform task scheduling through YARN of the Hadoop distributed file system to execute multiple data subsets in Map in parallel on multiple processing units; Preferably, the task scheduling in S2 is performed using a data locality optimization mechanism, and the data locality optimization mechanism includes the following: transfer the data subset to the ResourceManager (RM) of YARN, parse the data content in the data subset through RM to determine the required amount of parsing resources, determine the distribution of the data subset in HDFS through the ApplicationMaster (AM) of YARN, and allocate it to the processing units in the NodeManager (NM) of YARN for processing; Through the data locality optimization mechanism of YARN for task scheduling, the problems of data transmission delay and low resource utilization are solved, and the utilization efficiency of computing resources and the task processing speed are improved.

[0027] S3. Encrypt the data subsets in Map on the processing units through a multi-key homomorphic encryption algorithm to generate intermediate key-value pairs, and transfer the intermediate key-value pairs to Reduce in the MapReduce framework for processing; Furthermore, the multi-key homomorphic encryption algorithm for the data subset pair includes the following: key generation for the data subset, generate a pair of public key pk i and private key sk i for each data subset D i , and the formula is: (pk i , sk i ) ← KeyGen(λ), where λ is the security parameter and KeyGen is the key generation algorithm; read each data subset D i through Map, and encrypt the data item d i using the public key pk i , and the encryption formula is: c ij ← Encrypt(pk i , d j ), where c ij is the data item d jThe ciphertext, Encrypt is the encryption algorithm; generate the intermediate key-value pair <k i ,c ij >, where k i is the unique identifier of the data subset D i , and c ij is the ciphertext of the data item d j ; Generate the key pairs for each data subset through the multi-key homomorphic encryption algorithm, solve the key management and data security problems, and improve the flexibility and security of the system.

[0028] Furthermore, the multi-key homomorphic encryption algorithm also includes an intermediate key-value pair compression mechanism, which compresses the generated intermediate key-value pair <k i ,c ij > through the intermediate key-value pair compression mechanism to generate the compressed intermediate key-value pair <k i ,c ij > c , and the formula is: <k i ,c ij > c ←Snappy(<k i ,c ij >), where Snappy is the compression algorithm; Through the intermediate key-value pair compression mechanism, the transmission volume and storage space occupation of intermediate data are reduced, the problems of low data transmission efficiency and waste of storage resources are solved, and the overall performance of the system is improved.

[0029] S4. Reduce reads the intermediate key-value pairs generated by Map encryption, groups them according to the intermediate key-value pairs, and verifies the intermediate key-value pairs of each group to ensure all encryption; Furthermore, in the S4, the compressed intermediate key-value pair <k i ,c ij > c is transmitted to Reduce through the ShuffleandSort mechanism of Hadoop to decompress the intermediate key-value pair and restore the intermediate key-value pair <k i ,c ij >. Group according to the k i ,c ij key in the intermediate key-value pair <k i , and those with the same k i key are divided into the same group. The formula is: G i ={<k i ,c i1 ,<k i ,c i2 ,…,<k i ,c im >}, where Gi Denoted as k i of the group; By transmitting and decompressing the intermediate key-value pairs through the Shuffle and Sort mechanism of Hadoop, the efficiency problem of data grouping and transmission is solved, ensuring the correct grouping and efficient processing of data.

[0030] Furthermore, the ciphertext within each group G i is homomorphically encrypted, and the formula is: C i ← Add(c i1 , c i2 , …, c im ), where Add is the homomorphic encryption algorithm, and C i is the encrypted result of the group; Through homomorphic encryption operations, the intermediate key-value pairs <k i , C i > of the group are generated; By processing the ciphertext within each group through the homomorphic encryption algorithm, the computational complexity problem in the data encryption process is solved, improving the encryption efficiency and data security.

[0031] Furthermore, the verification of the intermediate key-value pairs <k i , C i > of the group in S4 includes integrity verification and correct value verification; The integrity verification is carried out by obtaining all the ciphertexts {c i , c i1 , c i2 ··· c im} in G i , re-executing the homomorphic encryption operation to generate a new ciphertext C' i , and the formula is: C' i1 ← Add(c i2 , c im ), comparing C i with C' i , and judging whether the two are equal. If the two are equal, the integrity verification passes; otherwise, the verification fails; The correct value verification uses the private key sk i to decrypt the ciphertext C i to obtain the decrypted d i , and the formula is: d i ← Decrypt(sk i , C i ), where Decrypt is the decryption algorithm; Obtain all the plaintexts {d i , d i1 , d i2 ··· d im} in the group Gi , the formula is: d' i ←Add(d i1 , d i2 , …, d im ), compare d i and d' i , and determine whether the two are equal. If the two are equal, the correct value verification passes; otherwise, the verification fails; Through the integrity verification and correct value verification mechanisms, the integrity and correctness of the data are ensured, the problems of data tampering and error propagation are solved, and the reliability and security of the system are improved.

[0032] S5. After the verification is completed, sort the groups processed by Reduce, merge them into an encrypted data file, and output through HDFS of the Hadoop distributed file system; Furthermore, the intermediate key-value pairs <k i , C i > in the S5 are sorted after passing the verification. All the intermediate key-value pairs <k i , C i > that pass the verification are collected into the list L. The formula is: L = {<k1, C1>, <k2, C2>, …, <k n , C n >}. Through quicksort, select a pivot element <k p , C p > from the list L, divide the list into two parts, one part contains the intermediate key-value pairs smaller than the pivot element, and the other part contains the intermediate key-value pairs larger than the pivot element. The formula is: L left = {<k i , C i >|k i < k p}, L right = {<k i , C i >|k i > k p}, where L left is the intermediate key-value pairs smaller than the pivot element, and L right is the intermediate key-value pairs larger than the pivot element; sort L left and L right through the recursive sorting mechanism, and merge the sorted left half, the pivot element, and the sorted right half to generate the final sorted list L s ; Sorting the intermediate key-value pairs through the quicksort algorithm solves the efficiency problem of data sorting and merging, and improves the orderliness and access efficiency of the data.

[0033] Further, the middle key-value pair <k i ,C i > in the sorted list is merged into an encrypted data file, which is divided into multiple data blocks of a fixed size through a data chunking mechanism, and a checksum is generated for each data block to ensure the integrity and consistency of the data. The data blocks and their checksums are written into HDFS through the API of Hadoop to achieve encrypted data storage; Through the data chunking mechanism and the generation of checksums, the integrity and consistency of the data are ensured, the reliability problems of data storage and access are solved, and the stability and data security of the system are improved.

[0034] In this embodiment, the dataset to be encrypted is divided into multiple data subsets by using the hash partitioning technique, and the efficient parallel processing of the data is realized by means of the Hadoop Distributed File System (HDFS) and the MapReduce framework. The multi-key homomorphic encryption algorithm is adopted to ensure the security of the data. The intermediate key-value pair compression mechanism is introduced to improve the data processing efficiency. Through the ShuffleandSort mechanism of Hadoop and the subsequent verification, sorting and merging processes, the accuracy and consistency of the encrypted data are ensured. The encrypted data file is output through HDFS, realizing the efficient and secure encrypted storage of large-scale datasets. This method significantly improves the speed and efficiency of homomorphic encryption calculation while maintaining the privacy and security of the data.

[0035] Embodiment 2. Based on Embodiment 1, this embodiment details an optimization scheme of the multi-key homomorphic encryption algorithm in a fast homomorphic encryption calculation method based on parallel computing in the present application. Although the multi-key homomorphic encryption algorithm solves the key holding problem, there are still problems of complex key management and low calculation efficiency in practical applications. To further improve the performance of the multi-key homomorphic encryption algorithm, an optimization scheme of the multi-key homomorphic encryption algorithm is introduced in the present application; A pair of public key pk i and private key sk i is generated for each data subset, and the formula is: (pk i ,sk i )←KeyGen(λ), where λ is a security parameter and KeyGen is a key generation algorithm. The generated key pair is registered on the blockchain, recording the generation time, generator and usage scope of the key, and ensuring the security and traceability of the key through a smart contract; when distributing the key through the blockchain network, each participating party can safely obtain and update the key, and the key distribution and update process is automatically executed through the smart contract to ensure the timeliness and effectiveness of the key; For the efficient encryption of data subsets, encryption calculations are performed using a GPU or a dedicated encryption chip TPM, including the following: In the Map stage, each data subset is read through Map, and the public key pk is used i to encrypt the data item d i , and the encryption formula is: c ij ←Encrypt(pk i , d j ), where c ij is the ciphertext of the data item d j , Encrypt is the encryption algorithm; generate the intermediate key-value pair <k i , c ij >, where k i is the unique identifier of the data subset D i , and c ij is the ciphertext of the data item d j ; In this embodiment, through an optimized multi-key homomorphic encryption algorithm, the key management mechanism based on the blockchain ensures the security and traceability of the keys, improves the security of the system, uses a GPU or a dedicated encryption chip for encryption calculations, significantly improves the encryption speed and calculation efficiency, and through optimized key management and an efficient encryption algorithm, reduces the latency of data processing and transmission, and improves the real-time performance of the system.

[0036] Embodiment 3. This embodiment details the compression optimization scheme and the application of the adaptive transmission mechanism in a fast homomorphic encryption calculation method based on parallel computing in this application; Regarding the traditional data compression and transmission methods, when dealing with large-scale data, there are problems such as large transmission latency and bandwidth consumption, which affect the overall performance of the system. In this embodiment, through an optimized intelligent compression and adaptive transmission mechanism, these problems are solved, and the security and performance of the system are further improved; In the Map stage, through machine learning algorithms, analyze the characteristics and patterns of the data, determine the most suitable compression algorithm and parameters, and perform dynamic compression; according to the analysis results, dynamically select the compression algorithm and parameters to compress each data subset. The specific steps are as follows: Obtain the size of the intermediate key-value pair data set. If the data set size is greater than the threshold, arithmetic coding is used for compression; if the data set size is less than the threshold, machine learning algorithms are used for compression.

[0037] Embodiment 4. Based on Embodiment 3, this embodiment details the arithmetic coding compression optimization scheme based on arithmetic coding in the Map stage, specifically as follows: Obtain the coding probability distribution of each group of intermediate key-value pairs, and perform arithmetic coding according to the coding probability distribution of each group of intermediate key-value pairs.

[0038] Arithmetic coding is a lossless data compression algorithm based on a probability model, and its compression process closely depends on the probability distribution of the data. First, a probability model needs to be established for the data to be compressed, and the probability of each symbol appearing is estimated, which is usually achieved through statistical analysis of the data. For example, in text compression, the occurrence frequencies of letters, words, or bytes can be counted to obtain the probability distribution of each symbol. Generally speaking, the greater the probability of a symbol, the fewer bits required for its encoding, and vice versa. In this way, compression can be achieved by taking advantage of the probability differences.

[0039] Next, arithmetic coding regards the entire encoding process as an operation on the probability interval [0, 1). The initial value range of this interval is from 0 to 1. For each symbol processed, the current interval is narrowed down according to the probability distribution of that symbol. For example, assume the first symbol is 'A' with a probability of 0.5. Then the initial encoding interval [0, 1) is divided into two parts: 0 - 0.5 represents 'A', and 0.5 - 1 represents other symbols. Then process the second symbol. Assume it is 'B' with a probability of 0.3. At this time, in the previous interval of 0 - 0.5 (the sub - interval corresponding to 'A'), it is further divided. For example, the sub - interval of 'B' may be 0.5*(0.7 - 0.4) = 0.5*0.3 = 0.15. Then the new interval range may become 0.2 to 0.35, depending on the cumulative probability distribution of each symbol. This process is repeated continuously, and each symbol will cause the current interval range to be further narrowed down.

[0040] As the data sequence is continuously processed, the interval range gradually decreases, and finally the encoding interval will converge to a small numerical range. Any value within this range can be selected as the encoding result. Usually, a binary number within the interval is selected as the final encoding value. For example, assume that after a series of symbol processing, the encoding interval is narrowed down to [0.333, 0.375). Then the binary representation of 0.333 (for example, it may be 0.01010101...) or a shorter approximation can be selected as the encoding. This value is then converted into a binary bit stream, thus completing the data compression.

[0041] During the entire compression process, the probability distribution plays a decisive role because it directly determines the proportion of each symbol in the interval division. Symbols with high probabilities will result in larger sub - intervals, so that these symbols can be represented with fewer bits during encoding, while symbols with low probabilities require more bits. In this way, arithmetic coding can utilize the statistical redundancy in the data to achieve efficient compression. Compared with traditional Huffman coding, this method has better compression performance when dealing with symbols with similar probability distributions because it can allocate the encoding space more precisely and is not restricted by the requirement that symbols in Huffman coding must correspond to integer numbers of bits.

[0042] In addition, arithmetic coding also involves the update of dynamic probability models. When processing a data stream, as data is continuously input, the probability model can be updated in real time to reflect the current probability distribution of symbols. This method, such as adaptive arithmetic coding, can automatically adjust probability estimates when processing different data sources, thereby improving compression efficiency. For example, when processing a text, the beginning may contain more common words, and then it may enter an area with a high density of technical terms. Adaptive arithmetic coding can automatically adjust the probability model so that the probabilities corresponding to common words remain relatively high, while the probabilities of technical terms gradually increase, thus achieving effective compression in both cases.

[0043] Arithmetic coding also utilizes the cumulative distribution function (CDF) of probabilities to determine the division points of each symbol within the current interval. For each symbol, its corresponding cumulative probability value determines the splitting position of the current interval. For example, if there are three symbols with probabilities P(A)=0.4, P(B)=0.3, and P(C)=0.3 respectively, then their cumulative probabilities are 0.4, 0.7, and 1. When processing symbols, the start and end points of the current interval are adjusted according to these cumulative probability values. This enables arithmetic coding to accurately divide the interval even when the probabilities of symbols are very close, thereby effectively utilizing the probability space and avoiding the problem of low coding efficiency for symbols with similar probabilities in Huffman coding.

[0044] During the decoding process, arithmetic coding also relies on the probability distribution. The decoder needs to have the same probability model as the encoder in order to accurately re-partition the compressed numerical value into the original symbol sequence. During decoding, each symbol is determined sequentially from left to right. Based on the range of the current numerical value and the probability distribution of the symbols, the current symbol is judged, and then the interval range is narrowed down to continue processing the remaining symbols. This process needs to be exactly the same as the interval division during the encoding process to ensure that the recovered symbol sequence is the same as the original data. Therefore, the accuracy of the probability distribution is also crucial during the decoding process, and any deviation in probability estimation may lead to decoding errors.

[0045] To further improve compression efficiency, arithmetic coding is often used in combination with preprocessors, such as predictive coding or pattern recognition algorithms. These preprocessors can change the statistical characteristics of the data to make it more suitable for the efficient compression conditions of arithmetic coding. For example, in image or video compression, the preprocessor can generate residual data by predicting pixel values or differences between frames. The probability distribution of this residual data may be more conducive to compression by arithmetic coding. In addition, arithmetic coding can also be combined with context models to adjust probability estimates according to the context environment of the current symbol. For example, in text data, the previous characters can be considered to predict the probability distribution of the next character, thereby achieving a higher compression ratio.

[0046] Regarding the problem of transmission delay, it is solved through an adaptive transmission mechanism: The adaptive transmission mechanism dynamically adjusts the priority and path of data transmission by monitoring the network status in real time, including indicators such as bandwidth, delay, and packet loss rate; Optimize the path, and select the optimal transmission path according to the network status and the actual situation of computing resources to reduce transmission delay and bandwidth consumption; The specific steps are as follows: Obtain network status data in real time through network monitoring tools; Select the optimal transmission path according to the network status to ensure the efficiency and reliability of data transmission; Regarding the priority adjustment of the adaptive transmission mechanism, dynamically adjust the priority of data transmission according to the importance and urgency of the data to ensure the priority transmission of important data; The specific steps are as follows: Classify the data and set priorities according to importance and urgency; During the transmission process, give priority to transmitting high-priority data to ensure the timeliness and accuracy of important data.

[0047] The adaptive transmission mechanism reduces transmission delay and bandwidth consumption by monitoring the network status in real time and dynamically adjusting the transmission path, improving the reliability and efficiency of data transmission.

[0048] The application of this embodiment in the cross-domain sharing of power data is fully demonstrated, highlighting the advantages and feasibility of the optimized intelligent compression and adaptive transmission mechanism in practical applications.

[0049] Example 5, this example is based on Example 3 and details the data machine learning compression optimization scheme based on machine learning in the Map stage, as follows: Segment the original data stream, divide a large amount of data into multiple smaller data subsets according to specific rules such as time series, content type, or geographical location, and by introducing a machine learning model, analyze the characteristics and patterns of each data subset to determine the appropriate compression algorithm and parameter settings; The machine learning model not only stays at the level of selecting algorithms, but can also go deep into the parameter adjustment stage, that is, fine-tune the specific parameters of the compression algorithm for different types of input data to achieve the best effect. During this process, the model will automatically generate a series of key-value pairs as an intermediate representation form. Among them, the key represents the position identifier of the original data segment, and the value is the binary data corresponding to these segments after being processed by a specific compression algorithm. This simplifies the data mapping process during subsequent decompression and ensures that even in the face of massive data, each compressed data unit can be effectively managed and traced.

[0050] By adopting an intelligent decision-making mechanism in data compression, the entire compression process can dynamically adapt according to the real-time changing data characteristics, avoiding the inefficiency problem caused by traditional static configurations. Through the generated compressed intermediate key-value pairs, both the information integrity of the original data structure is retained, and the actual occupied space size is greatly reduced, providing a solid foundation for the next step of operations.

[0051] Through preprocessing and postprocessing techniques, the execution efficiency of the compression algorithm is optimized, and the computing time and resource consumption during the compression process are reduced to further improve the overall performance of the machine learning-based data compression solution. We incorporated advanced preprocessing and postprocessing techniques into the design of the compression algorithm, aiming to optimize its execution efficiency and significantly reduce the computing time and resource consumption.

[0052] In the preprocessing stage, operations including but not limited to removing redundant information, standardizing data formats, and preliminary dimensionality reduction are implemented on the data subset about to enter the compression pipeline, accelerating the running speed of the compression algorithm and improving the quality of the compression result. After entering the compression link, the machine learning model will automatically select the optimal compression strategy according to the characteristics of the preprocessed data and start to execute the specific compression task; after compression is completed, postprocessing optimization is carried out. The core of postprocessing lies in evaluating the intermediate states generated during the compression process and finding potential improvement points. For example, it is found that some compression steps lead to unnecessary repeated calculations, and a caching mechanism is introduced to save the intermediate results to avoid re-computation when the same situation is encountered next time.

[0053] By using a parallel computing framework to accelerate the compression tasks in a multi-threaded environment, the multi-core or multi-processor capabilities provided by modern computer hardware are fully utilized.

[0054] For evaluating the quality and compression ratio of the compressed data, ensuring that the compressed data has a higher compression ratio while maintaining high quality, the difference degree before and after compression is measured by quantitative indicators such as peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) to evaluate the quality and compression ratio of the compressed data; The intelligent compression algorithm based on machine learning can dynamically select the optimal compression strategy, achieve a higher compression ratio and faster compression speed, and reduce the occupation of storage space; this process starts with the learning of a large amount of historical data. The machine learning model gradually accumulates rich empirical knowledge through in-depth analysis of past compression instances and can accurately predict the best compression methods for different types of data. Whenever a new data subset needs to be processed, the model will first extract its features, capture key attributes such as data distribution, frequency spectrum pattern, and correlation structure, and then combine the existing knowledge base to quickly formulate a personalized compression plan.

[0055] The intelligent compression algorithm does not complete the task in one go, but runs through the entire compression cycle, continuously monitoring the compression progress and making immediate responses according to the actual situation. For example, if it is found during compression that the current strategy fails to achieve the expected effect, the model can immediately switch to an alternative plan or even retrain itself to adapt to the newly emerging situation. To further improve the compression speed, technologies such as GPU acceleration, distributed computing, and memory-mapped I / O are utilized to shorten the time required for a single compression task.

[0056] This embodiment details the efficient compression effect and provides users with a flexible, reliable, and easily extensible data management platform. Whether dealing with daily operations or sudden large-traffic scenarios, it can handle them with ease, ensuring the security and economy of data transmission and storage.

[0057] The present invention deeply elaborates its purpose, technical solution, and beneficial effects through specific embodiments. However, these embodiments are only examples to demonstrate the application mode of the invention and do not constitute a limitation on the protection scope of the present invention. We clearly state that any reasonable modification, equivalent replacement, or technical improvement under the guidance of the spirit and principle of the present invention should be included in the protection scope of the present invention. This means that as long as these changes do not deviate from the core idea and basic functions of the invention, they should be protected by the patent right. The protection scope of the present invention should be broad, including all direct and obvious variants as well as non-obvious innovations that can be reasonably deduced by technical experts based on the disclosed content of the present invention. This broad protection aims to promote further research and development based on the present invention while ensuring that its innovation and practicality are comprehensively protected by law.

Claims

1. A distribution network data privacy protection method, characterized in that: include: Read the distribution network data, divide it into several data subsets through hash partitioning, and store the data subsets in HDFS; Task scheduling is performed through YARN, and several data subsets are transmitted to the corresponding processing units for homomorphic encryption, generating intermediate key-value pairs, and compressing them. Reduce reads the compressed intermediate key-value pairs, decompresses them, groups them according to the intermediate key values, and verifies each group of intermediate key-value pairs.

2. A distribution network data privacy protection method according to claim 1, characterized in that: The task scheduling steps include: transmitting the data subset to the master node RM of YARN, parsing the data content in the data subset through RM, determining the required parsing resources, determining the distribution of the data subset in HDFS through the task node AM of YARN, and allocating it to the processing unit in the slave node NM of YARN for processing.

3. A distribution network data privacy protection method according to claim 1, characterized in that: The verification step includes: obtaining the ciphertext sum and the plaintext sum of each group of intermediate key-value pairs, encrypting the plaintext sum by using the private key of each group of intermediate key-value pairs to obtain the decrypted ciphertext sum, comparing the decrypted ciphertext sum with the plaintext sum, and obtaining a comparison result.

4. A distribution network data privacy protection method according to claim 3, characterized in that: If the comparison results are consistent, the integrity verification passes; if the comparison results are inconsistent, the integrity verification fails.

5. A distribution network data privacy protection method according to any one of claims 1 to 4, characterized in that: The compression step includes: obtaining the coding probability distribution of each group of intermediate key-value pairs, and performing arithmetic coding according to the coding probability distribution to obtain the compressed code of each group of intermediate key-value pairs.

6. A distribution network data privacy protection method according to claim 5, characterized in that: The decompression step includes: decompressing the arithmetic coding of each group of intermediate key-value pairs according to the coding probability distribution of each group of intermediate key-value pairs.

7. A distribution network data privacy protection method according to any one of claims 1 to 4, characterized in that: The step of hash partitioning includes: obtaining a hash value of each data, constructing a hash ring according to the size of the hash values ​​of all data, and partitioning the data according to the hash ring.

8. A distribution network data privacy protection method according to any one of claims 1 to 4, characterized in that: The grouping step includes: according to the identification elements of the intermediate key-value pairs, grouping the intermediate key-value pairs with the same identification elements into the same group.

9. A distribution network data privacy protection method according to any one of claims 1 to 4, characterized in that: For each group of intermediate key-value pairs, quickly sort them according to their identifiers to generate a sorted list.

10. A distribution network data privacy protection method according to claim 9, characterized in that: The intermediate key-value pairs in the sorted list are merged into an encrypted data file, which is divided into multiple data blocks of fixed size through a data block mechanism, and a check code is generated for each data block to ensure the integrity and consistency of the data. The data block and its check code are written into HDFS through Hadoop's API to realize data encryption storage.