Data hierarchical aggregation and query method and system for privacy protection in 6G network
By adopting the data hierarchical aggregation and query method of secure multi-party computing technology in 6G network, the problems of data privacy protection and computing efficiency in AI services are solved, and efficient and secure data processing and query services are achieved.
Patent Information
- Application Number
- CN202510064077.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-15
Smart Images

Figure CN119997121A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technology, and specifically relates to a data hierarchical aggregation and query method and system for privacy protection in a 6G network. Background Art
[0002] With the widespread application of the fifth generation mobile communication technology (5G) and its gradual evolution to the sixth generation (6G), the low-latency network characteristics of 6G not only greatly meet the growing communication needs, but also provide strong support for the popularization of artificial intelligence (AI) services. The AI services deployed in the network can collect a large amount of user data through distributed edge nodes in real time, and use this data to continuously optimize the model capabilities, aiming to provide users with high-quality online intelligent services. In this context, hierarchical aggregation and query of data have become one of the key technologies, especially in scenarios where a large amount of user data needs to be processed and privacy protection is required.
[0003] In current research, data hierarchical aggregation technology has made some progress. These technologies usually take advantage of edge computing and cloud computing to distribute user data to multiple edge nodes and central nodes for processing. In the paper "Research on secure hierarchical multidimensional data aggregation assisted by fog computing", scholars from Nanjing University of Posts and Telecommunications applied hierarchical network models, fog computing, data aggregation, homomorphic encryption and blockchain technologies to the industrial Internet of Things, thereby detecting equipment failures and system weaknesses in real time, effectively solving problems such as latency, bandwidth consumption and data security. The patent "An edge-assisted privacy-preserving multidimensional data hierarchical aggregation method" (application number CN202310206321.6, application publication number CN116361827A) applied by Lanzhou University of Technology discloses a hierarchical aggregation framework assisted by edge computing. The framework combines the Paillier algorithm for encryption and decryption operations to ensure the privacy security of data during the aggregation process, and uses Horner's rule to efficiently parse the fine-grained results of multidimensional data from the aggregated ciphertext, thereby realizing privacy aggregation of multi-regional multidimensional data. However, the shortcomings of these solutions and inventions are: homomorphic encryption technology requires a large computing overhead and cannot guarantee the availability and efficiency of model training under massive data. At the same time, blockchain technology has problems such as poor scalability, insufficient privacy protection, slow transaction speed, high cost and security risks. These defects make it difficult for the solution to be widely used in industrial deployment.
[0004] With the popularization of AI services, especially the widespread application of classification tasks in various fields, higher requirements are placed on data security and privacy protection. As a classic classification method, the Naive Bayes algorithm has been widely used in data classification queries due to its simplicity and efficiency. However, in the existing hierarchical aggregation and query systems, the Naive Bayes algorithm is often directly applied to plaintext data, resulting in the risk of user privacy leakage. Scholars from the University of Leeds applied the Naive Bayes algorithm to 6G network health data analysis to predict stroke risk in the paper "Patient-Centric HetNets Powered by Machine Learning and Big Data Analyticsfor 6G Networks". The study used medical records to obtain feature variables, achieved predictions through a Naive Bayes classifier, and achieved good results. However, this method lacks security protection in data use and poses a risk of privacy leakage. The patent "Fast training and lightweight prediction method of multi-party privacy-preserving naive Bayesian model" (application number CN202310389980.8, publication number CN116405184A) applied by Hunan University discloses a multi-party privacy-preserving naive Bayesian model training method based on NewPai homomorphic encryption, which designs distributed threshold decryption and intermediate result protection. However, this method also faces the problem of excessive computational and communication overhead brought by homomorphic encryption, which seriously affects the efficiency and availability of model training.
[0005] In summary, the problems of existing technologies are as follows: (1) AI services in 6G networks have made progress in data processing and utilization, but have neglected the protection of original data, resulting in the risk of leakage of users' sensitive information, especially in scenarios where a large amount of user data needs to be processed and privacy needs to be guaranteed. (2) Data hierarchical aggregation schemes based on technologies such as homomorphic encryption and blockchain have solved the problems of data security and privacy protection to a certain extent, but they still face challenges such as high computational overhead, poor scalability, insufficient privacy protection, slow transaction speed, and high cost, which limit the widespread application of these schemes in industrial deployment. (3) As a classic classification method, the naive Bayes algorithm has been widely used in data classification queries, but in existing hierarchical aggregation and query systems, the algorithm is often directly applied to plaintext data and lacks a security protection mechanism, further exacerbating the risk of user privacy leakage. (4) Although some studies have attempted to combine homomorphic encryption and differential privacy technologies to protect naive Bayes models, these methods still face problems such as high computational overhead, reduced data practicality, and impaired model performance, making them difficult to be widely used in practical application scenarios. Summary of the invention
[0006] In order to solve the above problems existing in the prior art, the present invention provides a data hierarchical aggregation and query method and system for privacy protection in a 6G network. The technical problem to be solved by the present invention is achieved by the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides a privacy-preserving data hierarchical aggregation and query method in a 6G network, the method comprising:
[0008] System initialization phase: The central aggregation node sorts all edge nodes and groups them into pairs;
[0009] Data upload and edge aggregation stage: The data uploader encodes its local data and splits the encoded data into two data fragments, and uploads the two data fragments to a group of edge nodes after grouping;
[0010] Edge data preprocessing stage: The edge nodes holding the data slices of the same data uploader perform aggregation calculations to obtain the corresponding frequency addition slices containing frequency information, and convert the frequency addition slices into frequency multiplication slices, and upload the frequency addition slices to the central aggregation node;
[0011] Edge classification query request stage: the service requester uploads the data to be queried to a group of edge nodes after grouping, each edge node calculates the encrypted value of the corresponding classification result by using the frequency multiplication sharding, and returns the encrypted value to the service requester, and the service requester calculates the classification result according to the encrypted value;
[0012] Central data preprocessing stage: The central aggregation node performs aggregation calculation based on the frequency addition shards uploaded by the edge nodes to obtain the global frequency statistics;
[0013] Central node classification query request stage: the service requester uploads the data to be queried to the central aggregation node, and the central aggregation node calculates the naive Bayes prior value based on the uploaded data to be queried and the global frequency statistics, analyzes the naive Bayes prior value to obtain the classification result, and returns the classification result to the service requester;
[0014] Among them, the system initialization phase, data upload and edge aggregation phase, edge model training phase and central model aggregation training phase only need to be executed once, and the edge classification query request phase and the central node classification query request phase are infinitely responded to and executed according to the needs of the service requester.
[0015] In a second aspect, an embodiment of the present invention provides a data hierarchical aggregation and query system for privacy protection in a 6G network, including a data uploader, an edge training node, a central aggregation node, and a service requester; wherein,
[0016] The data uploader includes: a data preprocessing module for encoding its local data; a sharding processing module for splitting the encoded data into two data shards;
[0017] The edge node includes: a shard data aggregation module, which is used to aggregate and calculate the data shards of the same data uploader to obtain the corresponding frequency addition shards containing frequency information; a shard conversion module, which is used to convert the frequency addition shards into frequency multiplication shards; an edge query module, which is used to calculate the encrypted value of the corresponding classification result using the frequency multiplication shards, and return the encrypted value to the service requester; a shard upload module, which is used to upload the frequency addition shards to the central aggregation node;
[0018] The central aggregation node includes: a central grouping module, which is used to sort all edge nodes and group them in pairs; a frequency statistics module, which is used to aggregate and calculate the global frequency statistics according to the frequency addition slices uploaded by each edge node; a central query module, which is used to calculate the naive Bayesian prior value according to the query data uploaded by the service requester and the global frequency statistics, analyze the naive Bayesian prior value to obtain the classification result, and return the classification result to the service requester;
[0019] The service requester includes: an edge upload module, which is used to upload the data to be queried to a group of edge nodes after grouping; an edge receiving module, which is used to receive the encrypted value returned by each edge node and calculate the classification result according to the encrypted value; a central upload module, which is used to upload the data to be queried to the central aggregation node; and a central receiving module, which is used to receive the classification result returned by the central aggregation node.
[0020] Beneficial effects of the present invention:
[0021] The privacy-protected data hierarchical aggregation and query method in the 6G network proposed by the present invention realizes hierarchical aggregation and query of data by using secure multi-party computing technology, and effectively solves the privacy protection problem in the data hierarchical aggregation and query process of AI services in the 6G network; the method optimizes the data aggregation mechanism, so that the privacy of user data is guaranteed while distributed processing; in addition, by applying multi-party secure computing between edge nodes and central aggregation nodes, advanced data aggregation is realized, which not only improves processing efficiency, but also enhances data security; the method not only reduces computing and communication overhead, but also facilitates industrial deployment, and provides efficient and accurate support for AI data classification and query services in the 6G network environment, further improves user experience and ensures data information security, and further improves the efficiency of the solution without losing model accuracy, which is suitable for use in real scenarios. The method provided by the present invention will not affect the accuracy of the model in the process of privacy protection, and can collect user data in real time to update the network end model. Under the premise of ensuring user privacy security, the present invention can build a full-process efficient privacy protection framework that is compatible with the Naive Bayes algorithm, ensuring that the privacy of user data is strictly protected while providing accurate AI classification services; by optimizing the secure multi-party computing protocol and data processing flow, the present invention greatly reduces the amount of data calculation and transmission, reduces communication overhead, and improves the overall performance of the system; in addition, the algorithm-optimized secure multi-party computing technology replaces traditional homomorphic encryption and differential privacy technology, avoids the problems of high computational complexity and low efficiency, significantly improves query efficiency and accuracy, enhances data availability and classification accuracy, thereby further improving user experience and ensuring data information security.
[0022] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flowchart of a privacy-protected data hierarchical aggregation and query method in a 6G network provided by an embodiment of the present invention;
[0024] Figure 2 It is a flowchart of a privacy-protected data hierarchical aggregation and query method in a 6G network provided by an embodiment of the present invention in a feasible specific embodiment;
[0025] Figure 3 It is a schematic diagram of the edge node data uploading process provided by an embodiment of the present invention;
[0026] Figure 4 It is a structural diagram of a privacy-protected data hierarchical aggregation and query method and system in a 6G network provided by an embodiment of the present invention;
[0027] Figure 5 It is a schematic diagram of the working principle of the privacy-protected data hierarchical aggregation and query system in the 6G network provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0029] See also Figure 1 and Figure 2 The embodiment of the present invention provides a privacy-preserving data hierarchical aggregation and query method in a 6G network, the method comprising:
[0030] S10, system initialization phase: the central aggregation node sorts all edge nodes and groups them into pairs;
[0031] S20, data upload and edge aggregation stage: the data uploader encodes its local data and splits the encoded data into two data fragments, and uploads the two data fragments to a group of edge nodes after grouping;
[0032] S30, edge data preprocessing stage: the edge nodes holding the data shards of the same data uploader perform aggregation calculations to obtain the corresponding frequency addition shards containing frequency information, and convert the frequency addition shards into frequency multiplication shards, and upload the frequency addition shards to the central aggregation node;
[0033] S40, edge classification query request stage: the service requester uploads the data to be queried to a group of edge nodes after grouping, each edge node calculates the encrypted value of the corresponding classification result by frequency multiplication sharding, and returns the encrypted value to the service requester, and the service requester calculates the classification result according to the encrypted value;
[0034] S50, central data preprocessing stage: the central aggregation node performs aggregation calculation according to the frequency addition slices uploaded by the edge nodes to obtain the global frequency statistics;
[0035] S60, central node classification query request stage: the service requester uploads the data to be queried to the central aggregation node, and the central aggregation node calculates the naive Bayes prior value according to the uploaded data to be queried and the global frequency statistics, analyzes the naive Bayes prior value to obtain the classification result, and returns the classification result to the service requester;
[0036] Among them, the system initialization phase, data upload and edge aggregation phase, edge model training phase and central model aggregation training phase only need to be executed once, and the edge classification query request phase and the central node classification query request phase are infinitely responded to and executed according to the needs of the service requester.
[0037] Next, each stage is introduced in detail.
[0038] The system initialization stage of the embodiment of the present invention specifically includes:
[0039] The central aggregation node sorts all edge nodes and groups them in pairs according to the sorting results; each edge node independently selects the security parameters that it needs to keep confidential; based on the grouping results, the edge nodes in the group jointly select the security parameters required for data processing, and the edge nodes in the group jointly negotiate the security parameters required for uploading data to the central aggregation node. More specifically:
[0040] 1) The central aggregation node sorts all edge nodes. The specific sorting method is not limited. The communication time between adjacent edge nodes is shortened as much as possible. The central aggregation node groups the edge nodes in pairs according to the sorting results and numbers them in sequence. The subscript represents the serial number of the group, and the superscript represents the serial number of the edge node in the group. represents the edge node with sequence number 1 in the jth group, It represents the edge node with sequence number 2 in the jth group. Each group of edge nodes has its corresponding data uploader.
[0041] 2) Each edge node independently selects the security parameters that need to be kept confidential during the independent data processing phase, and ensures that these parameters are only known to each node and are not leaked to the outside. The edge node with the serial number 1 in the group randomly generates a random number R, large random numbers P1 and P2, and a perturbation μ1; the edge node with the serial number 2 in the group randomly generates a random number R', large random numbers p1 and p2, and a perturbation μ2, where μ1 and μ2 are two extremely small random numbers.
[0042] 3) The edge nodes in the group negotiate with each other to generate a scaling factor Δ, where Δ is a large random number that is an integer power of 10.
[0043] 4) All edge nodes negotiate in pairs in order ( and and and and and ), the symmetric key is negotiated using the DH (Diffie-Hellman) key exchange protocol. The key negotiated between the i∈{1,2}th edge node in the jth group and the previous edge node is recorded as The key negotiated with the next edge node is recorded as The key is stored securely locally to ensure that it will not be leaked. and For example, the exchange of DH keys includes the following steps:
[0044] a) Select common parameters, and First, you need to choose two common parameters: a large prime number p and a primitive root g (g is the multiplication primitive root modulo p, that is, the power of g can generate all integers from 1 to p-1).
[0045] b) Generate a private key, Choose a large random number a as the private key and make sure a is less than p; Choose a large random number b as the private key, and make sure b is less than p.
[0046] c) Calculate the public key, Calculate the public key A = g a modp; Calculate the public key B = g b mod p.
[0047] d) Exchange public keys, and Exchange public keys A and B.
[0048] e) Calculate the symmetric key, receive After obtaining the public key B, calculate the symmetric key receive After obtaining the public key A, calculate the shared key Easy to know
[0049] In the system initialization phase of the embodiment of the present invention, the central aggregation node sorts and groups all edge nodes so that subsequent collaborative processing can be carried out in an orderly manner. Then, each edge node autonomously selects the security parameters that need to be kept confidential during the data processing phase to ensure that these parameters are only known to each node and are not leaked to the outside. Then, the two edge nodes in the group jointly select the security parameters that need to be known to both parties during the data processing process. Then, based on the sorting results, each edge node conducts two-by-two security negotiations in order to establish a set of security parameters that are only known to the two participating parties and are required for uploading data to the central aggregation node.
[0050] Furthermore, the data uploading and edge aggregation stage in the embodiment of the present invention specifically includes:
[0051] The data uploader encodes the label value of its local data by 01 to obtain the label code; the data uploader maps the eigenvalue of its local data onto a matrix, encodes the matrix to obtain the matrix code, and constructs the feature code based on the matrix code; the data uploader performs bit-by-bit multiplication operation on each bit of the label code and feature code except the diagonal code bits to obtain the combined code; the data uploader splits the label code, feature code and combined code except the diagonal code bits to obtain two addition slices of the label code, two addition slices of the feature code, and two addition slices of the combined code respectively; the data uploader uploads the two addition slices of the label code, the two addition slices of the feature code, and the two addition slices of the combined code to a group of edge nodes after the corresponding grouping. More specifically:
[0052] 1) The i-th data uploader Uploader i The label value of the kth data in its local data is encoded by 01 to obtain the label code Then, the eigenvalue of the kth data is mapped to a d×d matrix, and OneHot encoding is performed on each row of the matrix; each diagonal of the matrix is directly encoded, and a unique code is generated for each eigenvalue by combining these two encoding methods. Finally, the matrix codes corresponding to all eigenvalues in the kth data are summarized to construct the feature code of the data.
[0053]
[0054] Where n is the feature dimension, Uploader i The diagonal code values corresponding to the 1st, 2nd, ..nth eigenvalues of the kth data, Uploader i The OneHot encoding value corresponding to the 1st, 2nd, ..nth eigenvalue of the kth data. d represents the encoding length of a single feature, which is obtained by summing up the square root of the dimension values of all feature vector attributes and then rounding up. Diagonal coding All the encoding bits outside are 0 or 1. Finally, the label is encoded With feature encoding Perform bit-by-bit multiplication on each bit except the diagonal coded bit to obtain the combined code of the data
[0055] in
[0056] 2) Data uploader Uploader i Encode the label Feature Encoding Combination Encoding Perform bit splitting outside the diagonal encoding bit to obtain two additive fragments of the label encoding. and Two additive slices of feature encoding
[0057] and
[0058] and two additive slices of the combined code
[0059] and Among them, the shards with subscript 1 are collectively referred to as The shards with subscript 2 are collectively called The addition shards after splitting satisfy
[0060] 3) Data uploader Uploader i The label encoding, feature encoding and combined encoding of the local data are uploaded to two different edge nodes of the corresponding group j. Uploaded to the edge node superior, Uploaded to the edge node so that these edge nodes can store, aggregate and provide secure query services.
[0061] In the data upload and edge aggregation stage of the embodiment of the present invention, data upload is integrated with edge functions. The data uploader encodes the local data and splits the encoded data. The split data fragments are then uploaded to different edge nodes so that these edge nodes can use these data fragments as preprocessed data sets.
[0062] Furthermore, the edge data preprocessing stage of the embodiment of the present invention specifically includes:
[0063] Each edge node counts the total number of data according to the label-coded addition slices uploaded by the data uploader, sums the label-coded addition slices uploaded by the data uploader to obtain the label frequency addition slices corresponding to the label value of 1, and calculates the label frequency addition slices corresponding to the label value of 0 according to the total number of data counted and the label frequency addition slices corresponding to the label value of 1; each edge node performs bitwise summation of the feature values with the same diagonal code in the feature-coded addition slices uploaded by the data uploader, counts the summation results of all feature values, and summarizes the statistical results to obtain the feature frequency addition slices. Each edge node performs bitwise summation of the combination values with the same diagonal code in the addition slices of the combination code uploaded by the data uploader, performs statistics on the summation results of all combination values, summarizes the statistical results to obtain the combination frequency addition slice corresponding to the label value 1, and calculates the combination frequency addition slice corresponding to the label value 0 based on the feature frequency addition slice and the combination frequency addition slice corresponding to the label value 1; each edge node converts the label frequency addition slice and the combination frequency addition slice corresponding to the respective label values 0 and 1 into the corresponding label frequency multiplication slice and the combination frequency multiplication slice. More specifically, taking the edge nodes in the jth group as and In order to include:
[0064] 1) and First, the total number of data N received by each uploader is counted according to the label-coded addition slices uploaded by the data uploader, and then the label-coded addition slices of each uploader are summed to obtain all The label frequency addition shard corresponding to the label value 1 and all The corresponding label value is 1, which is the aggregation addition shard Where j stands for Uploader i The group number corresponding to the edge node group to which it belongs. n represents features 1 to n, Y∈{0,1} represents the label value, and The label frequency addition shard for label value Y=1.
[0065] at the same time, Calculate N and The difference between the label value Y = 0 is obtained by adding the label frequency to the slice Same reason Calculate N and The difference between the two gets the label value Y = 0, and the label frequency addition fragmentation
[0066] 2) and Process the addition slices of their respective feature codes respectively. The specific operation is: and The feature codes in the addition slices are taken out, and the eigenvalue codes with the same diagonal code are summed bit by bit. After the summation is completed, the summation results of each eigenvalue code are counted, and these statistical results are summarized into a one-dimensional vector to express the aggregated results of all eigenvalues. Through this process, and Got all Corresponding feature frequency addition sharding and all Corresponding feature frequency addition sharding Where n is the feature dimension and m is the dimension of a single attribute of the feature vector. and According to the combined values with the same diagonal code in the addition slices of the combined code uploaded by the data uploader, the sum of all the combined values is counted, and the statistical results are summarized to obtain all Combined frequency addition sharding corresponding to label value 1 Combined frequency addition sharding corresponding to all label values 1 With X1,X2,...,X n represents features 1 to n, Y∈{0,1} represents the label value, and we can get and for X1,X2,...,X n The characteristic frequency addition slice of and for X1,X2,...,X n |Y=1 combination frequency addition slice; then calculate bit by bit and The difference is X1, X2, …, X n |Y=0 combination frequency addition sharding and bitwise calculations and The difference X1,X2,...,X n |Y=0 combination frequency addition sharding
[0067] 3) Within the group and Transforming protocol π through sharding STP Add the frequencies of each slice and Transform and replace with the corresponding frequency multiplication slice
[0068]
[0069] and
[0070]
[0071] To ensure the smooth progress of subsequent processes.
[0072] Among them, Sharding Transformation Protocol (π STP ) through the inner edge nodes of the group and The communication and collaborative computing between them can convert the addition data slices provided by the data uploader into multiplication data slices that are more suitable for subsequent processing; specifically:
[0073] 1) Convert the entire frequency addition slice into a frequency multiplication slice, that is, convert each value in the frequency addition slice vector. Each converted value must be stored separately. and On and corresponding to each other, that is and and and etc. and For example, the corresponding value after conversion is and Computational Generation b 12 =Δ·R; Calculate and generate b 21 =Δ·R′, Among them, R, R′, Δ, P1, P2, μ1, and μ2 are the security parameters generated in the first stage, which are
[0074] 2) Calculate and generate b1'1=b 11 +P1,b1'2=b 12 +P2, then send b1'1 and b1'2 Give Among them, P1 and P2 are the security parameters generated in the first stage.
[0075] 3) Calculate and generate b2'1=b 21 b1'1+p1, b2'2=b 22 b1'2+p2, then send b2'1 and b2'2 to Among them, p1 and p2 are the security parameters generated in the first stage.
[0076] 4) Calculate and generate b1"1 = b2'1modP1, b1"2 = b2'2modP2; then send the value of b1"1 + b1"2 to
[0077] 5) Calculate the value of b1”1+b1”2-p1-p2 and divide it by Δ 2 and R′, the result is
[0078] 6) and By converting all the values in the slice vector in turn, the entire addition slice vector can be converted into a multiplication slice vector.
[0079] In the edge data preprocessing stage of the embodiment of the present invention, edge nodes holding the same user data shards will collaborate to perform encrypted communication to ensure that while protecting the privacy of user data, key frequency addition shards required for training classifier models such as the Naive Bayes classifier model are jointly counted. Each edge node processes the data shards it holds respectively, and performs preliminary training of the classifier model such as the Naive Bayes classifier model based on them.
[0080] Furthermore, the edge classification query request stage of the embodiment of the present invention specifically includes:
[0081] The service requester performs OneHot encoding on the encrypted data to be queried provided by the service requester, and encrypts the encoding result and uploads it to a group of edge nodes after grouping through the edge upload protocol; wherein, each edge node calculates the first operation results corresponding to the label values of 0 and 1 respectively according to the encrypted value of the encoding result and the combination frequency multiplication slices corresponding to the label values of 0 and 1 during the uploading process of the edge upload protocol; each edge node calculates the first naive Bayesian prior comparison value of the label value of 1 using the first operation result corresponding to the label value of 1 and the label frequency multiplication slice corresponding to the label value of 0, and calculates the first naive Bayesian prior comparison value of the label value of 0 using the operation result corresponding to the label value of 0 and the label frequency multiplication slice corresponding to the label value of 1; the two edge nodes compare the relative size of the sum of the first naive Bayesian prior comparison values of all label values of 1 and the sum of the first naive Bayesian prior comparison values of all label values of 0 through the security comparison protocol, and obtains the pre-classification result corresponding to each edge node through the comparison result; each edge node returns its pre-classification result to the service requester; the service requester performs an XOR operation on the two pre-classification results, and determines the classification result according to the XOR operation result. More specifically:
[0082] 1) Service requester CLIENT i The feature vector D of the encrypted data to be queried is uploaded through the edge upload protocol i After encryption, it is securely transmitted to the jth group of edge nodes to which it belongs. and During this process, the protocol will perform a series of computational operations. After completing these operations, and Two different operation results will be calculated respectively. Among them, The eigenvector D i The encrypted value and the label value corresponding to the combination frequency multiplication shard Multiplication sharding of the combined frequency corresponding to the label value 0 The calculation results are recorded as the first calculation result corresponding to the label value 1 The first operation result corresponding to the label value 0 and The eigenvector D i The encrypted value and the label value corresponding to the combination frequency multiplication shard Multiplication sharding of the combined frequency corresponding to the label value 0 The calculation results are recorded as the first calculation results corresponding to the label value 1. The first operation result corresponding to the label value 0
[0083] 2) Use the first operation result corresponding to the label value of 1 Multiplication sharding of the label frequency corresponding to the label value 0 Calculate the first naive Bayesian prior comparison value of label value Y=1 Use the first operation result corresponding to the label value of 0 Multiplication sharding of the label frequency corresponding to the label value 1 Calculate the first naive Bayesian prior comparison value of label value Y=0 Similarly, Use the first operation result corresponding to the label value of 1 Multiplication sharding of the label frequency corresponding to the label value 0 Calculate the first naive Bayesian prior comparison value of label value Y=1 Use the first operation result corresponding to the label value of 0 Multiplication sharding of the label frequency corresponding to the label value 1 Calculate the first naive Bayesian prior comparison value of label value Y=0
[0084] 3) and Through the secure comparison protocol, the comparison data can be safely determined without exposing them. and During the entire protocol execution process, the two edge nodes cannot know the specific results of the comparison. After the protocol execution is completed, A pre-classification result will be obtained, denoted as k 1 ∈{0,1}, and A pre-classification result will also be obtained, denoted as k 2 ∈{0,1}.
[0085] 4) and The pre-classification results k are 1 and the pre-classification result k 2 Return to CLIENT i , CLIENT i K 1 and k 2 Perform XOR operation, if (k 1 )XOR(k 2 )=1, then the label Y of the data to be queried is considered to be 0; if (k 1 )XOR(k 2 )=0, it is considered that the label Y of the data to be queried is 1.
[0086] Among them, Edge Upload Protocol (π EUP ) Through communication and collaborative computing between the service requester and the edge node, the local data to be queried by the service requester is securely uploaded to the edge node in encrypted form; specifically:
[0087] 1) Service requester CLIENT i Encode the local data to be queried using OneHot encoding Where i is the number of the service requester, n is the feature dimension, m is the dimension of a single attribute of the feature vector, CLIENT i The two edge nodes are grouped as j. The protocol requires the service requester to perform vector-by-vector and feature-by-feature operations on the two corresponding edge nodes. and For example, it is easy to know and Corresponding to D i and The same feature in the calculation results are CLIENT i Randomly select a random number v, record is x1, is x2.
[0088] 2) CLIENT i Choose a random number seed and send to Choose a random number seed And send to CLIENT i , and then the two calculate the public random number seed seed = seed1 seed2 and generate a The random number matrix
[0089] 3) Generate a A random number vector And calculate x′1=r·M T +x1, where x′1=(x′ 1,1 ,x′ 1,2 ,…,x′ 1,m ),Then Send x′1 to CLIENT i .
[0090] 4) After receiving x′1, CLIENT i calculate And calculate x'2 = x2·M, where Then CLIENT i Send s1 and x'2 to
[0091] 5) After receiving s1 and x'2, First calculate Then we can calculate
[0092] 6) Similarly, CLIENT i and Continue to calculate the subsequent features, and we can get D i Combined frequency multiplication shards corresponding to the label value 1 Multiplication sharding of the combined frequency corresponding to the label value 0 The calculation results are the first calculation results corresponding to the label value 1 The first operation result corresponding to the label value 0 Then CLIENT i Let x2 be Performing the same operation, we can get D i Combined frequency multiplication shards corresponding to the label value 1 Multiplication sharding of the combined frequency corresponding to the label value 0 The calculation results are the first calculation results corresponding to the label value 1 The first operation result corresponding to the label value 0
[0093] Among them, the Security Comparison Protocol (π SCP )pass and The communication and collaborative calculation between them can calculate the comparison value and return it to the CLIENT without knowing the comparison value and comparison result. i , make CLIENT i Get the comparison result. Holds a priori comparison value and Holds a priori comparison value and Specifically:
[0094] 1) Calculate Calculate
[0095] 2) Randomly select two random numbers r1' and r2', and ensure that r1'>r2', then Then randomly select a number c∈{-1,1}, if c=1, let k 2 =1, if c = -1, then let k 2 = 0, then K 2 Return to CLIENT i , and calculate d (2) =c·r′1; Take d (1) =0.
[0096] 3) Generate a random number r', record x2 = (r', b (1) ,d (1) ,0), stored in Let x1=(1,b (2) ,d (2) ,0), stored in
[0097] 4) Choose a random number seed and send to Choose a random number seed and send to Then the two calculate the common random number seed seed = seed1·seed2 and generate a 4×2 random number matrix through seed
[0098] 5) Generate a 1×2 vector of random numbers And calculate x′1=r·M T +x1, where x′1=(x′ 1,1 ,x′ 1,2 ,x′ 1,3 ,x′ 1,4 ),Then Send x′1 to
[0099] 6) After receiving x1', calculate And calculate x'2 = x2·M, where x'2 = (x' 2,1 ,x' 2,2 ),Then Send s1 and x'2 to
[0100] 7) After receiving s1 and x'2, First calculate s2 = x' 2,1 ·r1+x' 2,2 ·r2, and then calculate s=s1-s2.
[0101] 8) Calculate z (1) =b (1) ·d (1) -r′, Calculate z (2) =s+b (2) ·d (2) ,Then z (2) The value of +r′2 is sent to
[0102] 9) Upon receiving z (2) After calculating the value of z = z (1) +z (2) +r′2, if z≤0, then let k 1 =0; if z>0, let k 1 =1, then K 1 Return to CLIENT i .
[0103] In the edge classification query request stage of the embodiment of the present invention, the service requester performs privacy encryption processing on the query data and uploads it to the edge node. The edge node uses the query data provided by the service requester and combines its own preliminarily trained classifier model, such as the naive Bayes classifier model, to perform calculations on the query data. On the premise of ensuring that the user data privacy and classification results are not leaked, the edge node calculates the encrypted value of the data classification and returns the value to the service requester, who can calculate the classification result based on this encrypted value.
[0104] Furthermore, the data preprocessing stage of the center in the embodiment of the present invention specifically includes:
[0105] Each edge node encrypts the label frequency addition slices corresponding to the respective label values of 0 and 1, and the combined frequency addition slices corresponding to the label values of 0 and 1, and uploads the encrypted label frequency addition slices corresponding to the label values of 0 and 1, and the combined frequency addition slices corresponding to the label values of 0 and 1 to the central aggregation node through the aggregation upload protocol; the central aggregation node sums up all the label frequency addition slices corresponding to the label value of 1 to obtain the aggregated label frequency addition slice with the label value of 1, sums up all the label frequency addition slices corresponding to the label value of 0 to obtain the aggregated label frequency addition slice with the label value of 0, sums up all the combined frequency addition slices corresponding to the label value of 1 to obtain the aggregated combined frequency addition slice with the label value of 1, and sums up all the combined frequency addition slices corresponding to the label value of 0 to obtain the aggregated combined frequency addition slice with the label value of 0. More specifically:
[0106] 1) Each edge node uploads protocol π through aggregation AUP Perform collaborative calculations with the previous and next edge nodes in sequence ( and and and and and ). In this process, each edge node, such as and Add the label frequencies corresponding to all label values 1 and shard them Addition sharding of label frequencies corresponding to label values of 0 Additive sharding of the combination frequency corresponding to the label value 1 Additive sharding of the combined frequency corresponding to the label value 0 Or the label frequency addition sharding corresponding to the label value 1 Addition sharding of label frequencies corresponding to label values of 0 Additive sharding of the combination frequency corresponding to the label value 1 Additive sharding of the combined frequency corresponding to the label value 0 The addition shard is encrypted to generate the label frequency addition shard corresponding to the encrypted label value of 1 Addition sharding of label frequencies corresponding to label values of 0 Additive sharding of the combination frequency corresponding to the label value 1 Additive sharding of the combined frequency corresponding to the label value 0 Or the label frequency addition sharding corresponding to the label value 1 Addition sharding of label frequencies corresponding to label values of 0 Additive sharding of the combination frequency corresponding to the label value 1
[0107] Additive sharding of the combination frequency corresponding to the label value 1
[0108] After encryption is completed, these encrypted frequency addition shards are uploaded to the central aggregation node to ensure the secure transmission of data and the smooth progress of subsequent calculation processes.
[0109] 2) After receiving the encrypted frequency addition shards uploaded by each edge computing node, the central aggregation node adds the label frequency addition shards corresponding to all encrypted label values 1 Add the label frequency corresponding to the encrypted label value 1 to the shard Sum the values to get the aggregated label frequency addition slice c with label value Y=1 1 , and then add the label frequency corresponding to all encrypted label values 0 to the fragment Add the label frequency corresponding to the encrypted label value 0 to the fragment Sum the aggregated label frequency addition slice c with label value Y=0 0 ; Then add the combined frequencies corresponding to all label values of 1 and shard them The combined frequency addition sharding corresponding to the label value 1 The elements at each position are added to obtain X1, X2, ..., X n | Aggregate combination frequency addition sharding with Y=1
[0110] Right now Then add the combined frequencies corresponding to all label values of 0 to split Combined frequency addition sharding corresponding to label value 0 The elements at each position are added to obtain X1, X2, ..., X n | Aggregate combination frequency addition sharding with Y=0 Right now
[0111]
[0112] Among them, Aggregate Upload Protocol (π AUP ) Through communication and collaborative computing between edge computing nodes, the data of each edge node is securely uploaded to the central aggregation node in encrypted form. Specifically:
[0113] 1) If Figure 3 As shown, the i-th edge node in the j-th group When uploading your own shard data to the central aggregation node, each value in the shard data will be subtracted Add Upload again, and is the security parameter generated in the first stage. This addition and subtraction process is the encryption process. or The value after frequency addition shard encryption is recorded as and The shard data includes the label frequency addition shards corresponding to the label values 0 and 1, and the combination frequency addition shards corresponding to the label values 0 and 1.
[0114] 2) If during the polymerization If a failure occurs and data is not uploaded, the final aggregation result needs to be subtracted The actual value of the corresponding shard data, and then add Subtract
[0115] In the central data preprocessing stage of the embodiment of the present invention, each edge node performs privacy encryption processing on its own frequency addition shard data and then uploads its own statistical data to the central aggregation node. After receiving these statistical data, the central aggregation node performs aggregation analysis to obtain global frequency statistics, and then a global classifier model such as a global naive Bayes classification model can be constructed based on the global frequency statistics.
[0116] Furthermore, the central node classification query request stage in the embodiment of the present invention specifically includes:
[0117] The service requester performs OneHot encoding on the encrypted data to be queried, and encrypts the encoding result and uploads it to the central aggregation node through the central upload protocol; wherein, during the upload process of the central upload protocol, the second operation results corresponding to the label values of 0 and 1 are calculated according to the encrypted value of the encoding result and the aggregated combination frequency addition slice corresponding to the label values of 0 and 1 respectively; the central aggregation node calculates the second naive Bayesian prior comparison value of the label value of 1 using the second operation result corresponding to the label value of 1 and the aggregated label frequency addition slice corresponding to the label value of 0, and calculates the second naive Bayesian prior comparison value of the label value of 0 using the second operation result corresponding to the label value of 0 and the aggregated label frequency addition slice corresponding to the label value of 1; compares the second naive Bayesian prior comparison value of the label value of 1 with the second naive Bayesian prior comparison value of the label value of 0, obtains the classification result of the central aggregation node through the comparison result, and returns the classification result to the service requester. More specifically:
[0118] 1) Service requester CLIENT i Encode the local data to be queried using OneHot encoding Where i is the ID of the user to be queried, n is the feature dimension, and m is the dimension of a single attribute of the feature vector. Then CLIENT i Through the central transmission protocol, the local data to be queried D i After encryption, it is securely transmitted to the central aggregation node. During this process, the protocol will perform a series of calculation operations. After completing these operations, the central aggregation node will calculate two different calculation result vectors, which are recorded as the second calculation result corresponding to the label value of 1. The second operation result corresponding to the label value 0
[0119] 2) The central aggregation node uses the second operation result G corresponding to the label value 1 i The aggregate combination frequency addition shard c corresponding to the label value 0 0 Calculate the second naive Bayesian prior comparison value of the label value Y = 1 Use the second operation result H corresponding to the label value of 0 i The aggregate combination frequency addition shard c corresponding to the label value 1 1 Calculate the second naive Bayesian prior comparison value of the label value Y = 0
[0120] 3) The second naive Bayesian prior comparison value of the central aggregation node comparison label value Y=1 Compare the value with the second naive Bayes prior of label value Y=0 The size of Then the label Y of the data to be queried is considered to be 0. The label Y of the query data is considered to be 1, and the classification result is returned to the CLIENT i .
[0121] Among them, the Center Upload Protocol (π CUP ) Through communication and collaborative computing between the service requester and the central aggregation node, the service requester's data to be queried is securely uploaded to the central aggregation node. Specifically:
[0122] 1) Service requester CLIENT i The local feature vector to be queried is encoded using OneHot encoding as Where i is the number of the query user, n is the feature dimension, m is the dimension of a single attribute of the feature vector, CLIENT i The two edge computing nodes are grouped together with the number j. The protocol requires CLIENT i Operate vector-by-vector and feature-by-feature with the central aggregation node. and For example, it is easy to know and Corresponding to D i and B 1 The same feature in the calculation results are CLIENT i Randomly select a random number v, record is x1, is x2.
[0123] 2) The central aggregation node selects a random number seed and send to Choose a random number seed And send to CLIENT i , and then the two calculate the public random number seed seed = seed1·seed2 and generate a The random number matrix
[0124] 3) The central aggregation node generates a A random number vector And calculate x′1=r·M T +x1, where x′1=(x′ 1,1 ,x′ 1,2 ,…,x′ 1,m ), then the central aggregation node sends x′1 to the CLIENT i .
[0125] 4) After receiving x′1, CLIENTi calculate And calculate x'2 = x2·M, where Then CLIENT i Send s1 and x'2 to the central aggregation node.
[0126] 5) After receiving s1 and x'2, the central aggregation node first calculates Then we can calculate
[0127] 6) Similarly, CLIENT i Continue to calculate the subsequent features with the central aggregation node, and get D i The aggregated combination frequency addition shard B corresponding to the label value 1 1 , the aggregate combination frequency addition shard B corresponding to the label value of 0 0 The result of the operation is recorded as the second operation result corresponding to the label value of 1 The second operation result corresponding to the label value 0
[0128] In the central node classification query request stage of the embodiment of the present invention, the service requester performs privacy encryption processing on the query data and uploads it to the central aggregation node. The central aggregation node performs calculations based on the uploaded query data and global frequency statistics based on the global naive Bayes classifier model, completes the classification task of the query data under the premise of privacy protection, and returns the classification results to the service requester.
[0129] In order to solve the problem of data hierarchical aggregation and secure query with privacy protection in the existing 6G network, and to improve the overall performance and security of the system, the present invention makes the following improvements to the existing technology:
[0130] 1) Data hierarchical aggregation framework: This invention designs a new data aggregation mechanism. The data uploader splits the local data and uploads it to the grouped edge nodes. The secure two-party computing technology is used to aggregate the data to achieve distributed processing and protect privacy. The edge nodes can also further integrate the data into the central aggregation node through multi-party secure computing to complete global data aggregation, thereby improving processing efficiency and security.
[0131] 2) Efficient privacy protection method: This invention uses secure multi-party computing technology to build a comprehensive privacy protection framework, covering all aspects of data uploading, aggregation and query, providing full-process privacy protection for user data, and ensuring the privacy security of user data during transmission and calculation. At the same time, the framework is compatible with the Naive Bayes algorithm, thus providing users with accurate AI classification services on the basis of strictly protecting data privacy.
[0132] 3) Optimize computing and communication overhead: This invention uses algorithm-optimized secure multi-party computing technology and data processing flow to simplify the entire data upload and computing process into linear processing, thereby significantly reducing the amount of data calculation and transmission. During data aggregation and query, only necessary encrypted data is transmitted, further reducing communication overhead.
[0133] 4) Improve query efficiency and accuracy: The algorithm-optimized secure multi-party computing technology is used to replace the homomorphic encryption and differential privacy technologies in traditional methods, avoiding the high computational complexity and low efficiency of homomorphic encryption, while solving the problem of reduced data availability and classification accuracy that may be caused by differential privacy. This method not only improves computational efficiency, but also improves data availability and classification accuracy.
[0134] In a second aspect, an embodiment of the present invention provides a data hierarchical aggregation and query system for privacy protection in a 6G network, including a data uploader, an edge training node, a central aggregation node, and a service requester; wherein,
[0135] Data Uploader i :Responsible for collecting local data from data uploaders, preprocessing the collected local data, splitting the preprocessed local data, and encrypting the split data fragments. Then, upload the encrypted data fragments to the edge nodes of their corresponding groups.
[0136] Edge nodes: such as edge nodes or Responsible for aggregating the data fragments uploaded by each data uploader to generate edge node training data. According to the security query protocol, the edge nodes in the same group communicate collaboratively to calculate the model parameters and provide the service requester with a full-process privacy protection classification service. At the same time, the edge node encrypts the necessary training data according to the aggregation upload protocol and sends the encrypted data to the central aggregation node.
[0137] Central aggregation node: responsible for sorting and grouping all edge nodes. Receives encrypted training data from each edge node and performs aggregation calculations to generate global training data. Then, the central node calculates the global classifier model parameters and provides classification services to service requesters in real time according to the central upload protocol without obtaining the user's real data information.
[0138] Service requester CLIENT i:A user who initiates a classification service request to an edge node or a central aggregation node. According to the edge upload protocol or the central upload protocol, the service requester securely uploads the data to be queried to the two edge nodes or the central aggregation node of the group to which it belongs without revealing the specific true value of the data. Then, the service requester can calculate the classification result of the data to be queried based on the return value of the edge node, or directly obtain the classification result based on the return value of the central aggregation node.
[0139] Furthermore, if Figure 4 As shown, the data uploader of the embodiment of the present invention includes: a data preprocessing module, which is used to encode its local data; a fragmentation processing module, which is used to split the encoded data into two data fragments;
[0140] The edge node includes: a shard data aggregation module, which is used to aggregate and calculate the data shards of the same data uploader to obtain the corresponding frequency addition shards containing frequency information; a shard conversion module, which is used to convert the frequency addition shards into frequency multiplication shards; an edge query module, which is used to calculate the encrypted value of the corresponding classification result using the frequency multiplication shards, and return the encrypted value to the service requester; a shard upload module, which is used to upload the frequency addition shards to the central aggregation node;
[0141] The central aggregation node includes: a central grouping module, which is used to sort all edge nodes and group them in pairs; a frequency statistics module, which is used to aggregate and calculate the global frequency statistics based on the frequency addition slices uploaded by each edge node; a central query module, which is used to calculate the naive Bayesian prior value based on the query data uploaded by the service requester and the global frequency statistics, analyze the naive Bayesian prior value to obtain the classification result, and return the classification result to the service requester;
[0142] The service requester includes: an edge upload module, which is used to upload the data to be queried to a group of edge nodes after grouping; an edge receiving module, which is used to receive the encrypted value returned by each edge node and calculate the classification result based on the encrypted value; a central upload module, which is used to upload the data to be queried to the central aggregation node; and a central receiving module, which is used to receive the classification result returned by the central aggregation node.
[0143] Furthermore, the data uploader in the embodiment of the present invention also includes: a data collection module, which is used to collect local data.
[0144] Furthermore, the edge node of the embodiment of the present invention also includes a security parameter initialization module, which is used to autonomously select the security parameters that each edge node needs to keep confidential; based on the grouping results, the edge nodes in the group jointly select the security parameters required for data processing, and the edge nodes in the group jointly negotiate the security parameters required for uploading data to the central aggregation node.
[0145] Figure 5 The working principle of the privacy-preserving data hierarchical aggregation and query system in 6G networks is illustrated in detail.
[0146] As for the system embodiment of the second aspect, since it is basically similar to the method embodiment of the first aspect, the description is relatively simple, and the relevant parts may refer to the partial description of the method embodiment of the first aspect.
[0147] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0148] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art may understand and implement other variations of the disclosed embodiments by viewing the specification and its drawings. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "one" does not exclude multiple situations. Certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0149] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A privacy-preserving data hierarchical aggregation and query method in a 6G network, characterized in that: The method comprises: System initialization phase: The central aggregation node sorts all edge nodes and groups them into pairs; Data upload and edge aggregation stage: The data uploader encodes its local data and splits the encoded data into two data fragments, and uploads the two data fragments to a group of edge nodes after grouping; Edge data preprocessing stage: The edge nodes holding the data slices of the same data uploader perform aggregation calculations to obtain the corresponding frequency addition slices containing frequency information, and convert the frequency addition slices into frequency multiplication slices, and upload the frequency addition slices to the central aggregation node; Edge classification query request stage: the service requester uploads the data to be queried to a group of edge nodes after grouping, each edge node calculates the encrypted value of the corresponding classification result by using the frequency multiplication sharding, and returns the encrypted value to the service requester, and the service requester calculates the classification result according to the encrypted value; Central data preprocessing stage: The central aggregation node performs aggregation calculation based on the frequency addition shards uploaded by the edge nodes to obtain the global frequency statistics; Central node classification query request stage: the service requester uploads the data to be queried to the central aggregation node, and the central aggregation node calculates the naive Bayes prior value based on the uploaded data to be queried and the global frequency statistics, analyzes the naive Bayes prior value to obtain the classification result, and returns the classification result to the service requester; Among them, the system initialization phase, data upload and edge aggregation phase, edge model training phase and central model aggregation training phase only need to be executed once, and the edge classification query request phase and the central node classification query request phase are infinitely responded to and executed according to the needs of the service requester.
2. The privacy-preserving data hierarchical aggregation and query method in a 6G network according to claim 1, characterized in that: The system initialization phase specifically includes: The central aggregation node sorts all edge nodes and groups them into pairs according to the sorting results; Each edge node independently selects the security parameters that it needs to keep confidential; Based on the grouping results, the edge nodes in the group jointly select the security parameters required for data processing, and the edge nodes in the group jointly negotiate the security parameters required for uploading data to the central aggregation node.
3. The privacy-preserving data hierarchical aggregation and query method in a 6G network according to claim 1, characterized in that: The data upload and edge aggregation phase specifically includes: The data uploader encodes the label value of its local data by 01 to obtain the label code; The data uploader maps the eigenvalues of its local data onto a matrix, encodes the matrix to obtain a matrix code, and constructs a eigencode based on the matrix code; The data uploader performs bit-by-bit multiplication operation on each bit of the label code and the feature code except the diagonal code bit to obtain a combined code; The data uploader splits the label code, the feature code, and the combination code at the bit positions other than the diagonal code positions, and obtains two addition slices of the label code, two addition slices of the feature code, and two addition slices of the combination code respectively; The data uploader uploads two addition slices of label coding, two addition slices of feature coding, and two addition slices of combination coding to a group of edge nodes after corresponding grouping.
4. The privacy-preserving data hierarchical aggregation and query method in a 6G network according to claim 3, characterized in that: The edge data preprocessing stage specifically includes: Each edge node counts the total number of data received by each edge node according to the label-coded addition slices uploaded by the data uploader, sums the label-coded addition slices uploaded by the data uploader to obtain the label frequency addition slice corresponding to the label value of 1, and calculates the label frequency addition slice corresponding to the label value of 0 according to the total number of data counted and the label frequency addition slice corresponding to the label value of 1; Each edge node performs bitwise summation of the eigenvalues with the same diagonal code in the addition slice of the feature code uploaded by the data uploader, performs statistics on the summation results of all eigenvalues, and summarizes the statistical results to obtain the feature frequency addition slice; Each edge node performs bitwise summation of the combination values with the same diagonal code in the addition slices of the combination code uploaded by the data uploader, performs statistics on the summation results of all combination values, summarizes the statistical results to obtain the combination frequency addition slice corresponding to the label value 1, and calculates the combination frequency addition slice corresponding to the label value 0 based on the feature frequency addition slice and the combination frequency addition slice corresponding to the label value 1; Each edge node converts the label frequency addition sharding and combined frequency addition sharding corresponding to the respective label values 0 and 1 into the corresponding label frequency multiplication sharding and combined frequency multiplication sharding.
5. The privacy-preserving data hierarchical aggregation and query method in a 6G network according to claim 4, characterized in that: The edge classification query request stage specifically includes: The service requester performs OneHot encoding on the data to be queried, and encrypts the encoding result and uploads it to a group of edge nodes after grouping through the edge upload protocol; wherein, each edge node calculates the first operation result corresponding to the label value of 0 and 1 respectively according to the encrypted value of the encoding result and the combination frequency multiplication corresponding to the label value of 0 and 1 during the uploading process of the edge upload protocol; Each edge node calculates the first naive Bayesian prior comparison value of the label value of 1 by using the first operation result corresponding to the label value of 1 and the label frequency multiplication slice corresponding to the label value of 0, and calculates the first naive Bayesian prior comparison value of the label value of 0 by using the operation result corresponding to the label value of 0 and the label frequency multiplication slice corresponding to the label value of 1; The two edge nodes compare the relative size of the sum of the first naive Bayesian prior comparison values of all label values 1 and the sum of the first naive Bayesian prior comparison values of all label values 0 through a secure comparison protocol, and obtain the pre-classification result corresponding to each edge node through the comparison result; Each edge node returns its pre-classification result to the service requester; the service requester performs an XOR operation on the two pre-classification results and determines the classification result according to the XOR operation result.
6. The privacy-preserving data hierarchical aggregation and query method in a 6G network according to claim 4, characterized in that: The central data preprocessing stage specifically includes: Each edge node encrypts the label frequency addition slices corresponding to the respective label values 0 and 1, and the combined frequency addition slices corresponding to the label values 0 and 1, and uploads the encrypted label frequency addition slices corresponding to the label values 0 and 1, and the combined frequency addition slices corresponding to the label values 0 and 1 to the central aggregation node through the aggregation upload protocol; The central aggregation node sums up the label frequency addition slices corresponding to all label values 1 to obtain the aggregated label frequency addition slice with label value 1, sums up the label frequency addition slices corresponding to all label values 0 to obtain the aggregated label frequency addition slice with label value 0, sums up the combined frequency addition slices corresponding to all label values 1 to obtain the aggregated combined frequency addition slice with label value 1, and sums up the combined frequency addition slices corresponding to the label value 0 to obtain the aggregated combined frequency addition slice with label value 0.
7. The privacy-preserving data hierarchical aggregation and query method in a 6G network according to claim 6, characterized in that: The central node classification query request stage specifically includes: The service requester performs OneHot encoding on the data to be queried, and encrypts the encoding result and uploads it to the central aggregation node through the central upload protocol; wherein, during the upload process of the central upload protocol, the second operation result corresponding to the label value of 0 and 1 is calculated by adding the aggregate combination frequency corresponding to the encrypted value of the encoding result and the label value of 0 and 1 respectively; The central aggregation node calculates the second naive Bayesian prior comparison value of the label value of 1 by using the second operation result corresponding to the label value of 1 and the aggregated label frequency addition shard corresponding to the label value of 0, and calculates the second naive Bayesian prior comparison value of the label value of 0 by using the second operation result corresponding to the label value of 0 and the aggregated label frequency addition shard corresponding to the label value of 1; The second naive Bayesian prior comparison value with a label value of 1 is compared with the second naive Bayesian prior comparison value with a label value of 0, and the classification result of the central aggregation node is obtained through the comparison result, and the classification result is returned to the service requester.
8. A privacy-preserving data hierarchical aggregation and query system in a 6G network, characterized in that: Including data uploader, edge training node, central aggregation node, service requester; among them, The data uploader includes: a data preprocessing module for encoding its local data; a sharding processing module for splitting the encoded data into two data shards; The edge node includes: a shard data aggregation module, which is used to aggregate and calculate the data shards of the same data uploader to obtain the corresponding frequency addition shards containing frequency information; a shard conversion module, which is used to convert the frequency addition shards into frequency multiplication shards; an edge query module, which is used to calculate the encrypted value of the corresponding classification result using the frequency multiplication shards, and return the encrypted value to the service requester; a shard upload module, which is used to upload the frequency addition shards to the central aggregation node; The central aggregation node includes: a central grouping module, which is used to sort all edge nodes and group them in pairs; a frequency statistics module, which is used to aggregate and calculate the global frequency statistics according to the frequency addition slices uploaded by each edge node; a central query module, which is used to calculate the naive Bayesian prior value according to the query data uploaded by the service requester and the global frequency statistics, analyze the naive Bayesian prior value to obtain the classification result, and return the classification result to the service requester; The service requester includes: an edge upload module, which is used to upload the data to be queried to a group of edge nodes after grouping; an edge receiving module, which is used to receive the encrypted value returned by each edge node and calculate the classification result according to the encrypted value; a central upload module, which is used to upload the data to be queried to the central aggregation node; and a central receiving module, which is used to receive the classification result returned by the central aggregation node.
9. The privacy-preserving data hierarchical aggregation and query system in a 6G network according to claim 8, characterized in that: The data uploader also includes: a data collection module for collecting local data.
10. The privacy-preserving data hierarchical aggregation and query system in a 6G network according to claim 8, characterized in that: The edge node also includes a security parameter initialization module, which is used to autonomously select the security parameters that each edge node needs to keep confidential. Based on the grouping results, the edge nodes in the group jointly select the security parameters required for data processing, and the edge nodes in the group jointly negotiate the security parameters required for uploading data to the central aggregation node.
Citation Information
Patent Citations
Edge-assisted privacy protection multi-dimensional data hierarchical aggregation method
CN116361827A
Multi-party privacy protection naive Bayesian model rapid training and lightweight prediction method
CN116405184A
Multi-layer and multi-level personalized local differential privacy method for classifying and grading data frequency estimation
CN117390683A
Safe acquisition method of naive Bayesian model in 6G network
CN117651272A
Cited By
Cloud edge-end collaborative data security joint computing method and system
CN121396440A
Cloud edge-end collaborative data security joint computing method and system
CN121396440B
Secure search system, secure search apparatus, secure search method, and program
US20250348613A1