A Selective Data Aggregation Method and System Based on Conditional Classification Encoding

Through the selective data aggregation method based on conditional classification encoding, the user data is encrypted and processed and compliant evaluation is used to evaluate user data using ElGamal and Paillier encryption algorithms, which solves the problems of privacy protection and data accuracy in multi-attribute data aggregation, and realizes efficient and flexible data analysis.

CN120030577BActive Publication Date: 2025-07-22ZHEJIANG SCI-TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510520997.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-22
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

Existing data aggregation methods are difficult to efficiently process when processing multi-attribute data while ensuring data accuracy and privacy protection. Especially for diversified attribute features of non-numeric types, they cannot meet the complex scenario requirements of multi-attribute simultaneous aggregation.

Method used

The selective data aggregation method based on conditional classification encoding is adopted, and the encoding mechanism and conditional screening strategy are used to encrypt user data using ElGamal and Paillier encryption algorithms to generate encryption matrix sequences, and selective aggregation is achieved through the compliance evaluation of authoritative institutions and multi-level security mechanisms to ensure that data privacy is not violated.

Benefits of technology

It realizes accurate aggregation of multi-attribute data, improves the flexibility and practicality of data analysis, optimizes computing efficiency, provides reliable privacy guarantees, and meets the diverse needs in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030577B_ABST
    Figure CN120030577B_ABST
Patent Text Reader

Abstract

The present solution provides a selective data aggregation method and system based on conditional classification coding, including the steps of: obtaining a data analysis request, where the data analysis request records the task conditions and task requirements of the current data analysis task; collecting segmented user data of users in the same area based on the data analysis request, and performing encryption processing on the segmented user data to obtain an encrypted matrix sequence, where the segmented user data includes coded user data corresponding to ordinary attributes and private attribute user data of private attributes, selectively aggregating the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data; performing compliance evaluation on the preliminary aggregated data and screening the qualified screened data; and performing secondary aggregation on the screened data to obtain final aggregated data, which can be based on specific task conditions set by data requesters on the premise of ensuring that the privacy of the private attributes and ordinary attributes of data owners is not violated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data aggregation, and particularly to a selective data aggregation method and system based on conditional classification coding. Background Art

[0002] With the rapid development of information technology, data plays a crucial role in all fields of today's society. Whether it is business decision-making, medical research, or social governance, it highly depends on the accuracy and integrity of the data uploaded by devices by data users. In this context, data users are no longer satisfied with simple data uploading and basic processing, but put forward more stringent requirements, expecting to extract valuable information from the vast and complex data to provide strong support for various decisions.

[0003] In the process of data statistical analysis, the choice of data uploading and processing method is crucial. Traditional data uploading methods often directly upload a large amount of raw data to the server. This method not only occupies a large amount of communication resources, resulting in high communication overhead, but also occupies a huge storage cost at the storage end, bringing huge pressure to data storage and management. To solve these problems, the use of data aggregation for uploading and processing has gradually become the mainstream trend. Data aggregation can sort out and merge scattered data in a targeted manner, effectively reducing the data transmission volume during communication, thereby reducing communication overhead; at the same time, at the storage level, the aggregated data also greatly reduces the required storage space and storage cost. Moreover, data aggregation can also remove redundant information in the data, improve the overall efficiency of data processing, and make subsequent data analysis and mining work more efficient and accurate.

[0004] However, in the context of the distributed privacy protection model, data aggregation faces more severe and complex challenges. In a distributed system, user groups in each region often have unique attribute characteristics and private data due to many factors such as their geographical location, social environment, and consumption habits. This requires that when performing data aggregation, the privacy information of users must be strictly protected from leakage on the premise of ensuring data accuracy and integrity.

[0005] Existing methods often find it difficult to accurately extract effective features when processing multi-attribute data, especially for diversified attribute features of non-numerical types. How to process them in a unified and efficient manner is still a key issue that needs to be solved urgently. Chinese invention patent CN202411395708.1 proposes a method that can accurately aggregate data to attributes, aiming to solve the problem of non-aggregate field data loss in the prior art. The solution can distinguish between aggregated attributes and non-aggregate attributes through attribute separation and identification, ensuring the integrity and accuracy of reported data. However, the invention can only identify a single attribute at a time, which has certain limitations in practical applications. In actual data scenarios, in many cases, it is necessary to process the aggregation of multiple attributes at the same time. For example, when comprehensively analyzing user purchasing behavior, it may be necessary to consider multiple attributes such as the user's age, occupation, purchase time, purchase amount, etc. The invention cannot be efficiently processed when facing complex scenarios where multiple attributes are aggregated at the same time. This not only limits its promotion scope in practical applications, but also affects the effect and efficiency of data aggregation to a certain extent.

[0006] Therefore, how to provide a data aggregation method that meets the data demander's balance needs between privacy protection and data availability is a key issue that needs to be solved urgently. Summary of the invention

[0007] The purpose of the present invention is to provide a selective data aggregation method and system based on conditional classification coding, which, through a coding mechanism and a conditional screening strategy, can accurately implement the selective aggregation operation of data according to the specific task conditions set by the data demander while ensuring that the privacy of the private attributes and common attributes of the data owner is not violated.

[0008] To achieve the above objectives, the present technical solution provides a selective data aggregation method based on conditional classification coding, comprising the following steps:

[0009] Obtaining a data analysis request, wherein the data analysis request records the task conditions and task requirements of the current data analysis task;

[0010] Based on the data analysis request, segmented user data of users in the same area are collected, and the segmented user data are encrypted to obtain an encrypted matrix sequence, wherein the segmented user data includes coded user data corresponding to common attributes and private attribute user data corresponding to private attributes.

[0011] Selectively aggregate the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data;

[0012] Conduct compliance assessment on the preliminary aggregated data and screen the eligible screening data;

[0013] The filtered data are aggregated twice to obtain the final aggregated data.

[0014] Compared with the prior art, the technical solution of the present invention has the following features and beneficial effects:

[0015] 1. The present invention proposes an innovative selective data aggregation method based on conditional classification coding, which can achieve data security aggregation for various statistical function requirements, and at the same time support customized statistical analysis and attribute-based precise aggregation. By introducing an innovative privacy protection mechanism and efficient algorithm design, while ensuring data privacy and security, the flexibility and practicality of data analysis are significantly improved, meeting the diverse needs in different scenarios.

[0016] 2. The present invention supports selective data aggregation with random coding and enhanced permissions. By encoding attributes and combining encryption algorithms to generate an encrypted matrix sequence, the present invention can support secure and flexible decision-making based on user attributes. This technology allows complex calculations to be performed on encrypted data while protecting data privacy, ensuring the security and integrity of data during processing. This design not only optimizes the calculation efficiency but also provides more reliable privacy protection for data-driven decision-making.

[0017] 3. The present invention optimizes the application of encryption algorithms and adopts homomorphic encryption algorithms. This algorithm not only has high-intensity security but also can operate efficiently in resource-constrained environments. This design not only optimizes the calculation efficiency but also provides more reliable privacy protection for data-driven decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic flowchart of the selective data aggregation method based on conditional classification coding of the present solution.

[0019] Figure 2 is a logical flowchart of the selective data aggregation method based on conditional classification coding of the present solution.

[0020] Figure 3 is a schematic hardware structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention shall fall within the protection scope of the present invention.

[0022] Those skilled in the art should understand that in the disclosure of the present invention, the orientation or positional relationships indicated by the terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting the present invention.

[0023] Embodiment 1

[0024] This solution provides a selective data aggregation method based on conditional classification coding. By introducing a classification coding system, parallel processing of multi-attribute data is achieved, significantly improving the efficiency and accuracy of data aggregation. At the same time, by dynamically adjusting the conditional screening strategy, the flexibility and adaptability of the data aggregation process are ensured, enabling it to handle data requirements in different scenarios. It not only meets diverse data requirements but also avoids unnecessary data exposure, thus achieving the effect of maximizing data privacy protection.

[0025] Specifically, as Figure 1 shown, the selective data aggregation method based on conditional classification coding includes the following steps:

[0026] Obtain a data analysis request, where the data analysis request records the task conditions and task requirements of the current data analysis task;

[0027] Based on the data analysis request, collect the segmented user data of users in the same area, and perform encryption processing on the segmented user data to obtain an encrypted matrix sequence. The segmented user data includes encoded user data corresponding to ordinary attributes and private attribute user data of private attributes.

[0028] Perform selective aggregation on the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data;

[0029] Conduct a compliance assessment on the preliminary aggregated data and screen the eligible screened data;

[0030] Perform secondary aggregation on the screened data to obtain the final aggregated data.

[0031] The system for implementing the selective data aggregation method based on conditional classification coding in this solution consists of five entities: users in different regions, an aggregation server, an authoritative institution, data requesters, and a trusted third party. Among them, data requesters refer to users who issue data analysis requests. When a data requester needs to obtain user data that meets specific conditions to support statistical analysis decisions, it sends a corresponding data analysis request to the aggregation server. Among them, users in different regions refer to users to be analyzed and decision-making, who provide user data for the aggregation server to support statistical analysis decisions. Among them, the aggregation server performs data processing work for data aggregation. Among them, the authoritative institution conducts compliance evaluation on the preliminarily aggregated data. Among them, the trusted third party provides random parameters for the cryptographic system in the data encryption process.

[0032] Specifically, when a data requester needs to obtain user data that meets specific task conditions to support statistical analysis decisions, it sends a data analysis request to the aggregation server. Users within the same region perform segmentation coding and encryption processing on personal attributes according to the requirements of the preset task conditions to generate an encrypted matrix sequence. The aggregation server performs specific aggregation operations on the received encrypted matrix sequence according to the task requirements and sends the preliminarily aggregated data to the authoritative institution for compliance evaluation. The authoritative institution filters out eligible users based on the preset conditions and feeds back the filtered data to the aggregation server. The aggregation server then performs secondary aggregation on the encrypted data of the eligible users to ensure that the data requester only receives a data set that meets the conditions and has undergone privacy protection processing. This process, through multi-level security mechanisms and strict permission control, can not only efficiently screen out the target user group but also comprehensively guarantee user privacy and data security throughout the entire data processing process, providing reliable technical support for data-driven decision-making.

[0033] In the "Obtain Data Analysis Request" step, the data requester sends a data analysis request to the aggregation server, where the data analysis request records the task conditions and task requirements of the current data analysis task.

[0034] Specifically, the data analysis task is to obtain user data that meets specific task conditions and specific task requirements. Task conditions refer to the data attributes of the user data to be aggregated. Task requirements specify that the classification of the data attributes within the task conditions is a common attribute or a private attribute, where common attributes are divided into boolean attributes and numerical attributes.

[0035] Exemplarily, when the data analysis request sent by the data requester is "Obtain the average daily working hours of adolescent males in a certain region", where the task conditions are "region, adolescent, male, working hours", and the task requirements specify that "region, adolescent, male" are boolean attributes among common attributes, and "working hours" is a private attribute.

[0036] Similarly, when the data analysis request sent by the data requester is "obtain the daily Internet access time of users aged [30, 50] in a certain region", where the task conditions are "region, [r, s]=[30, 50], Internet access time", and the task requirements stipulate that "region" is a boolean attribute in the general attributes, "[r, s]=[30, 50]" is a numerical attribute in the general attributes, and "Internet access time" is a private attribute.

[0037] In the step of "collecting segmented user data of users in the same region based on the data analysis request":

[0038] Based on the data analysis request, obtain the general attribute user data and private attribute user data uploaded by the user, and encode the general attribute user data to obtain encoded user data, where the general attribute user data is the user data corresponding to the general attribute.

[0039] It should be noted that since the general attributes and private attributes are specified in the task requirements, this solution can directly divide the obtained user data into general attribute user data and private attribute user data according to the task requirements. In other words, the user uploads the attribute value corresponding to the general attribute as the general attribute user data, and uploads the attribute value corresponding to the private attribute as the private attribute user data.

[0040] Furthermore, since the general attributes include boolean attributes and numerical attributes, different encoding methods are selected for different types of general attributes, as follows:

[0041] When the general attribute is selected as a boolean attribute, in the step of "encoding the general attribute user data to obtain encoded user data", compare the general attribute user data with the task request. When the general attribute user data is the same as the value of the general attribute in the task request, it is encoded as 1. When the general attribute user data is different from the value of the general data in the task request, it is encoded as any integer other than 1 and 0.

[0042] The corresponding method for encoding the general attribute user data to obtain encoded user data is as follows:

[0043] ;

[0044] where is the attribute value of the obtained general attribute user data, τ is 1, is any integer not equal to τ, is the encoded user data.

[0045] Exemplarily, for example, if the user v1 is "middle-aged male in this region", then the encoded user data , and the private attribute user data is x1.

[0046] When the ordinary attribute is selected as a numerical attribute, in the step of "encoding the user data of the ordinary attribute to obtain the encoded user data", construct the full numerical value of the ordinary attribute corresponding to the current ordinary attribute user data, compare the numerical value of the ordinary attribute user data with the full numerical value, replace the position corresponding to the current ordinary attribute user data in the full numerical value with 1, and set other positions to any integer other than 1 and 0 to obtain the encoded user data.

[0047] The corresponding method for encoding the user data of the ordinary attribute to obtain the encoded user data is as follows:

[0048] ;

[0049] where is 1, τ is any integer not equal to , is the encoded user data, is the attribute value of the obtained ordinary attribute user data, is the j-th numerical value in the full numerical value.

[0050] Exemplarily, for example, the age of user v1 is "35 years old", and the full numerical value is defined as Then the encoded user data , where y 35 is τ.

[0051] The encryption processes for both boolean attributes and numerical attributes are the same. Specifically, in the step of "encrypting the segmented user data to obtain the encrypted matrix sequence", use the ElGamal encryption algorithm to encrypt all the encoded user data to obtain the first encrypted data, use the Paillier algorithm to encrypt all the private attribute user data to obtain the second encrypted data, and integrate the first encrypted data and the second encrypted data of all users to obtain the encrypted matrix sequence, where the encrypted matrix sequence records the first encrypted data and the second encrypted data of each user.

[0052] Furthermore, the ElGamal encryption algorithm includes the public key and private key of the ElGamal encryption system for encrypting ordinary attributes and the ElGamal random parameters sent by a trusted third party; the Paillier cryptosystem includes the public key of the Paillier cryptosystem for encrypting private attributes and the Paillier random parameters sent by a trusted third party.

[0053] Specifically, before the step of "encrypting user data to obtain an encrypted matrix sequence", the steps include: generating the public key and private key of the ElGamal encryption system for encrypting ordinary attributes and the public key of the Paillier cryptosystem for encrypting private attributes, and obtaining the ElGamal random parameter and Paillier random parameter sent by a trusted third party.

[0054] Further, select a security parameter and a generator g, and generate the public key and private key of the ElGamal encryption system for encrypting ordinary attributes based on the generator g, where the security parameter is used to control the security of encryption, and the larger the security parameter, the higher the corresponding security.

[0055] The formula for generating the public key of the ElGamal encryption system for encrypting ordinary attributes is as follows: ;

[0056] The formula for generating the private key of the ElGamal encryption system for encrypting ordinary attributes is as follows:

[0057] ;

[0058] where g is the generator, α is the randomly selected private key, is the public key of the ElGamal encryption system, is the public key parameter, is the private key of the ElGamal encryption system.

[0059] Further, select large prime numbers p and q, and generate the public key of the Paillier cryptosystem for encrypting private attributes based on the large prime numbers. The formula is as follows:

[0060] ;

[0061] where p and q are large prime numbers, is the public key of the Paillier cryptosystem, and N is the public key parameter of the Paillier cryptosystem.

[0062] The ElGamal encryption system of this solution is used to protect the privacy of ordinary attributes, and the Paillier encryption system is used to protect the private attributes of users.

[0063] Further, the trusted third party is responsible for generating random parameters for subsequent encryption and data aggregation. The trusted third party selects the ElGamal random parameter k from the integer ring of modulo j , where 1 ≤ j ≤ l, and l represents the number of ordinary attributes. The ElGamal random parameter k jUsed to encrypt ordinary attributes based on the ElGamal encryption system; where the trusted third party selects random parameters of the Paillier cryptosystem from the multiplicative group of modulo and where 1 ≤ i ≤ n and n is the modulus. Additionally, regarding the random parameters of the Paillier cryptosystem compute special random numbers where and the random parameters of the Paillier cryptosystem satisfy the following constraint formula:

[0064] ;

[0065] where λ is the security parameter of the Paillier encryption system, usually related to N This constraint ensures the correctness and consistency of the encrypted data.

[0066] The ElGamal encryption algorithm adopted in this scheme is an asymmetric encryption algorithm based on discrete logarithms. A random number is required during the encryption process to ensure the uniqueness of each encryption result. The randomness of the random parameter ensures that the encrypted result of ordinary attributes cannot be inferred by attackers by observing the ciphertext; Paillier is a symmetric encryption algorithm that supports additive homomorphisms, and a random number is also required during the encryption process to ensure security. The randomness of the random parameter ensures that the user's private attributes cannot be directly decrypted after encryption, while the random number ensures that the correctness of the encrypted data can also be guaranteed during data aggregation.

[0067] Correspondingly, when the ordinary attribute is a boolean attribute, the formula for encrypting all encoded user data using the ElGamal encryption algorithm to obtain the first encrypted data is as follows:

[0068] ;

[0069] where represents the first encrypted data of the j-th ordinary attribute of the i-th user, is the random parameter of the ElGamal encryption algorithm for the j-th ordinary attribute of the i-th user, g is the generator, β is the core parameter of the public key, is the j-th encoded user data of the i-th user, g is the primitive root of the generator, is the random parameter, is the -th power result of the generator g, β is the public key of the receiver, is the -th power result of the public key, and p is a large prime number;

[0070] The formula for encrypting all private attribute user data using the Paillier algorithm to obtain the second encrypted data is as follows:

[0071] ;

[0072] ;

[0073] where is the second encrypted data of the i-th user, p and q are large prime numbers, is the random parameter of the Paillier algorithm for the i-th user, x i is the private attribute user data, h is a random number, is the power of the random number, used to ensure that the attacker cannot infer the plaintext or the random number used for encryption by observing the ciphertext , N is the public key parameter of the Paillier cryptosystem;

[0074] The encrypted matrix sequence obtained by integrating the first encrypted data and the second encrypted data is as follows:

[0075] ;

[0076] where represents the first encrypted data of the l-th ordinary attribute of the n-th user, is the second encrypted data of the n-th user, and n is the total number of users.

[0077] When the ordinary attribute is a boolean attribute, the formula for encrypting all encoded user data using the ElGamal encryption algorithm to obtain the first encrypted data is as follows:

[0078] ;

[0079] where represents the first encrypted data of the j-th ordinary attribute of the i-th user, is the random parameter of the ElGamal encryption algorithm for the j-th ordinary attribute of the i-th user, g is the generator, β is the core parameter of the public key, is the encoded user data obtained by segmenting and encoding the ordinary attribute data of the i-th user. g is the primitive root of the generator, is the random parameter, is the power result of the generator g, β is the public key of the receiver, is the power result of the public key, p is a large prime number.

[0080] The formula for encrypting all private attribute user data using the Paillier algorithm to obtain the second encrypted data is as follows:

[0081] ;

[0082] ;

[0083] where is the second encrypted data of the i-th user, p and g are large prime numbers, is the random parameter of the Paillier algorithm for the i-th user, x i is the private attribute user data, h is a random number, is the power of the random number, used to ensure that the attacker cannot infer the plaintext or the random number used for encryption by observing the ciphertext , N is the public key parameter of the Paillier cryptosystem;

[0084] The encrypted matrix sequence obtained by integrating the first encrypted data and the second encrypted data is as follows:

[0085] ;

[0086] where the first l columns of the encrypted matrix sequence are the encrypted values of the attribute segmentation encoding of the t-th user, corresponding to the first encrypted data of different users is the second encrypted data of the t-th user, 1 ≤ t ≤ n, and n is the total number of users.

[0087] Furthermore, when the ordinary attribute is a boolean attribute:

[0088] In the step of "selectively aggregating the encrypted matrix sequence based on task requirements to obtain the preliminary aggregated data", an aggregation operation is performed on the first encrypted data of each user in the encrypted matrix sequence to obtain an encrypted sequence as the preliminary aggregated data.

[0089] The obtained preliminary aggregated data is where where represents the first encrypted data of the j-th ordinary attribute of the i-th user, p is a large prime number, represents the aggregation of the i-th user, n is the total number of users, and P is the preliminary aggregated data. It should be noted that in this scheme, it is necessary to judge whether the i-th user meets all the conditions in the task conditions, then represents the aggregated product value of all the first encrypted data of the current user as the preliminary aggregated data, and subsequent encryption of this preliminary aggregated data can be performed for judgment.

[0090] In the step of "performing compliance assessment on the preliminary aggregated data and screening eligible screened data", the authoritative institution decrypts the preliminary aggregated data using the private key of the ElGamal encryption algorithm to obtain the cumulative product value of the ordinary attributes of each user. If the cumulative product value is 1, the first encrypted data and the second encrypted data of the current user are retained; otherwise, the first encrypted data and the second data of the current user are discarded to obtain the screened data.

[0091] In the step of "performing secondary aggregation on the screened data to obtain the final aggregated data", the second encrypted data in the screened data is aggregated to obtain the final aggregated data.

[0092] Correspondingly, the final aggregated data is , where m is the number of users in the screened data, 1 ≤ i ≤ m, N is the public key parameter of the Paillier cryptosystem, is the final aggregated data, is the second encrypted data of the i-th user. This formula means that assuming that after screening, the data of m users meet the requirements, then the private attributes of these m users are aggregated.

[0093] Meanwhile, the trusted third party sends the private key of the Paillier cryptosystem and the parameters of the users who do not meet the task conditions to the data requester. The data requester uses the parameters to complete the decryption operation of the final aggregated result. The specific decryption process is as follows: ; and calculate to obtain the aggregated value of the private attributes.

[0094] Among them is the encoded plaintext part, is the randomization factor used to protect the plaintext information, is the private key of the paillier encryption algorithm, is the random parameter of the unselected users, satisfying . As can be seen from the above, is . The meaning of the formula of this decryption process is to make up this private key . Because in the paillier system this is 1, the equation in the decryption process can be expressed as .

[0095] Furthermore, when the ordinary attribute is a numerical attribute:

[0096] In the step of "performing selective aggregation on the encrypted matrix sequence based on the task requirements to obtain the preliminary aggregated data", the first encrypted data of each user in the encrypted matrix sequence that meets the task conditions is aggregated to obtain an encrypted sequence as the preliminary aggregated data.

[0097] Correspondingly, the obtained preliminary aggregated data is , where , where represents the first encrypted data of the j-th ordinary attribute of the i-th user, 1 ≤ i ≤ n, H is the preliminary aggregated data, is the preliminary aggregation or data of the i-th user, s represents the position of the right endpoint, r represents the position of the left endpoint, p represents a large prime number.

[0098] However, assuming the universal set U = 100, then the ordinary attribute data of each user has to be divided into 100 parts and encoded and encrypted to obtain the first encrypted data. If we want to select users satisfying [30, 50], then when aggregating the encrypted matrix sequence, we need to aggregate the ciphertexts at positions [30, 50], that is, we need to aggregate and , corresponding to s = 30 and r = 50.

[0099] In the step of "conducting compliance assessment on the preliminary aggregated data and screening the qualified screened data", the authoritative institution decrypts the preliminary aggregated data using the private key of the ElGamal encryption algorithm to obtain the product value of the ordinary attributes of each user. If the product value is not 1, then retain the first encrypted data and the second encrypted data of the current user; otherwise, discard the first encrypted data and the second data of the current user to obtain the screened data.

[0100] In the step of "conducting secondary aggregation on the screened data to obtain the final aggregated data", aggregate the second encrypted data in the screened data to obtain the final aggregated data.

[0101] The final aggregated data is , where m is the number of users in the screened data, 1 ≤ i ≤ m, is the second encrypted data of the i-th user, N represents the public key parameter of the Paillier cryptosystem, is the final aggregated data. This formula means that assuming that after screening, the data of m users meet the requirements, then aggregate the private attributes of these m users.

[0102] Meanwhile, the trusted third party sends the private key of the Paillier cryptosystem and the parameters of the users who do not meet the task conditions to the data requester, and the data requester uses the parameters to complete the decryption operation of the final aggregation result. The specific decryption process is the same as above.

[0103] In summary, the selective data aggregation method based on conditional classification coding provided by this solution can complete the selective aggregation operation without revealing the ordinary attributes and private attributes of users. The logic flowchart of the selective data aggregation method based on conditional classification coding is as Figure 2 shown: Collect user data of users and perform classification coding on the user data to obtain segmented user data, where the segmented user data includes coded user data corresponding to ordinary attributes and private attribute user data of private attributes; perform encryption processing on the coded user data and private attribute user data to obtain an encrypted matrix sequence; the aggregation server aggregates the coded user data to obtain initial aggregated data and is judged by the authoritative agency. If the conditions are met, the aggregation server then aggregates the private attribute user data to obtain the final aggregated data, and the final aggregated data is sent to the data requester for statistical analysis.

[0104] Embodiment 2

[0105] This solution provides a selective data aggregation system based on conditional classification coding, which consists of a trusted third party, users, an aggregation server, an authoritative agency, and data requesters:

[0106] The data requester sends a data analysis request to the aggregation server, where the data analysis request records the task conditions and task requirements of the current data analysis task;

[0107] When receiving the data analysis request, collect the segmented user data of users in the same area based on the data analysis request, and perform encryption processing on the segmented user data to obtain an encrypted matrix sequence, where the segmented user data includes coded user data corresponding to ordinary attributes and private attribute user data of private attributes, and the trusted third party provides the parameters of the encryption algorithm to the users;

[0108] The aggregation server performs selective aggregation on the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data, and sends the preliminary aggregated data to the authoritative agency;

[0109] The authoritative structure conducts a compliance assessment on the preliminary aggregated data and screens the qualified screened data, and sends the screened data to the aggregation server again;

[0110] The aggregation server performs secondary aggregation on the screened data to obtain the final aggregated data, and sends the final aggregated data to the data requester.

[0111] When receiving a data analysis request from a data requester, the trusted third party initializes and sends the parameters of the ElGamal encryption algorithm and the Paillier algorithm to the user. The user encrypts the segmented user data based on the ElGamal encryption algorithm and the Paillier algorithm to obtain an encrypted matrix sequence, and sends the encrypted matrix sequence to the aggregation server. The aggregation server aggregates the encoded user data in the encrypted matrix sequence to obtain initial aggregated data, and sends the initial aggregated data to the authoritative institution. After the authoritative institution makes a judgment, it obtains screened data, and the screened data is sent to the aggregation server again for aggregation. At this time, the aggregation server aggregates the private user data to obtain the final aggregated data, and sends the final aggregated data to the data requester. At the same time, the trusted third party also sends the parameters of the ElGamal encryption algorithm and the Paillier algorithm to the data requester.

[0112] Embodiment III

[0113] This embodiment also provides an electronic device. Refer to Figure 3 , including a memory 404 and a processor 402. A computer program is stored in the memory 404, and the processor 402 is configured to run the computer program to execute the steps in any of the above embodiments of the selective data aggregation method based on conditional classification coding.

[0114] Specifically, the above-mentioned processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC for short), or may be configured as one or more integrated circuits implementing the embodiments of the present application.

[0115] Among them, the memory 404 may include a large-capacity memory 404 for data or instructions. The memory 404 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 402.

[0116] The processor 402 reads and executes the computer program instructions stored in the memory 404 to implement any one of the above-mentioned selective data aggregation methods based on conditional classification coding.

[0117] Optionally, the above-mentioned electronic device may further include a transmission device 406 and an input / output device 408. Among them, the transmission device 406 is connected to the above-mentioned processor 402, and the input / output device 408 is connected to the above-mentioned processor 402.

[0118] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above-mentioned network can include wired or wireless networks provided by a communication provider of an electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 406 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0119] The input / output device 408 is used to input or output information. In this embodiment, the input information can be a data analysis request, etc., and the output information can be the final aggregated data, etc.

[0120] Optionally, in this embodiment, the above-mentioned processor 402 can be set to execute the following steps through a computer program:

[0121] Obtain a data analysis request, where the data analysis request records the task conditions and task requirements of the current data analysis task;

[0122] Collect the segmented user data of users in the same area based on the data analysis request, and perform encryption processing on the segmented user data to obtain an encrypted matrix sequence, where the segmented user data includes encoded user data corresponding to ordinary attributes and private attribute user data of private attributes,

[0123] Selectively aggregate the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data;

[0124] Conduct a compliance assessment on the preliminary aggregated data and screen the qualified screened data;

[0125] Perform secondary aggregation on the screened data to obtain the final aggregated data.

[0126] It should be noted that the specific examples in this embodiment can refer to the examples described in the above-mentioned embodiment and optional implementation manners, and will not be elaborated herein.

[0127] In general, various embodiments can be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. Some aspects of the present invention can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, a microprocessor, or other computing devices, but the present invention is not limited thereto. Although various aspects of the present invention may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, as a non-limiting example, the blocks, devices, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or a controller or other computing devices, or some combination thereof.

[0128] Embodiments of the present invention can be implemented by computer software that is executable by a data processor of a mobile device, such as in a processor entity, or by hardware, or by a combination of software and hardware. A computer software or program (also referred to as a program product), including software routines, applets, and / or macros, can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product can include one or more computer-executable components that are configured to perform the embodiments when the program runs. One or more computer-executable components can be at least one software code or a part thereof. Additionally, in this regard, it should be noted that any block of the logical flow in the figures can represent a program step, or an interconnected logical circuit, block, and function, or a combination of program steps and logical circuits, blocks, and functions. The software can be stored on physical media such as memory chips or storage blocks implemented within the processor, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and their data variants, CDs. The physical media are non-transitory media.

[0129] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0130] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A selective data aggregation method based on conditional classification coding, characterized in that, The steps include: Obtain a data analysis request, in which the task conditions and task requirements of the current data analysis task are recorded. The data analysis task is to obtain user data that meets the task conditions and task requirements. The task conditions refer to the data attributes of the user data to be aggregated, and the task requirements specify that the classification of the data attributes in the task conditions is a common attribute or a private attribute, where the common attributes are divided into boolean attributes and numerical attributes; Based on the data analysis request, collect the segmented user data of users in the same area, and perform encryption processing on the segmented user data to obtain an encrypted matrix sequence, where the segmented user data includes the encoded user data corresponding to the common attributes and the private attribute user data of the private attributes; Perform selective aggregation on the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data; Conduct a compliance assessment on the preliminary aggregated data and filter the qualified filtered data; Perform secondary aggregation on the filtered data to obtain the final aggregated data.

2. The selective data aggregation method based on conditional classification coding according to claim 1, wherein Based on the data analysis request, obtain the common attribute user data and private attribute user data uploaded by the user, and encode the common attribute user data to obtain encoded user data, where the common attribute user data is the user data corresponding to the common attributes.

3. The selective data aggregation method based on conditional classification coding according to claim 1, wherein When the common attribute is selected as a boolean attribute, compare the common attribute user data with the task request. When the common attribute user data is consistent with the value of the common attribute in the task request, it is encoded as 1. When the common attribute user data is inconsistent with the value of the common data in the task request, it is encoded as any integer other than 1 and 0.

4. The selective data aggregation method based on conditional classification coding according to claim 1, wherein When the common attribute is selected as a numerical attribute, construct the numerical full value of the common attribute corresponding to the current common attribute user data, compare the value of the common attribute user data with the numerical full value, replace the position corresponding to the current common attribute user data in the numerical full value with 1, and the other positions are any integer other than 1 and 0 to obtain the encoded user data.

5. The selective data aggregation method based on conditional classification coding according to claim 1, characterized in that Use the ElGamal encryption algorithm to encrypt all the encoded user data to obtain the first encrypted data, use the Paillier algorithm to encrypt all the private attribute user data to obtain the second encrypted data, and integrate the first encrypted data and the second encrypted data of all users to obtain an encrypted matrix sequence, where the encrypted matrix sequence records the first encrypted data and the second encrypted data of each user.

6. The selective data aggregation method based on conditional classification coding according to claim 5, wherein When the common attribute is a boolean attribute, perform an aggregation operation on the first encrypted data of each user in the encrypted matrix sequence to obtain an encrypted sequence as the preliminary aggregated data. The authoritative agency decrypts the preliminary aggregated data using the private key of the ElGamal encryption algorithm to obtain the cumulative product value of the common attributes of each user. If the cumulative product value is 1, retain the first encrypted data and the second encrypted data of the current user; otherwise, discard the first encrypted data and the second data of the current user to obtain the filtered data, and aggregate the second encrypted data in the filtered data to obtain the final aggregated data.

7. The selective data aggregation method based on conditional classification coding according to claim 5, wherein When the ordinary attribute is a numerical attribute, perform the first encrypted data aggregation operation that meets the task conditions for each user in the encrypted matrix sequence to obtain an encrypted sequence as the preliminary aggregated data. The authoritative agency decrypts the preliminary aggregated data using the private key of the ElGamal encryption algorithm to obtain the product value of the ordinary attributes of each user. If the product value is not 1, retain the first encrypted data and the second encrypted data of the current user; otherwise, discard the first encrypted data and the second data of the current user to obtain the filtered data, and aggregate the second encrypted data in the filtered data to obtain the final aggregated data.

8. A selective data aggregation system based on conditional classification coding, characterized in that, It consists of a trusted third party, users, an aggregation server, an authoritative agency, and data requesters: The data requester sends a data analysis request to the aggregation server. The data analysis request records the task conditions and task requirements of the current data analysis task. The data analysis task is to obtain user data that meets the task conditions and task requirements. The task conditions refer to the data attributes of the user data to be aggregated. The task requirements specify that the classification of the data attributes in the task conditions is an ordinary attribute or a private attribute, where the ordinary attributes are divided into boolean attributes and numerical attributes; When receiving the data analysis request, collect the segmented user data of users in the same area based on the data analysis request, and perform encryption processing on the segmented user data to obtain an encrypted matrix sequence. The segmented user data includes the encoded user data corresponding to the ordinary attributes and the private attribute user data of the private attributes. The trusted third party provides the parameters of the encryption algorithm to the users; The aggregation server performs selective aggregation on the encrypted matrix sequence based on the data analysis request to obtain the preliminary aggregated data, and sends the preliminary aggregated data to the authoritative agency; The authoritative structure conducts a compliance assessment on the preliminary aggregated data and filters the qualified filtered data, and sends the filtered data to the aggregation server again; The aggregation server performs secondary aggregation on the filtered data to obtain the final aggregated data, and sends the final aggregated data to the data requester.

9. A readable storage medium, characterized in that, The readable storage medium stores a computer program. The computer program includes program code for controlling a process to execute the process. The process includes the selective data aggregation method based on conditional classification coding according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • A method, device, terminal and medium for data aggregation

    CN118916241B

  • Information coding method and device, equipment and storage medium

    CN115204111A