Selective data aggregation method and system based on conditional classification coding
Through the selective data aggregation method based on conditional classification encoding, the feature extraction problem in multi-attribute data processing is solved, and efficient and flexible data aggregation and privacy protection are achieved.
Patent Information
- Application Number
- CN202510520997.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-24
AI Technical Summary
It is difficult for the prior art to accurately extract effective features when processing multi-attribute data, especially for diversified attribute features of non-numeric types. How to carry out unified and efficient processing is still a key issue that needs to be solved urgently.
The selective data aggregation method based on conditional classification encoding is adopted to realize the selective data aggregation operation through the encoding mechanism and conditional screening strategy to ensure the security and integrity of data privacy.
It realizes efficient aggregation of multi-attribute data, improves the flexibility and practicality of data analysis, meets the diverse needs in different scenarios, and ensures the protection of data privacy.
Smart Images

Figure CN120030577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data aggregation, and in particular to a selective data aggregation method and system based on conditional classification coding. Background Art
[0002] With the rapid development of information technology, data plays a vital role in all fields of today's society. Whether it is business decision-making, medical research or social governance, it is highly dependent on the accuracy and completeness of data uploaded by data users. In this context, data users are no longer satisfied with simple data uploading and basic processing, but have put forward more stringent requirements, hoping to mine valuable information from massive and complex data to provide strong support for various decisions.
[0003] In the process of data statistical analysis, the choice of data upload processing method is crucial. Traditional data upload methods often directly upload a large amount of raw data to the server. This method not only occupies a large amount of communication resources, resulting in high communication overhead, but also occupies a huge storage cost on the storage side, which brings huge pressure to data storage and management. In order to solve these problems, the use of data aggregation for upload processing has gradually become a mainstream trend. Data aggregation can organize and merge scattered data in a targeted manner, effectively reducing the amount of data transmission during the communication process, thereby reducing communication overhead; at the same time, at the storage level, aggregated data also greatly reduces the required storage space and reduces storage costs. Moreover, data aggregation can also remove redundant information in the data, improve the overall efficiency of data processing, and make subsequent data analysis and mining work more efficient and accurate.
[0004] However, in the context of a distributed privacy protection model, data aggregation faces even more severe and complex challenges. In a distributed system, user groups in different regions often have unique attributes and private data due to their geographical location, social environment, consumption habits and many other factors. This requires that when data aggregation is performed, the user's privacy information must be strictly protected from being leaked while ensuring the accuracy and integrity of the data.
[0005] Existing methods often find it difficult to accurately extract effective features when processing multi-attribute data, especially for diversified attribute features of non-numerical types. How to process them in a unified and efficient manner is still a key issue that needs to be solved urgently. Chinese invention patent CN202411395708.1 proposes a method that can accurately aggregate data to attributes, aiming to solve the problem of non-aggregate field data loss in the prior art. The solution can distinguish between aggregated attributes and non-aggregate attributes through attribute separation and identification, ensuring the integrity and accuracy of reported data. However, the invention can only identify a single attribute at a time, which has certain limitations in practical applications. In actual data scenarios, in many cases, it is necessary to process the aggregation of multiple attributes at the same time. For example, when comprehensively analyzing user purchasing behavior, it may be necessary to consider multiple attributes such as the user's age, occupation, purchase time, purchase amount, etc. The invention cannot be efficiently processed when facing complex scenarios where multiple attributes are aggregated at the same time. This not only limits its promotion scope in practical applications, but also affects the effect and efficiency of data aggregation to a certain extent.
[0006] Therefore, how to provide a data aggregation method that meets the data demander's balance needs between privacy protection and data availability is a key issue that needs to be solved urgently. Summary of the invention
[0007] The purpose of the present invention is to provide a selective data aggregation method and system based on conditional classification coding, which, through a coding mechanism and a conditional screening strategy, can accurately implement the selective aggregation operation of data according to the specific task conditions set by the data demander while ensuring that the privacy of the private attributes and common attributes of the data owner is not violated.
[0008] To achieve the above objectives, the present technical solution provides a selective data aggregation method based on conditional classification coding, comprising the following steps: Obtaining a data analysis request, wherein the data analysis request records the task conditions and task requirements of the current data analysis task; Based on the data analysis request, segmented user data of users in the same area are collected, and the segmented user data are encrypted to obtain an encrypted matrix sequence, wherein the segmented user data includes coded user data corresponding to common attributes and private attribute user data corresponding to private attributes. Selectively aggregate the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data; Conduct compliance assessment on the preliminary aggregated data and screen the eligible screening data; The filtered data are aggregated twice to obtain the final aggregated data.
[0009] Compared with the prior art, this technical solution has the following characteristics and beneficial effects: 1. The present invention proposes an innovative selective data aggregation method based on conditional classification coding, which can realize data security aggregation for multiple statistical function requirements, and support customized statistical analysis and attribute-based precise aggregation. By introducing innovative privacy protection mechanisms and efficient algorithm design, the flexibility and practicality of data analysis are significantly improved while ensuring data privacy and security, thus meeting the diverse needs in different scenarios.
[0010] 2. The present invention supports random coding and selective data aggregation with enhanced permissions. By encoding attributes and combining encryption algorithms to generate encryption matrix sequences, the present invention can support secure and flexible decision-making based on user attributes. This technology allows complex calculations to be performed on encrypted data while protecting data privacy, ensuring the security and integrity of data during processing. This design not only optimizes computing efficiency, but also provides more reliable privacy protection for data-driven decision-making.
[0011] 3. The present invention optimizes the application of encryption algorithms and adopts homomorphic encryption algorithms. This algorithm not only has high-intensity security, but also can run efficiently in resource-constrained environments. This design not only optimizes computing efficiency, but also provides more reliable privacy protection for data-driven decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a flowchart of the selective data aggregation method based on conditional classification coding of this scheme.
[0013] Figure 2 It is a logical flow chart of the selective data aggregation method based on conditional classification coding of this scheme.
[0014] Figure 3 It is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0015] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention.
[0016] Those skilled in the art should understand that, in the disclosure of the present invention, the terms "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the drawings, which are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the above terms should not be understood as limiting the present invention.
[0017] Embodiment 1 This solution provides a selective data aggregation method based on conditional classification coding. By introducing a classification coding system, it realizes the parallel processing of multi-attribute data, significantly improving the efficiency and accuracy of data aggregation. At the same time, by dynamically adjusting the conditional screening strategy, it ensures the flexibility and adaptability of the data aggregation process, and can cope with data requirements in different scenarios. It not only meets diverse data needs, but also avoids unnecessary data exposure, thereby protecting data privacy to the greatest extent.
[0018] Specifically, Figure 1 As shown, the selective data aggregation method based on conditional classification coding includes the following steps: Obtaining a data analysis request, wherein the data analysis request records the task conditions and task requirements of the current data analysis task; Based on the data analysis request, segmented user data of users in the same area are collected, and the segmented user data are encrypted to obtain an encrypted matrix sequence, wherein the segmented user data includes coded user data corresponding to common attributes and private attribute user data corresponding to private attributes. Selectively aggregate the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data; Conduct compliance assessment on the preliminary aggregated data and screen the eligible screening data; The filtered data are aggregated twice to obtain the final aggregated data.
[0019] The system for implementing the selective data aggregation method based on conditional classification coding in this scheme consists of five entities: users in different regions, aggregation servers, authoritative institutions, data demanders and trusted third parties, wherein data demanders refer to users who issue data analysis requests. When data demanders need to obtain user data that meets specific conditions to support statistical analysis decisions, they send responsive data analysis requests to the aggregation server; wherein users in different regions refer to users whose decisions are to be analyzed, and they provide user data to the aggregation server to support statistical analysis decisions; wherein the aggregation server performs data processing for data aggregation; wherein the authoritative institutions conduct compliance assessments on preliminary aggregated data; wherein the trusted third party provides random parameters for the cryptographic system in the data encryption processing process.
[0020] Specifically, when the data demander needs to obtain user data that meets specific task conditions to support statistical analysis decisions, it will issue a data analysis request to the aggregation server; users in the same area segment and encode personal attributes according to the requirements of the preset task conditions and implement encryption processing to generate an encrypted matrix sequence. The aggregation server performs specific aggregation operations on the received encrypted matrix sequence according to the task requirements, and sends the preliminary aggregated data to the authority for compliance assessment. The authority selects qualified users based on the preset conditions and feeds the selected data back to the aggregation server. The aggregation server then performs secondary aggregation on the encrypted data of qualified users to ensure that the data demander only receives qualified and privacy-protected data sets. This process, through multi-level security mechanisms and strict authority control, can not only efficiently screen out the target user group, but also fully protect user privacy and data security throughout the entire data processing process, providing reliable technical support for data-driven decision-making.
[0021] In the step of "obtaining data analysis request", the data demander sends a data analysis request to the aggregation server, wherein the data analysis request records the task conditions and task requirements of the current data analysis task.
[0022] Specifically, the data analysis task is to obtain user data that meets specific task conditions and specific task requirements. The task conditions refer to the data attributes of the user data that need to be aggregated. The task requirements stipulate that the data attributes within the task conditions are classified as common attributes or private attributes, where common attributes are divided into Boolean attributes and numerical attributes.
[0023] For example, when the data demander sends a data analysis request such as "obtain the average daily working hours of teenage males in a certain region", where the task conditions are "region, teenagers, males, and online time", the task requirements stipulate that "region, teenagers, males" are Boolean attributes among common attributes, and "online time" is a private attribute.
[0024] Similarly, for example, when the data analysis request sent by the data demander is "obtain the daily online time of users in the [30,50] age group in a certain region", where the task conditions are "region, [r,s]=[30,50], online time", the task requirement stipulates that "region" is a Boolean attribute among common attributes, "[r,s]=[30,50]" is a numerical attribute among common attributes, and "online time" is a private attribute.
[0025] In the "Collect segmented user data for users in the same region based on data analysis requests" step: Based on the data analysis request, common attribute user data and private attribute user data uploaded by the user are obtained, and the common attribute user data is encoded to obtain encoded user data, wherein the common attribute user data is user data corresponding to common attributes.
[0026] It should be noted that since common attributes and private attributes are specified in the task requirements, this solution can directly divide the acquired user data into common attribute user data and private attribute user data according to the task requirements. In other words, the user uploads the attribute values corresponding to the common attributes as common attribute user data, and uploads the attribute values corresponding to the private attributes as private attribute user data.
[0027] Furthermore, since common attributes include Boolean attributes and numerical attributes, different encoding methods are selected for different types of common attributes, as follows: When the common attribute is selected as a Boolean attribute, in the step of "encoding the common attribute user data to obtain encoded user data", the common attribute user data is compared with the task request. When the common attribute user data is consistent with the value of the common attribute in the task request, it is encoded as 1; when the common attribute user data is inconsistent with the value of the common data in the task request, it is encoded as any integer other than 1 and 0.
[0028] The corresponding method of encoding the common attribute user data to obtain the encoded user data is as follows: ; in is the attribute value of the common attribute user data obtained, τ is 1, is any integer not equal to τ, To encode user data.
[0029] For example, user v 1 If it is "middle-aged male in the area", then encode the user data , private attribute user data is x 1 .
[0030] When the common attribute is selected as a numerical attribute, in the step of "encoding the common attribute user data to obtain encoded user data", the numerical full value of the common attribute corresponding to the current common attribute user data is constructed, the numerical value of the common attribute user data is compared with the numerical full value, the position corresponding to the current common attribute user data in the numerical full value is replaced with 1, and the other positions are any integers other than 1 and 0 to obtain the encoded user data.
[0031] The corresponding method of encoding the common attribute user data to obtain the encoded user data is as follows: ; in is 1, τ is not equal to Any integer of To encode user data, To obtain the attribute value of the common attribute user data, is the jth value in the total value of the value.
[0032] For example, user v 1 The age is "35 years old", and the full value of the definition value is Encode user data , where y 35 is τ.
[0033] The encryption process is consistent regardless of whether it is a Boolean attribute or a numerical attribute. Specifically, in the step of "encrypting the segmented user data to obtain an encryption matrix sequence", the ElGamal encryption algorithm is used to encrypt all encoded user data to obtain first encrypted data, the Paillier algorithm is used to encrypt all private attribute user data to obtain second encrypted data, and the first encrypted data and the second encrypted data of all users are integrated to obtain an encryption matrix sequence, wherein the encryption matrix sequence records the first encrypted data and the second encrypted data of each user.
[0034] Furthermore, the ElGamal encryption algorithm includes the public key and private key of the ElGamal encryption system for encrypting common attributes and the ElGamal random parameters sent by a trusted third party; the Paillier cryptographic system includes the public key of the Paillier cryptographic system for encrypting private attributes and the Paillier random parameters sent by a trusted third party.
[0035] Specifically, before the step of "encrypting user data to obtain an encrypted matrix sequence", the steps include: generating a public key and a private key of an ElGamal encryption system for encrypting common attributes and a public key of a Paillier cryptographic system for encrypting private attributes, and obtaining ElGamal random parameters and Paillier random parameters sent by a trusted third party.
[0036] Further, select security parameters And generator g, based on generator g, generate the public key and private key of the ElGamal encryption system for encrypting common attributes, where the security parameter is used to control the security of encryption, and the larger the security parameter, the higher the corresponding security.
[0037] The formula for generating the public key of the ElGamal encryption system that encrypts common attributes is as follows: ; The formula for generating the private key of the ElGamal encryption system that encrypts common attributes is as follows: ; Where g is the generator, α is a randomly selected private key, is the public key of the ElGamal encryption system, is the public key parameter, It is the private key of the ElGamal encryption system.
[0038] Furthermore, large prime numbers p and q are selected, and the public key of the Paillier cryptosystem for encrypting private attributes is generated based on the large prime numbers. The formula is as follows: ; Where p and q are large prime numbers, is the public key of the Paillier cryptosystem, and N is the public key parameter of the Paillier cryptosystem.
[0039] The ElGamal encryption system of this scheme is used to protect the privacy of common attributes, and the Paillier encryption system is used to protect the private attributes of users.
[0040] Furthermore, the trusted third party is responsible for generating random parameters for subsequent encryption and data aggregation, where the trusted third party generates random parameters from the integer ring of the modulus. Select the ElGamal random parameter k j , where 1≤j≤l, l represents the number of common attributes, and ElGamal random parameter k j Used to encrypt common attributes based on the ElGamal encryption system; the trusted third party uses the multiplication group of the module Select the random parameters of the Paillier cryptosystem , where 1≤i≤n, n is the modulus, and the random parameter of the Paillier cryptosystem Calculate special random numbers ,in and the random parameters of the Paillier cryptosystem The following constraint formula is satisfied: ; in λ is the security parameter of the Paillier encryption system, usually N Related, this constraint ensures the correctness and consistency of the encrypted data.
[0041] The ElGamal encryption algorithm used in this scheme is an asymmetric encryption algorithm based on discrete logarithms. A random number is required during the encryption process to ensure the uniqueness of each encryption result. The random parameter The randomness ensures that the encryption result of common attributes cannot be inferred by attackers by observing the ciphertext; Paillier is a symmetric encryption algorithm that supports additive homomorphism. Random numbers are also required in the encryption process to ensure security. The randomness ensures that the user's private attributes cannot be directly decrypted after encryption, while the random number ensures that the correctness of the encrypted data can also be guaranteed during data aggregation.
[0042] Correspondingly, when the common attribute is a Boolean attribute, the formula for encrypting all the encoded user data using the ElGamal encryption algorithm to obtain the first encrypted data is as follows: ; in represents the first encrypted data of the jth common attribute of the i-th user, is the random parameter of the ElGamal encryption algorithm for the jth common attribute of the i-th user, g is the generator, β is the core parameter of the public key, is the jth encoded user data of the i-th user, g is the primitive root of the generator, is a random parameter, is the generator g The result of the power, β is the public key of the recipient, For public key The result of the power, p is a large prime number; The formula for encrypting all private attribute user data using the Paillier algorithm to obtain the second encrypted data is as follows: ; ; in is the second encrypted data of the i-th user, p and q are large prime numbers, is the random parameter of the Paillier algorithm for the ith user, x i is private attribute user data, h is a random number, It is a random number The power is used to ensure that attackers cannot infer the plaintext or the random number used for encryption by observing the ciphertext. , N is the public key parameter of the Paillier cryptosystem; The encryption matrix sequence obtained by integrating the first encrypted data and the second encrypted data is as follows: ; in The first encrypted data representing the lth common attribute of the nth user, is the second encrypted data of the nth user, where n is the total number of users.
[0043] When the common attribute is a Boolean attribute, the formula for encrypting all the encoded user data using the ElGamal encryption algorithm to obtain the first encrypted data is as follows: ; in represents the first encrypted data of the jth common attribute of the i-th user, is the random parameter of the ElGamal encryption algorithm for the jth common attribute of the i-th user, g is the generator, β is the core parameter of the public key, The encoded user data g is the primitive root of the generator obtained by segmenting and encoding the common attribute data of the i-th user i. is a random parameter, is the generator g The result of the power, β is the public key of the recipient, For public key The result of the power is that p is a large prime number.
[0044] The formula for encrypting all private attribute user data using the Paillier algorithm to obtain the second encrypted data is as follows: ; ; in is the second encrypted data of the i-th user, p and g are large prime numbers, is the random parameter of the Paillier algorithm for the ith user, x i is private attribute user data, h is a random number, It is a random number The power is used to ensure that attackers cannot infer the plaintext or the random number used for encryption by observing the ciphertext. , N is the public key parameter of the Paillier cryptosystem; The encryption matrix sequence obtained by integrating the first encrypted data and the second encrypted data is as follows: ; The first l columns of the encryption matrix sequence are the encrypted values of the attribute segmentation code of the t-th user, corresponding to the first encrypted data of different users. is the second encrypted data of the t-th user, 1≤t≤n, and n is the total number of users.
[0045] Furthermore, when the normal attribute is a Boolean attribute: In the step of "selectively aggregating the encryption matrix sequence based on task requirements to obtain preliminary aggregated data", an aggregation operation is performed on the first encrypted data of each user in the encryption matrix sequence to obtain an encryption sequence as preliminary aggregated data.
[0046] Get the initial aggregate data as ,in ,in represents the first encrypted data of the jth common attribute of the i-th user, p is a large prime number, represents the aggregation of the i-th user, n is the total number of users, and P is the preliminary aggregation data. It should be noted that this solution needs to determine whether the i-th user meets all the conditions in the task conditions. This means that the aggregated cumulative value of all first encrypted data of the current user is used as preliminary aggregate data, and the preliminary aggregate data can be subsequently encrypted to make a judgment.
[0047] In the step of "performing compliance assessment on preliminary aggregated data and screening qualified screening data", the authority uses the private key of the ElGamal encryption algorithm to decrypt the preliminary aggregated data to obtain the cumulative multiplication value of each user's common attributes. If the cumulative multiplication value is 1, the first encrypted data and the second encrypted data of the current user are retained; otherwise, the first encrypted data and the second data of the current user are discarded to obtain the screening data.
[0048] In the step of "performing secondary aggregation on the screened data to obtain final aggregated data", the second encrypted data in the screened data is aggregated to obtain final aggregated data.
[0049] Correspondingly, the final aggregated data is , where m is the number of users in the screened data, 1≤i≤m, and N is the public key parameter of the Paillier cryptosystem. For the final aggregate data, is the second encrypted data of the i-th user. This formula means that, assuming that after the screening, the data of m users meet the requirements, then the private attributes of these m users are aggregated.
[0050] At the same time, the trusted third party sends the private key of the Paillier cryptographic system and the parameters of users who do not meet the task conditions to the data demander, who uses the parameters to complete the decryption operation of the final aggregation result. The specific decryption process is as follows: ; and calculate Get the aggregate value of a private attribute.
[0051] in is the encoded plaintext part, is a randomization factor used to protect plaintext information. Is the private key of the Paillier encryption algorithm, is the random parameter of the unselected user, satisfying , from the above, we can know that that is The formula for the decryption process means to assemble this private key , since this is 1 in the Paillier system, the equation for the decryption process can be expressed as .
[0052] Furthermore, when the normal attribute is a numeric attribute: In the step of "selectively aggregating the encrypted matrix sequence based on the task requirements to obtain preliminary aggregated data", the first encrypted data of each user in the encrypted matrix sequence that meets the task conditions is aggregated to obtain an encrypted sequence as preliminary aggregated data.
[0053] Correspondingly, the preliminary aggregated data obtained is ,in ,in The first encrypted data representing the jth common attribute of the i-th user, 1≤i≤n, H is the preliminary aggregate data, is the preliminary aggregation or data of the i-th user, s represents the position of the right endpoint, r represents the position of the left endpoint, p Represents large prime numbers.
[0054] However, assuming that the set U=100, then the common attribute data of each user must be divided into 100 parts and encoded and encrypted to obtain the first encrypted data. If the user who satisfies [30, 50] is to be selected, then when aggregating the encryption matrix sequence, the ciphertext at the position [30, 50] needs to be aggregated, which is and For aggregation, the corresponding value is s=30, r=50.
[0055] In the step of "performing compliance assessment on preliminary aggregated data and screening qualified screening data", the authority uses the private key of the ElGamal encryption algorithm to decrypt the preliminary aggregated data to obtain the cumulative multiplication value of each user's common attributes. If the cumulative multiplication value is not 1, the first encrypted data and the second encrypted data of the current user are retained; otherwise, the first encrypted data and the second data of the current user are discarded to obtain the screening data.
[0056] In the step of "performing secondary aggregation on the screened data to obtain final aggregated data", the second encrypted data in the screened data is aggregated to obtain final aggregated data.
[0057] The final aggregated data is , where m is the number of users in the filtered data, 1≤i≤m, is the second encrypted data of the i-th user, N represents the public key parameter of the Paillier cryptosystem, To finally aggregate the data, this formula means that, assuming that after the screening, the data of m users meet the requirements, then the private attributes of these m users are aggregated.
[0058] At the same time, the trusted third party sends the private key of the Paillier cryptographic system and the parameters of users who do not meet the task conditions to the data demander. The data demander uses the parameters to complete the decryption operation of the final aggregation result. The specific decryption process is the same as above.
[0059] In summary, the selective data aggregation method based on conditional classification coding provided by the present solution can complete the selective aggregation operation without leaking the common attributes and private attributes of the user. Figure 2 As shown: user data of users is collected and classified and encoded to obtain segmented user data, wherein the segmented user data includes encoded user data corresponding to common attributes and private attribute user data of private attributes; the encoded user data and private attribute user data are encrypted to obtain an encrypted matrix sequence; the aggregation server aggregates the encoded user data to obtain initial aggregated data and is judged by an authoritative organization. If the conditions are met, the aggregation server aggregates the private attribute user data to obtain final aggregated data, and the final aggregated data is sent to the data demander for statistical analysis.
[0060] Embodiment 2 This solution provides a selective data aggregation system based on conditional classification coding, which consists of a trusted third party, users, aggregation servers, authoritative institutions, and data demanders: The data demander sends a data analysis request to the aggregation server, where the data analysis request records the task conditions and task requirements of the current data analysis task; When a data analysis request is received, segmented user data of users in the same area are collected based on the data analysis request, and the segmented user data are encrypted to obtain an encryption matrix sequence, wherein the segmented user data includes encoded user data corresponding to common attributes and private attribute user data corresponding to private attributes, wherein the trusted third party provides the user with parameters of the encryption algorithm; The aggregation server selectively aggregates the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data, and sends the preliminary aggregated data to the authority; The authoritative structure conducts compliance assessment on the preliminary aggregated data and screens the qualified filtered data, and sends the filtered data to the aggregation server again; The aggregation server performs secondary aggregation on the filtered data to obtain final aggregated data, and sends the final aggregated data to the data demander.
[0061] When receiving a data analysis request from a data demander, the trusted third party initializes and sends the parameters of the ElGamal encryption algorithm and the Paillier algorithm to the user. The user encrypts the segmented user data based on the ElGamal encryption algorithm and the Paillier algorithm to obtain an encryption matrix sequence, and sends the encryption matrix sequence to the aggregation server. The aggregation server aggregates the encoded user data in the encryption matrix sequence to obtain initial aggregated data, and sends the initial aggregated data to the authority. The authority makes a judgment and obtains the filtered data. The filtered data is sent to the aggregation server again for aggregation. At this time, the aggregation server aggregates the private user data to obtain the final aggregated data, and sends the final aggregated data to the data demander. At the same time, the trusted third party also sends the parameters of the ElGamal encryption algorithm and the Paillier algorithm to the data demander.
[0062] Embodiment 3 This embodiment also provides an electronic device, referring to Figure 3 , including a memory 404 and a processor 402, wherein the memory 404 stores a computer program, and the processor 402 is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the selective data aggregation method based on conditional classification coding.
[0063] Specifically, the processor 402 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0064] The memory 404 may include a large-capacity memory 404 for data or instructions. The memory 404 may be used to store or cache various data files that need to be processed and / or used for communication, as well as possible computer program instructions executed by the processor 402.
[0065] The processor 402 implements any one of the selective data aggregation methods based on conditional classification coding in the above embodiments by reading and executing computer program instructions stored in the memory 404 .
[0066] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408 , wherein the transmission device 406 is connected to the processor 402 , and the input / output device 408 is connected to the processor 402 .
[0067] The transmission device 406 can be used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0068] The input / output device 408 is used to input or output information. In this embodiment, the input information may be a data analysis request, etc., and the output information may be final aggregated data, etc.
[0069] Optionally, in this embodiment, the processor 402 may be configured to perform the following steps through a computer program: Obtaining a data analysis request, wherein the data analysis request records the task conditions and task requirements of the current data analysis task; Based on the data analysis request, segmented user data of users in the same area are collected, and the segmented user data are encrypted to obtain an encrypted matrix sequence, wherein the segmented user data includes coded user data corresponding to common attributes and private attribute user data corresponding to private attributes. Selectively aggregate the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data; Conduct compliance assessment on the preliminary aggregated data and screen the eligible screening data; The filtered data are aggregated twice to obtain the final aggregated data.
[0070] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.
[0071] In general, various embodiments may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects of the invention may be implemented in hardware, while other aspects may be implemented in firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that, as non-limiting examples, the boxes, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0072] Embodiments of the present invention can be implemented by computer software, which is executable by a data processor of a mobile device, such as in a processor entity, or implemented by hardware, or implemented by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets and / or macros can be stored in any device readable data storage medium, and they include program instructions for performing specific tasks. Computer program products can include one or more computer executable components configured to perform embodiments when the program is running. One or more computer executable components can be at least one software code or a part thereof. In addition, at this point, it should be noted that any box of the logic flow in the figure can represent a program step, or interconnected logic circuits, boxes and functions, or a combination of program steps and logic circuits, boxes and functions. Software can be stored in physical media such as memory chips or storage blocks implemented in processors, magnetic media such as hard disks or floppy disks, and optical media such as, for example, DVDs and data variants thereof, CDs. Physical media are non-transient media.
[0073] Those skilled in the art should understand that the technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0074] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A selective data aggregation method based on conditional classification coding, characterized in that: The following steps are involved: Obtaining a data analysis request, wherein the data analysis request records the task conditions and task requirements of the current data analysis task; Based on the data analysis request, segmented user data of users in the same area are collected, and the segmented user data are encrypted to obtain an encrypted matrix sequence, wherein the segmented user data includes coded user data corresponding to common attributes and private attribute user data corresponding to private attributes. Selectively aggregate the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data; Conduct compliance assessment on the preliminary aggregated data and screen the eligible screening data; The filtered data are aggregated twice to obtain the final aggregated data.
2. The selective data aggregation method based on conditional classification coding according to claim 1 is characterized in that: The data analysis task is to obtain user data that meets specific task conditions and specific task requirements. Task conditions refer to the data attributes of user data that need to be aggregated. Task requirements stipulate that data attributes within task conditions are classified as common attributes or private attributes, where common attributes are divided into Boolean attributes and numerical attributes.
3. The selective data aggregation method based on conditional classification coding according to claim 1, characterized in that: Based on the data analysis request, common attribute user data and private attribute user data uploaded by the user are obtained, and the common attribute user data is encoded to obtain encoded user data, wherein the common attribute user data is user data corresponding to common attributes.
4. The selective data aggregation method based on conditional classification coding according to claim 2 is characterized in that: When the common attribute is selected as a Boolean attribute, the common attribute user data is compared with the task request. When the common attribute user data is consistent with the value of the common attribute in the task request, it is encoded as 1. When the common attribute user data is inconsistent with the value of the common data in the task request, it is encoded as any integer other than 1 and 0.
5. The selective data aggregation method based on conditional classification coding according to claim 2, characterized in that: When the common attribute is selected as a numerical attribute, construct the numerical full value of the common attribute corresponding to the current common attribute user data, compare the numerical value of the common attribute user data with the numerical full value, replace the position corresponding to the current common attribute user data in the numerical full value with 1, and replace the other positions with any integers other than 1 and 0 to obtain the encoded user data.
6. The selective data aggregation method based on conditional classification coding according to claim 1, characterized in that: All encoded user data are encrypted using the ElGamal encryption algorithm to obtain first encrypted data, all private attribute user data are encrypted using the Paillier algorithm to obtain second encrypted data, and the first encrypted data and second encrypted data of all users are integrated to obtain an encryption matrix sequence, wherein the first encrypted data and second encrypted data of each user are recorded in the encryption matrix sequence.
7. The selective data aggregation method based on conditional classification coding according to claim 6 is characterized in that: When the common attribute is a Boolean attribute, the first encrypted data of each user in the encryption matrix sequence is aggregated to obtain an encrypted sequence as preliminary aggregated data. The authority uses the private key of the ElGamal encryption algorithm to decrypt the preliminary aggregated data to obtain the cumulative multiplication value of the common attribute of each user. If the cumulative multiplication value is 1, the first encrypted data and the second encrypted data of the current user are retained. Otherwise, the first encrypted data and the second data of the current user are discarded to obtain the filtered data. The second encrypted data in the filtered data is aggregated to obtain the final aggregated data.
8. The selective data aggregation method based on conditional classification coding according to claim 6, characterized in that: When the common attribute is a numerical attribute, the first encrypted data of each user in the encrypted matrix sequence that meets the task conditions is aggregated to obtain an encrypted sequence as preliminary aggregated data. The authoritative organization uses the private key of the ElGamal encryption algorithm to decrypt the preliminary aggregated data to obtain the cumulative multiplication value of the common attribute of each user. If the cumulative multiplication value is not 1, the first encrypted data and the second encrypted data of the current user are retained. Otherwise, the first encrypted data and the second data of the current user are discarded to obtain the filtered data, and the second encrypted data in the filtered data is aggregated to obtain the final aggregated data.
9. A selective data aggregation system based on conditional classification coding, characterized in that: It consists of trusted third parties, users, aggregation servers, authoritative institutions and data demanders: The data demander sends a data analysis request to the aggregation server, where the data analysis request records the task conditions and task requirements of the current data analysis task; When a data analysis request is received, segmented user data of users in the same area are collected based on the data analysis request, and the segmented user data are encrypted to obtain an encryption matrix sequence, wherein the segmented user data includes encoded user data corresponding to common attributes and private attribute user data corresponding to private attributes, wherein the trusted third party provides the user with parameters of the encryption algorithm; The aggregation server selectively aggregates the encrypted matrix sequence based on the data analysis request to obtain preliminary aggregated data, and sends the preliminary aggregated data to the authority; The authoritative structure conducts compliance assessment on the preliminary aggregated data and screens the qualified filtered data, and sends the filtered data to the aggregation server again; The aggregation server performs secondary aggregation on the filtered data to obtain final aggregated data, and sends the final aggregated data to the data demander.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which includes a program code for controlling a process to execute a process, wherein the process includes the selective data aggregation method based on conditional classification coding according to any one of claims 1 to 8.
Citation Information
Patent Citations
A method, device, terminal and medium for data aggregation
CN118916241B
Information coding method and device, equipment and storage medium
CN115204111A
Archive management method and system based on big data
CN118551414A
Multi-attribute aggregation equivalent connection query method and device based on homomorphic encryption
CN118734331A
Crowdsourcing surveying and mapping architecture and method based on cloud side end
CN119128954A