A method, apparatus and system for processing data
By receiving data requests and generating noise data as target data, the problems of high computing resource consumption and poor processing flexibility under multiple data demanders are solved, and data security and processing efficiency are improved.
Patent Information
- Application Number
- CN202210499637.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-05-09
AI Technical Summary
In the case of multiple data demanders and small-granularity statistical requirements, existing technologies have the problems of high computing resource consumption, poor data processing flexibility and inability to meet group data processing requirements.
Receive requests from data demanders, obtain original data of the corresponding data range, and generate noise data as target data based on the data processing type, and send it to the data demanders to meet grouping requirements and improve data security and flexibility.
By generating noise data as target data, the group processing requirements of data demanders are met, thereby improving the data security of data providers and the processing efficiency and flexibility of data demanders.
Smart Images

Figure CN114817330B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security, and in particular to a method, device and system for processing data. Background Art
[0002] Usually, when a development platform provides access rights to other entities, it is often necessary to provide raw data to be processed so that other entities can perform statistical analysis based on the acquired raw data and receive the analysis results of other entities for further processing. However, in order to ensure the data security of the raw data, the general development platform often needs to process the raw data and provide the processed data to other entities.
[0003] At present, the main methods for general development platforms to process processed data are: 1) directly providing statistical data generated based on the original data according to the statistical needs of other entities; 2) adding noise to the original data to obtain processed data and provide processed data; however, when the number of other entities is large, the statistical needs are many, and the granularity of the needs is small (for example: group statistics), method 1) has the problem of consuming large computing resources and poor data processing flexibility; method 2) has the problem that the processed data obtained cannot meet the group data processing needs of other entities. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method and apparatus for processing data, which are capable of receiving a data acquisition request sent by a data demander, acquiring multiple pieces of raw data corresponding to the requested data range from a data source, generating corresponding target data for each piece of the raw data based on the data processing type and noise data determined for the data processing type, and sending the multiple target data to the data demander, so that the data demander can group the multiple target data based on the data processing type. This overcomes the problem of being unable to meet the data demander's needs for grouped data processing, improves the data security of the data provider, and enhances the flexibility and efficiency of the data demander in processing data.
[0005] To achieve the above-mentioned purpose, according to one aspect of an embodiment of the present invention, a method for processing data is provided, which is characterized by comprising: receiving a data acquisition request sent by a data demander; the data acquisition request indicates one or more data ranges and a data processing type; according to the one or more data ranges indicated by the data acquisition request, obtaining multiple original data belonging to one or more data ranges from a data source; according to the data processing type and the noise data corresponding to the data processing type determined for the multiple original data, generating corresponding target data for each of the original data, wherein the noise data corresponding to the data processing type enables the multiple target data to meet the grouping requirements; sending the multiple target data to the data demander so that the data demander processes the multiple target data based on the data processing type.
[0006] Optionally, the method for processing data includes: the data acquisition request further indicates the number of data groups required by the data demander and the first data volume threshold of each data group; after receiving the data acquisition request sent by the data demander, further including: when it is determined that the first data volume threshold of any data group is less than the second data volume threshold, sending information indicating a request abnormality to the data demander, wherein the second data volume threshold indicates the minimum data volume required for noise data of any data processing type to meet the grouping requirements.
[0007] Optionally, the method for processing data further includes: determining the number of the multiple original data when it is determined that the first data volume threshold of any of the data groups is not less than the second data volume threshold; sending information indicating a request abnormality to the data demander when it is determined that the number of the multiple original data is less than or equal to the second data volume threshold; and executing the step of generating corresponding target data for each of the original data when it is determined that the number of the multiple original data is greater than the second data volume threshold.
[0008] Optionally, the method for processing data, after obtaining multiple pieces of original data belonging to one or more data ranges from the data source, further includes: calculating a mean corresponding to the data processing type based on the multiple original data; generating noise data of the data processing type based on the mean corresponding to the data processing type, a first data volume threshold of data of any data group further indicated by the data acquisition request, and a processing error.
[0009] Optionally, in the method for processing data, when the data processing type indicates calculating a sum or a mean, calculating the mean corresponding to the data processing type includes: calculating the mean of the multiple original data.
[0010] Optionally, the method for processing data further includes: for the case where the data processing type is variance or standard deviation, calculating the mean corresponding to the data processing type includes: for each piece of the original data, performing a square operation on the original data to generate first data corresponding to the original data; and calculating the mean of multiple first data.
[0011] Optionally, generating noise data of the data processing type includes: calculating a privacy budget parameter based on a mean corresponding to the data processing type, a first data volume threshold of data of any data group further indicated by the data acquisition request, and a processing error; and inputting the privacy budget parameter into a Laplace distribution function to obtain the noise data.
[0012] To achieve the above-mentioned purpose, according to the second aspect of an embodiment of the present invention, a method for processing data is provided, which is characterized in that it includes: sending a data acquisition request to a data provider, wherein the data acquisition request indicates one or more data ranges and a data processing type; upon receiving multiple target data sent by the data provider, dividing the multiple target data into one or more data groups, and processing the target data in the data groups, wherein the target data is formed based on the original data of the data provider and noise data corresponding to the data processing type, and the noise data corresponding to the data processing type enables the multiple target data to meet the grouping requirements.
[0013] Optionally, the method for processing data further includes: setting the number of required data groups and the first data volume threshold of each data group, and dividing the multiple target data into one or more data groups, including: dividing the multiple target data according to the number of data groups and the first data volume threshold of each data group when the number of the multiple target data is not less than the sum of the first data volume threshold of each data group.
[0014] Optionally, the method for processing data further includes: when the number of the multiple target data is less than the sum of the first data volume thresholds of each of the data groups, obtaining one or more target data groups whose first data volume threshold is less than the number of the multiple target data, and processing the target data in the target data groups.
[0015] Optionally, the processing of the target data in the data group includes: for the case where the data processing type is variance or standard deviation, determining for the data group first target data corresponding to the data processing type of sum or mean provided by the data provider, and second target data corresponding to the data processing type of variance or standard deviation; and calculating the variance or standard deviation of the data group based on the calculation relationship between the square of the mean of the first target data and the mean of the second target data.
[0016] To achieve the above-mentioned purpose, according to a third aspect of an embodiment of the present invention, there is provided a device for processing data, characterized in that it is applied to a data provider and comprises: a data acquisition module, a data processing module and a data sending module; wherein,
[0017] The data acquisition module is configured to receive a data acquisition request sent by a data requester; the data acquisition request indicates one or more data ranges and a data processing type; and acquire, from a data source, a plurality of pieces of original data within the one or more data ranges indicated by the data acquisition request;
[0018] The data processing module is configured to generate corresponding target data for each piece of the original data according to the data processing type and the noise data corresponding to the data processing type determined for the plurality of original data, wherein the noise data corresponding to the data processing type enables the plurality of target data to meet a grouping requirement;
[0019] The data sending module is used to send the multiple target data to the data demander, so that the data demander processes the multiple target data based on the data processing type.
[0020] Optionally, the device for processing data includes: the data acquisition request further indicates the number of data groups required by the data demander and the first data volume threshold of each data group; after receiving the data acquisition request sent by the data demander, further includes: when it is determined that the first data volume threshold of any data group is less than the second data volume threshold, sending information indicating that the request is abnormal to the data demander, wherein the second data volume threshold indicates the minimum data volume required for noise data of any data processing type to meet the grouping requirements.
[0021] Optionally, the data processing device is further used to determine the number of the multiple original data when it is determined that the first data volume threshold of any of the data groups is not less than the second data volume threshold; when it is determined that the number of the multiple original data is less than or equal to the second data volume threshold, send information indicating a request abnormality to the data demander; when it is determined that the number of the multiple original data is greater than the second data volume threshold, execute the step of generating corresponding target data for each of the original data.
[0022] Optionally, the data processing device, after obtaining multiple original data belonging to one or more data ranges from the data source, further includes: calculating the mean corresponding to the data processing type based on the multiple original data; generating noise data of the data processing type based on the mean corresponding to the data processing type, the first data volume threshold of the data of any data group further indicated by the data acquisition request, and the processing error.
[0023] Optionally, the data processing device, in response to the data processing type indicating a sum or a mean, calculates the mean corresponding to the data processing type, including: calculating the mean of the multiple original data.
[0024] Optionally, the device for processing data is further used to calculate the mean corresponding to the data processing type when the data processing type is variance or standard deviation, including: performing a square operation on each piece of the original data to generate first data corresponding to the original data; and calculating the mean of multiple first data.
[0025] Optionally, the data processing device is used to generate noise data of the data processing type, including: calculating a privacy budget parameter based on the mean corresponding to the data processing type, a first data volume threshold of data of any data group further indicated by the data acquisition request, and a processing error; and inputting the privacy budget parameter into a Laplace distribution function to obtain the noise data.
[0026] To achieve the above-mentioned purpose, according to a fourth aspect of an embodiment of the present invention, a device for processing data is provided, characterized in that it is applied to a data demand side and comprises: a data request module and a data processing module; wherein,
[0027] The data request module is configured to send a data acquisition request to a data provider, wherein the data acquisition request indicates one or more data ranges and data processing types;
[0028] The data processing module is used to divide the multiple target data sent by the data provider into one or more data groups and process the target data in the data groups, wherein the target data is formed based on the original data of the data provider and noise data corresponding to the data processing type, and the noise data corresponding to the data processing type enables the multiple target data to meet the grouping requirements.
[0029] Optionally, the data processing device is also used to set the number of required data groups and the first data volume threshold of each data group, and the dividing the multiple target data into one or more data groups includes: when the number of the multiple target data is not less than the sum of the first data volume threshold of each data group, dividing the multiple target data according to the number of data groups and the first data volume threshold of each data group.
[0030] Optionally, the data processing device is further used to obtain one or more target data groups whose first data volume threshold is smaller than the number of the multiple target data when the number of the multiple target data is smaller than the sum of the first data volume thresholds of each of the data groups, and process the target data in the target data groups.
[0031] Optionally, the data processing device is used to process the target data in the data group, including: for the case where the data processing type is variance or standard deviation, determining for the data group the first target data corresponding to the data processing type of sum or mean provided by the data provider, and the second target data corresponding to the data processing type of variance or standard deviation; based on the calculation relationship between the square of the mean of the first target data and the mean of the second target data, calculating the variance or standard deviation of the data group.
[0032] To achieve the above objectives, according to the fifth aspect of an embodiment of the present invention, a system for processing data is provided, comprising a data provider having the device for processing data according to the third aspect and a data demander having the device for processing data according to the fourth aspect.
[0033] To achieve the above-mentioned purpose, according to the sixth aspect of an embodiment of the present invention, there is provided an electronic device for processing data, characterized in that it includes: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement any of the methods described above for processing data.
[0034] To achieve the above-mentioned purpose, according to the seventh aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored, characterized in that when the program is executed by a processor, any method as described in the above-mentioned method for processing data is implemented.
[0035] One embodiment of the above invention has the following advantages or beneficial effects: it is capable of receiving a data acquisition request sent by a data demander and acquiring multiple pieces of raw data corresponding to the requested data range from a data source; generating corresponding target data for each piece of the raw data based on the data processing type and noise data determined for the data processing type, and sending the multiple target data to the data demander, so that the data demander can group the multiple target data based on the data processing type. This overcomes the problem of being unable to meet the data demander's needs for grouped data processing, improves the data security of the data provider, and enhances the flexibility and efficiency of the data demander in processing data.
[0036] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings are provided for a better understanding of the present invention and are not intended to limit the present invention.
[0038] Figure 1 This is a flowchart of a method for processing data applied to a data provider provided by an embodiment of the present invention;
[0039] Figure 2 This is a flowchart of a method for processing data applied to a data demand side provided by an embodiment of the present invention;
[0040] Figure 3 This is a structural diagram of a data providing end of a data processing device provided by an embodiment of the present invention;
[0041] Figure 4 This is a structural diagram of a data demand side of a data processing device provided by one embodiment of the present invention;
[0042] Figure 5 This is a schematic diagram of the structure of a data processing system provided by one embodiment of the present invention;
[0043] Figure 6 is an exemplary system architecture diagram in which embodiments of the present invention may be applied;
[0044] Figure 7 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, in which various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0046] like Figure 1 As shown, an embodiment of the present invention provides a method for processing data, which is applied to a data provider. The method may include the following steps:
[0047] Step S101: Receive a data acquisition request sent by a data demander; the data acquisition request indicates one or more data ranges and a data processing type; and acquire multiple pieces of original data belonging to one or more data ranges from a data source according to the one or more data ranges indicated by the data acquisition request.
[0048] Specifically, for application scenarios in which the software open platform provides an access interface, one or more other entities can develop business modules based on the access interface provided by the software open platform. Usually, other entities, as data demanders, need to obtain data from the software open platform (i.e., data providers); further, other entities can provide data to the software open platform through the business modules they develop, so that the software open platform can further process the data from other entities.
[0049] Furthermore, a data acquisition request is received, and the data acquisition request may indicate a data range and a data processing type; wherein the data range may be a data range of the data to be processed, and the data range is associated with a query condition for obtaining the data, for example, the data range of the data to be processed is: the number of abc within one month; the data processing types include: sum (or mean), variance (or standard deviation); that is, the data demander may further calculate the corresponding sum (or mean) and variance (or standard deviation) of the statistical data based on the acquired data.
[0050] Furthermore, the data provider can obtain multiple pieces of raw data within one or more data ranges indicated by the data acquisition request (e.g., the data range indicated by the data query condition) from an internal data source (e.g., a database) within the data provider. Raw data refers to internal data that has not been processed with noise. Furthermore, target data is generated based on the raw data and provided to the data requester to improve data security.
[0051] Preferably, before obtaining the original data based on the data acquisition request of the data demander, the validity of the data acquisition request and the original data is judged; in an embodiment of the present invention, the data demander can calculate the corresponding sum (or mean) and variance (or standard deviation) based on the data provided by the data provider, and can perform group calculations on the data received from the data provider, thereby improving the flexibility of data processing.
[0052] Specifically, the number of data groups required by the data demander and the first data volume threshold of each data group are parsed from the data acquisition request, that is, the data acquisition request further indicates the number of data groups required by the data demander and the first data volume threshold of each data group; wherein the first data volume threshold is the amount of data contained in the data group; after receiving the data acquisition request sent by the data demander, the data provider further includes: when it is determined that the first data volume threshold of any data group is less than the second data volume threshold, sending information indicating that the request is abnormal to the data demander, wherein the second data volume threshold indicates the minimum data volume required for noise data of any data processing type to meet the grouping requirements. Among them, the second data volume threshold is a set quantity threshold, which is determined according to the application scenario and the distribution of data types. For example, the second data volume threshold is set to 30. It can be understood that when the data volume of any data group reaches the second data volume threshold, the result of adding noise data to the original data can make the calculation result within the set error range (that is, the minimum data volume that meets the grouping requirements), that is, the second data volume threshold indicates the minimum data volume required for noise data of any data processing type to meet the grouping requirements; therefore, when it is determined that the first data volume threshold of any data group (for example, 20) is less than the second data volume threshold (for example, 30), information indicating that the request is abnormal is sent to the data demander. By sending the abnormal information of the request, the effectiveness and security of the data provided by the data provider are improved.
[0053] Furthermore, when it is determined that the first data volume threshold of any of the data groups is not less than the second data volume threshold, the number of the plurality of original data is determined; when it is determined that the number of the plurality of original data is less than or equal to the second data volume threshold, information indicating a request exception is sent to the data demander; when it is determined that the number of the plurality of original data is greater than the second data volume threshold, the step of generating corresponding target data for each of the original data is performed. Specifically, when the first data volume threshold satisfies the condition of being not less than the second data volume threshold, the step of obtaining the original data from the data source is performed, and the number of the original data is determined. If the number of the original data obtained is less than the second data volume threshold (for example, 30), the step of generating target data based on the original data and providing the data to the data demander is not performed, and information indicating a request exception is sent to the data demander. Otherwise, the step of generating corresponding target data for each of the original data is performed. The description of generating corresponding target data based on the original data is consistent with that of step S102 and will not be repeated here. The effectiveness and security of the data provided by the data provider are improved by sending the abnormal information of the request.
[0054] Step S102: Generate corresponding target data for each piece of the original data according to the data processing type and the noise data corresponding to the data processing type determined for the multiple pieces of original data, wherein the noise data corresponding to the data processing type enables the multiple pieces of target data to meet grouping requirements.
[0055] Specifically, after obtaining multiple pieces of raw data belonging to one or more data ranges from the data source, the method further includes: calculating a mean value corresponding to the data processing type based on the multiple pieces of raw data; and generating noise data of the data processing type based on the mean value corresponding to the data processing type, a first data volume threshold of any data group further indicated in the data acquisition request, and a processing error, and performing noise processing using local differential privacy.
[0056] Specifically, the method for generating noise data for different data processing types is:
[0057] 1) When the data processing type is sum (or mean):
[0058] For example, the multiple pieces of original data obtained are represented as data sequence R1; the corresponding target data generated are represented as R2;
[0059] The calculated mean of R1 is expressed as avg(R1); that is, for the case where the data processing type indicates finding the sum or the mean, the calculating of the mean corresponding to the data processing type includes: calculating the mean of the multiple original data.
[0060] Combined with the first data volume threshold (e.g., n) and processing error (e.g., a, which can be set to 0.1%-20% according to the application scenario) of any data group further indicated in the data acquisition request; further, the privacy budget parameter ε is determined as:
[0061]
[0062] Furthermore, the privacy budget parameter is input into a Laplace distribution function to obtain the noise data L(0, ε); wherein L is a Laplace distribution function. That is, generating the noise data for the data processing type includes: calculating the privacy budget parameter based on the mean corresponding to the data processing type (i.e., summation or mean), the first data volume threshold of the data of any data group further indicated by the data acquisition request, and the processing error; and inputting the privacy budget parameter into a Laplace distribution function to obtain the noise data. That is, generating the noise data L(0, ε) for the data processing type based on the mean corresponding to the data processing type, the first data volume threshold of the data of any data group further indicated by the data acquisition request, and the processing error.
[0063] Furthermore, assuming that a piece of original data R1 is represented by X, and the data corresponding to the generated target data R2 is represented by Y, then Y=X+L(0,ε).
[0064] The data demander can perform group summing, averaging and other operations based on multiple target data provided to the data demander. The error value between the obtained statistical result and the result based on the summation and averaging of the original data is less than the error value set by the data demander, thereby achieving data security by adding noise data to the original data, and enabling the data demander to obtain a statistical value (sum or mean) that meets the error conditions based on the processed data.
[0065] The following describes the process of determining noise data and the usability proof of the sum (or mean) of the target data of the present invention: For the Laplace distribution model L(u,b), the mean of the distribution is u and the variance is 2b 2 ; The noise data constructed in the present invention is L(0,b) generated based on the Laplace distribution model; wherein b=f / c; wherein f is the difference between the maximum value and the minimum value in the R1 sequence.
[0066] For a data group in R2, sum(Y)=sum(X)+sum(L(0,b)). To ensure the availability of the sum, sum(X) and sum(Y) need to be infinitely close, that is, sum(L(0,b)) is infinitely close to 0. When the data volume n of a group (that is, the first data volume threshold) meets the requirement (for example, n is greater than or equal to 30), according to the central limit theorem, sum(L(0,b)) obeys the normal distribution N(0,2nb 2 ). According to the normal distribution y=N(u,δ 2 ), then the probability of |yu|>3δ is less than a, for example, a=0.3%, that is, The probability is less than 0.3%. For example, if the error value indicated by the data demander is a, then sum(L(0,b))=asum(X)=a*n*avg(X). Let You can get Where b is the privacy budget parameter of the present invention. Therefore, That is, the probability that |sum(Y)-sum(X)|≤αsum(X) is greater than 99.7%. Thus, the method of the embodiment of the present invention can ensure the availability of summation with an error of a and the availability of averaging. The target data R2 generated by the method of the embodiment of the present invention can also achieve the availability of group summation or averaging.
[0067] 2) For the case where the data processing type is variance or standard deviation:
[0068] For example, the multiple pieces of original data obtained are data sequence R1; the corresponding target data generated is R3;
[0069] Calculate the mean avg(R1') of the R1' column corresponding to R1; wherein, the result obtained by performing a square operation on the original data contained in R1 is the data of the R1' column (i.e., the first data); the mean of R1' calculated through the data of R1' is expressed as avg(R1') (i.e., the mean of multiple first data); that is, for the case where the data processing type is variance or standard deviation, calculating the mean corresponding to the data processing type includes: for each piece of the original data, performing a square operation on the original data to generate first data corresponding to the original data; and calculating the mean of multiple first data.
[0070] Combined with the first data volume threshold (e.g., n) of any data group further indicated in the data acquisition request and the processing error a, the privacy budget parameter ε is determined as:
[0071]
[0072] Furthermore, the privacy budget parameter is input into a Laplace distribution function to obtain the noise data L(0,ε); wherein L is a Laplace distribution function. That is, generating the noise data of the data processing type includes: calculating the privacy budget parameter based on the mean corresponding to the data processing type (i.e., calculating the variance or standard deviation), the first data volume threshold of the data of any data group further indicated by the data acquisition request, and the processing error; and inputting the privacy budget parameter into the Laplace distribution function to obtain the noise data. Further, the privacy budget parameter is input into a Laplace distribution function to obtain the noise data L(0,ε); wherein L is a Laplace distribution function. That is, generating the noise data of the data processing type includes: calculating the privacy budget parameter based on the mean corresponding to the data processing type (i.e., calculating the variance or standard deviation), the first data volume threshold of the data of any data group further indicated by the data acquisition request, and the processing error; and inputting the privacy budget parameter into the Laplace distribution function to obtain the noise data. That is, noise data of the data processing type is generated based on the mean value corresponding to the data processing type, the first data amount threshold of the data of any data group further indicated by the data acquisition request, and the processing error.
[0073] Furthermore, for a piece of original data R1 denoted as X, the corresponding data of the generated target data R3 is denoted as Y, then Y=X+L(0,ε).
[0074] Based on this processing, the multiple target data R3 provided to the data demander can enable the data demander to perform grouping operations such as variance and standard deviation based on R3 and R2. The error value between the obtained statistical results and the results based on the variance and standard deviation of the original data is less than the error value set by the data demander, thereby achieving data security by adding noise data to the original data, and enabling the data demander to obtain statistical values that meet the error conditions based on the processed data.
[0075] Similarly, the usability proof of the present invention for determining the variance of noise data for a group is described as follows: Since |avg(R1)-avg(R2)|≤a*avg(R1), the following expression can be obtained: |avg(R1) 2 -avg(R2) 2 |≤a(a+2)*avg(R1) 2 ; and since |avg(R3)-avg(R1 2 )|≤a*avg(R1 2 ), according to the error summation formula, the data demander can calculate the variance variance(R1)=avg(R3)-avg(R2) 2Compared with the variance value variance(R1)=avg(R1) calculated directly from the original data 2 )-avg(R1) 2 The error is a(a+3).
[0076] Step S103: sending the plurality of target data to a data demander, so that the data demander processes the plurality of target data based on the data processing type.
[0077] Specifically, after determining one or more corresponding target data according to the data processing type indicated by the request for obtaining data, the target data is sent to the data demander.
[0078] Exemplarily, a plurality of original data are represented as R1;
[0079] When the data processing type is summation or average, the first target data R2 corresponding to R1 is sent to the data demander.
[0080] In the case where the data processing type is to calculate the variance or standard deviation, the first target data R2 corresponding to R1 and the second target data R3 corresponding to R1 are sent to the data demander.
[0081] Furthermore, the plurality of target data are sent to the data demander, so that the data demander processes the plurality of target data based on the data processing type. The description of the data demander processing the plurality of target data based on the data processing type is consistent with the description of steps S201-S202 and will not be repeated here.
[0082] Preferably, the data demander receives processing results (such as statistical analysis, etc.) on multiple target data, and further processes the processing results and related data in combination.
[0083] like Figure 2 As shown, an embodiment of the present invention provides a method for processing data, which is applied to a data demander. The method may include the following steps:
[0084] Step S201: Sending a data acquisition request to a data provider, wherein the data acquisition request indicates one or more data ranges and data processing types.
[0085] Specifically, the descriptions about the data acquisition request, the data acquisition request indicating one or more data ranges, and the data processing type are consistent with the description of step S101 and are not repeated here.
[0086] Step S202: Upon receiving multiple pieces of target data sent by the data provider, the multiple pieces of target data are divided into one or more data groups, and the target data in the data groups are processed, wherein the target data is formed based on the original data of the data provider and noise data corresponding to the data processing type, and the noise data corresponding to the data processing type enables the multiple pieces of target data to meet the grouping requirements.
[0087] Specifically, the description of generating target data is consistent with the description of step S102 and will not be repeated here.
[0088] Furthermore, receiving target data sent by a data provider further includes: setting the number of required data groups and a first data volume threshold for each of the data groups, and dividing the multiple target data into one or more data groups, including: dividing the multiple target data according to the number of data groups and the first data volume threshold for each of the data groups, when the number of the multiple target data is not less than the sum of the first data volume thresholds of the respective data groups. For example: the number of data groups is any value between 1 and 10, the first data volume threshold is 40 (greater than 30, for example, the second data volume threshold is 30, and the second data volume threshold indicates the minimum data volume required for noise data of any data processing type to meet the grouping requirements), dividing the received target data into one or more data groups, and calculating the statistical data corresponding to the data processing type for each data group, for example, sum (or mean) and variance (or standard deviation); wherein the target data can be grouped based on multiple dimensions, for example, based on the grouping dimensions included in the application scenario, dividing the target data into data groups based on age, region, gender, etc.
[0089] Furthermore, in the case where the number of the multiple target data is less than the sum of the first data volume thresholds of each of the data groups, one or more target data groups whose first data volume thresholds are less than the number of the multiple target data are obtained, and the target data in the target data groups are processed. For example: the number of target data obtained is 40, and the sum of the first data volume thresholds of the two data groups is 80 (35+45); that is, in the case where the number of the multiple target data is less than the sum of the first data volume thresholds of each of the data groups, a target data group with a first data volume threshold of 35 is obtained, and the target data in the target data group is processed. It can be understood that the above-mentioned first data volume threshold and the number of original data are only examples. In actual big data application scenarios, the number of target data is usually of a larger order of magnitude; by calculating statistical values in groups, the flexibility of data demanders in obtaining statistical values for different dimensions is improved, and the availability and security of data are improved.
[0090] Furthermore, the methods for processing target data according to different data processing types are:
[0091] Still taking the target data sequences R2 and R3 generated based on the original data sequence R1 described in step S102 as an example:
[0092] 1) For the case where the data processing type is summation or mean, for example: the target data sequence is R2, divided into 1...N data groups, for example: the target data contained in data group 1 is (x1...x100), then the sum of the target data contained in data group 1 is SUM(x1...x100) or the mean AVG=SUM(x1...x100) / 100
[0093] 2) For the case where the data processing type is variance or standard deviation, for example: the target data series is R2 and R3, divided into 1…N data groups. For example: the target data R2 contained in data group 1 is (x1…x100) and R3 is (y1…y100). Then the formula for calculating the variance of the target data contained in data group 1 can be: avg(y1…y100)-avg(x1…x100) 2 , further, obtaining the corresponding standard deviation based on the calculated variance square root. That is, the processing of the target data in the data group includes: for the case where the data processing type is variance or standard deviation, determining for the data group the first target data corresponding to the data processing type of sum or mean provided by the data provider (for example, x1…x100 included in R2), and the second target data corresponding to the data processing type of variance or standard deviation (for example, y1…y100 included in R3); based on the square of the mean of the first target data (i.e., avg(x1…x100) 2 ), and the calculation relationship between the mean of the second target data (i.e., avg(y1…y100)) (i.e., avg(y1…y100)-avg(x1…x100) 2 ), calculate the variance or standard deviation of the data set.
[0094] like Figure 3 As shown, an embodiment of the present invention provides a data processing device 300, which is applied to a data provider, including: a data acquisition module 301, a data processing module 302 and a data sending module 303; wherein,
[0095] The data acquisition module 301 is configured to receive a data acquisition request from a data requester; the data acquisition request indicates one or more data ranges and a data processing type; and acquire, from a data source, a plurality of pieces of original data within the one or more data ranges indicated by the data acquisition request;
[0096] The data processing module 302 is configured to generate corresponding target data for each piece of the original data according to the data processing type and the noise data corresponding to the data processing type determined for the plurality of original data, wherein the noise data corresponding to the data processing type enables the plurality of target data to meet a grouping requirement;
[0097] The data sending module 303 is configured to send the plurality of target data to a data demander, so that the data demander processes the plurality of target data based on the data processing type.
[0098] like Figure 4 As shown, an embodiment of the present invention provides a data processing device 400, which is applied to a data demand side and includes: a data request module 401 and a data processing module 402; wherein,
[0099] The data request module 401 is configured to send a data acquisition request to a data provider, wherein the data acquisition request indicates one or more data ranges and data processing types;
[0100] The data processing module 402 is used to divide the multiple target data sent by the data provider into one or more data groups and process the target data in the data groups, wherein the target data is formed based on the original data of the data provider and noise data corresponding to the data processing type, and the noise data corresponding to the data processing type enables the multiple target data to meet the grouping requirements.
[0101] like Figure 5 As shown, an embodiment of the present invention provides a system 500 for processing data, including: a data provider 502 having a device for processing data and a data demander 501 having a device for processing data.
[0102] An embodiment of the present invention also provides an electronic device for processing data, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in any of the above embodiments.
[0103] An embodiment of the present invention further provides a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, the method provided in any of the above embodiments is implemented.
[0104] Figure 6 An exemplary system architecture 600 is shown to which the method for processing data or the apparatus for processing data according to the embodiments of the present invention can be applied.
[0105] like Figure 6 As shown, system architecture 600 may include terminal devices 601, 602, 603, a network 604, and a server 605. Network 604 is used to provide a medium for communication links between terminal devices 601, 602, 603 and server 605. Network 604 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0106] Users can use terminal devices 601, 602, 603 to interact with server 605 via network 604 to receive or send messages, etc. Various client applications can be installed on terminal devices 601, 602, 603, such as various application modules developed based on open platforms.
[0107] The terminal devices 601 , 602 , and 603 may be various electronic devices having a display screen and supporting various client applications, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0108] The server 605 may be a server that provides various services, such as a background management server that provides support for client applications used by users using the terminal devices 601, 602, and 603. The background management server may process the received data acquisition request and feed back the processed target data to the terminal device.
[0109] It should be noted that the data demand side of the method for processing data provided in the embodiment of the present invention is generally executed by the terminal devices 601, 602, and 603, and the data provider side of the method for processing data provided in the embodiment of the present invention is generally executed by the server 605. Accordingly, the device for processing data on the data demand side is generally set in the terminal devices 601, 602, and 603; the device for processing data on the data provider side is generally set in the server 605.
[0110] It should be understood that Figure 6 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0111] Reference below Figure 7 , which shows a schematic structural diagram of a computer system 700 of a terminal device suitable for implementing an embodiment of the present invention. Figure 7 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0112] like Figure 7As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the system 700 are also stored in the RAM 703. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0113] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, and the like; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 708 including a hard disk; and a communication section 709 including a network interface card such as a LAN card or a modem. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 710 as needed, so that computer programs read therefrom can be installed into the storage section 708 as needed.
[0114] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from a removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above-mentioned functions defined in the system of the present invention are executed.
[0115] It should be noted that the computer-readable medium described in the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0116] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0117] The modules and / or units described in the embodiments of the present invention may be implemented in software or hardware. The modules and / or units described may also be provided in a processor. For example, they may be described as comprising a processor including a data acquisition module, a data processing module, and a data sending module. The names of these modules do not, in some cases, limit the modules themselves. For example, a data sending module may also be described as a module for sending a plurality of target data to a data requester.
[0118] As another aspect, the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiments, or may exist independently without being incorporated into the device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the device, the device includes:
[0119] For the data provider, a data acquisition request sent by a data demander is received; the data acquisition request indicates one or more data ranges and a data processing type; according to the one or more data ranges indicated by the data acquisition request, multiple original data belonging to one or more data ranges are acquired from a data source; according to the data processing type and the noise data corresponding to the data processing type determined for the multiple original data, corresponding target data is generated for each of the original data, wherein the noise data corresponding to the data processing type enables the multiple target data to meet the grouping requirements; the multiple target data are sent to the data demander so that the data demander processes the multiple target data based on the data processing type.
[0120] For the data demand side, a data acquisition request is sent to the data provider, wherein the data acquisition request indicates one or more data ranges and data processing types; when multiple target data are received from the data provider, the multiple target data are divided into one or more data groups, and the target data in the data groups are processed, wherein the target data is formed based on the original data of the data provider and noise data corresponding to the data processing type, and the noise data corresponding to the data processing type enables the multiple target data to meet the grouping requirements.
[0121] Embodiments of the present invention can receive data acquisition requests from a data demander, retrieve multiple pieces of raw data corresponding to the requested data range from a data source, generate corresponding target data for each piece of raw data based on the data processing type and noise data determined for that data processing type, and send the multiple target data to the data demander, enabling the data demander to group and process the multiple target data based on the data processing type. This overcomes the problem of being unable to meet the data demander's needs for grouped data processing, improves the data security of the data provider, and enhances the flexibility and efficiency of the data demander in processing data.
[0122] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for processing data, characterized in that include: Receiving a data acquisition request sent by a data demander; the data acquisition request indicates one or more data ranges and data processing types; Acquire, from a data source, a plurality of pieces of original data that fall within one or more data ranges indicated by the data acquisition request; generating corresponding target data for each piece of the original data according to the data processing type and the noise data corresponding to the data processing type determined for the plurality of original data, wherein the noise data corresponding to the data processing type enables the plurality of target data to meet grouping requirements; sending the plurality of target data to a data demander, so that the data demander processes the plurality of target data based on the data processing type; After acquiring a plurality of pieces of original data belonging to one or more of the data ranges from the data source, the method further includes: calculating a mean value corresponding to the data processing type based on the plurality of original data; generating noise data of the data processing type based on the mean value corresponding to the data processing type, a first data volume threshold of data of any data group further indicated by the data acquisition request, and a processing error; Generating the noise data of the data processing type includes: calculating a privacy budget parameter based on a mean corresponding to the data processing type, a first data volume threshold of any data group further indicated by the data acquisition request, and a processing error; and inputting the privacy budget parameter into a Laplace distribution function to obtain the noise data; wherein the formula for calculating the privacy budget parameter is as follows: Where ε represents the privacy budget parameter, n represents the first data volume threshold, a represents the processing error, avg(R1) represents the mean corresponding to the data processing type, and R1 represents the data sequence of multiple original data.
2. The method according to claim 1, characterized in that The data acquisition request further indicates the number of data groups required by the data demander and a first data volume threshold for each of the data groups; After receiving the data acquisition request sent by the data demander, the method further includes: When it is determined that the first data volume threshold of any of the data groups is less than the second data volume threshold, information indicating a request abnormality is sent to the data demander, wherein the second data volume threshold indicates the minimum data volume required for noise data of any data processing type to meet the grouping requirements.
3. The method according to claim 2, characterized in that Further including: When it is determined that the first data volume threshold of any of the data groups is not less than the second data volume threshold, determining the number of the plurality of pieces of original data; If it is determined that the number of the plurality of original data pieces is less than or equal to the second data amount threshold, sending information indicating that the request is abnormal to the data requester; When it is determined that the number of the plurality of original data pieces is greater than the second data amount threshold, the step of generating corresponding target data for each piece of the original data is performed.
4. The method according to claim 1, wherein In case of indicating the sum or average of the data processing type, The calculating the mean corresponding to the data processing type includes: Calculate the mean of the plurality of original data.
5. The method according to claim 1, wherein Further including: For the case where the data processing type is variance or standard deviation, The calculating the mean corresponding to the data processing type includes: performing a square operation on each piece of the original data to generate first data corresponding to the original data; Calculate the mean of the plurality of first data.
6. A method for processing data, characterized in that include: Sending a data acquisition request to a data provider, wherein the data acquisition request indicates one or more data ranges and data processing types; Upon receiving a plurality of target data pieces sent by the data provider, dividing the plurality of target data pieces into one or more data groups, and processing the target data within the data groups, wherein the target data pieces are formed based on original data from the data provider and noise data corresponding to the data processing type, and the noise data corresponding to the data processing type enables the plurality of target data pieces to meet grouping requirements; After the data provider obtains multiple pieces of raw data belonging to one or more of the data ranges from the data source, the method further includes: calculating a mean corresponding to the data processing type based on the multiple pieces of raw data; generating noise data of the data processing type based on the mean corresponding to the data processing type, a first data volume threshold of data of any data group further indicated by the data acquisition request, and a processing error; generating the noise data of the data processing type includes: calculating a privacy budget parameter based on the mean corresponding to the data processing type, a first data volume threshold of data of any data group further indicated by the data acquisition request, and a processing error; and inputting the privacy budget parameter into a Laplace distribution function to obtain the noise data; The formula for calculating the privacy budget parameter is as follows: Where ε represents the privacy budget parameter, n represents the first data volume threshold, a represents the processing error, avg(R1) represents the mean corresponding to the data processing type, and R1 represents the data sequence of multiple original data.
7. The method for processing data according to claim 6, characterized in that: Also includes: The number of required data groups and a first data volume threshold for each data group are set. The dividing the plurality of target data into one or more data groups includes: When the number of the plurality of target data is not less than the sum of the first data amount thresholds of the respective data groups, The plurality of target data are divided according to the number of the data groups and a first data volume threshold of each of the data groups.
8. The method for processing data according to claim 7, characterized in that: Further including: When the number of the plurality of target data is less than the sum of the first data amount thresholds of the respective data groups, One or more target data groups whose first data volume threshold is smaller than the number of the plurality of target data are acquired, and the target data in the target data groups are processed.
9. The method for processing data according to claim 6, characterized in that: The processing of the target data in the data group includes: For the case where the data processing type is variance or standard deviation, Determining, for the data group, first target data corresponding to a data processing type of sum or mean and second target data corresponding to a data processing type of variance or standard deviation provided by the data provider; The variance or standard deviation of the data group is calculated based on a calculated relationship between the square of the mean of the first target data and the mean of the second target data.
10. A device for processing data, characterized in that: Applied to the data provider, including: data acquisition module, data processing module and data sending module; among them, The data acquisition module is configured to receive a data acquisition request sent by a data requester; the data acquisition request indicates one or more data ranges and a data processing type; and acquire, from a data source, a plurality of pieces of original data within the one or more data ranges indicated by the data acquisition request; The data processing module is configured to generate corresponding target data for each piece of the original data according to the data processing type and the noise data corresponding to the data processing type determined for the plurality of original data, wherein the noise data corresponding to the data processing type enables the plurality of target data to meet a grouping requirement; The data sending module is used to send the plurality of target data to a data demander, so that the data demander processes the plurality of target data based on the data processing type; The apparatus is configured to, after acquiring a plurality of pieces of original data belonging to one or more of the data ranges from the data source, further comprise: calculating a mean value corresponding to the data processing type based on the plurality of original data; and generating noise data of the data processing type based on the mean value corresponding to the data processing type, a first data volume threshold value of data of any data group further indicated by the data acquisition request, and a processing error; Generating the noise data of the data processing type includes: calculating a privacy budget parameter based on a mean corresponding to the data processing type, a first data volume threshold of any data group further indicated by the data acquisition request, and a processing error; and inputting the privacy budget parameter into a Laplace distribution function to obtain the noise data; wherein the formula for calculating the privacy budget parameter is as follows: Where ε represents the privacy budget parameter, n represents the first data volume threshold, a represents the processing error, avg(R1) represents the mean corresponding to the data processing type, and R1 represents the data sequence of multiple original data.
11. A device for processing data, characterized in that: Applied to the data demand side, including: data request module and data processing module; among them, The data request module is configured to send a data acquisition request to a data provider, wherein the data acquisition request indicates one or more data ranges and data processing types; The data processing module is configured to, upon receiving a plurality of target data sent by the data provider, divide the plurality of target data into one or more data groups and process the target data within the data groups, wherein the target data is formed based on the original data of the data provider and noise data corresponding to the data processing type, and the noise data corresponding to the data processing type enables the plurality of target data to meet the grouping requirements; wherein, after the data provider obtains a plurality of original data belonging to one or more data ranges from the data source, the module further comprises: calculating a mean corresponding to the data processing type based on the plurality of original data; generating noise data of the data processing type based on the mean corresponding to the data processing type, a first data volume threshold of data of any data group further indicated by the data acquisition request, and a processing error; generating noise data of the data processing type includes: calculating a privacy budget parameter based on the mean corresponding to the data processing type, a first data volume threshold of data of any data group further indicated by the data acquisition request, and a processing error; and inputting the privacy budget parameter into a Laplace distribution function to obtain the noise data; The formula for calculating the privacy budget parameter is as follows: Where ε represents the privacy budget parameter, n represents the first data volume threshold, a represents the processing error, avg(R1) represents the mean corresponding to the data processing type, and R1 represents the data sequence of multiple original data.
12. A data processing system comprising a data providing end of the data processing device according to claim 10 and a data demanding end of the data processing device according to claim 11.
13. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.
14. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Data transmission security control method and device, computer equipment and Internet of Things system
CN110430218A
Data processing method and data processing device
CN112835904A