Data processing method and device, electronic equipment, storage medium and program product

By selecting appropriate processing strategies based on the dimensions and attributes of user data, and performing perturbation processing and shuffling on multi-dimensional user data, the problem of low availability of user data after privacy protection is solved, and the validity of data is maintained while protecting privacy.

CN120705903APending Publication Date: 2025-09-26CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510742013.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When performing privacy protection on multi-dimensional user data, the existing technology has the problem of low availability of user data after privacy protection.

Method used

According to the dimension and privacy budget of user data, different processing strategies for categorical attributes and numerical attributes are determined, the user data is perturbed by calculating the probability, and the data is shuffled using a shuffler to generate perturbed user data.

Benefits of technology

While ensuring privacy protection, it improves the availability of user data and ensures that personal privacy is not leaked during data transmission and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705903A_ABST
    Figure CN120705903A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, electronic equipment, a storage medium and a program product. The method comprises the steps that after to-be-processed user data is obtained, a target processing strategy corresponding to the to-be-processed user data is determined according to the data dimension of the to-be-processed user data or a preset privacy budget, and the target processing strategy is used for conducting disturbance processing on the to-be-processed user data according to the attribute and the value range of each piece of user sub-data; the attributes of the user sub-data comprise classification attributes or numerical attributes. And finally, executing the corresponding target processing strategy on the to-be-processed user data to generate disturbed user data. According to the method, the availability of the user data after privacy protection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information security technology, and in particular to a data processing method, device, electronic device, storage medium and program product. Background Art

[0002] With the rapid development of information technology, user data has become a key resource driving the development of an intelligent society. This data can be used for statistical analysis, model training, and other tasks, thereby analyzing user behavior and improving the user experience. However, while utilizing this data, how to protect its privacy is a key concern.

[0003] In the existing technology, differential privacy (DP) is used independently on the terminal device to protect the privacy of user data. Noise is injected into each user's data through the local DP model and then sent to the server, so that the server obtains the blurred data.

[0004] However, for multi-dimensional user data, existing technologies suffer from the problem of low availability of user data after privacy protection. Summary of the Invention

[0005] Embodiments of the present application provide a data processing method, apparatus, electronic device, storage medium, and program product to improve the availability of privacy-protected user data.

[0006] In a first aspect, an embodiment of the present application provides a data processing method, comprising:

[0007] Obtaining user data to be processed, which includes user sub-data in multiple dimensions;

[0008] Determine the target processing strategy for the user data to be processed based on the data dimensions or the preset privacy budget. The target processing strategy is used to perturb the user data to be processed based on the attributes and value range of each user sub-data. The attributes of the user sub-data may include categorical attributes or numerical attributes.

[0009] Execute the corresponding target processing strategy on the user data to be processed to generate disturbed user data.

[0010] In one possible implementation, if the data dimension of the user data to be processed is greater than a preset data dimension, or if the preset privacy budget of the user data to be processed does not exceed a preset privacy budget threshold, a first processing strategy is determined as the target processing strategy. The first processing strategy is used to sequentially calculate a first probability for each user sub-data based on the attribute of each user sub-data and the corresponding value range, and perform perturbation processing on the corresponding user sub-data in the user data to be processed based on the first probability.

[0011] In one possible implementation, when the data dimension of the user data to be processed does not exceed the preset data dimension, or the preset privacy budget of the user data to be processed is less than the preset privacy budget threshold, the second processing strategy is determined as the target processing strategy; the second processing strategy is used to calculate the second probability of the user data to be processed based on the attributes of each user sub-data and the corresponding value range, and perform perturbation processing on the user data to be processed according to the second probability.

[0012] In one possible implementation, the first processing strategy includes a first sub-processing strategy for processing user sub-data of categorical attributes and a second sub-processing strategy for processing user sub-data of numerical attributes;

[0013] For the first sub-processing strategy, a first probability is determined based on the preset sub-privacy budget of the user sub-data and the number of categories within the value range; based on the first probability, the user sub-data, and the value range, the perturbed user sub-data is determined, where the probability that the perturbed user sub-data is the user sub-data is the first probability, and the probability that the perturbed user sub-data is other data within the value range is the third probability, where the sum of the first and third probabilities is 1;

[0014] For the second sub-processing strategy, a first probability is determined based on the preset sub-privacy budget of the user sub-data; based on the user sub-data, a first sub-value range including the user sub-data is determined from the value range; based on the first probability, the user sub-data, the value range and the first sub-value range, the perturbed user sub-data is determined, the probability that the perturbed user sub-data is data within the first sub-value range is the first probability, the probability that the perturbed user sub-data is other data within the second sub-value range is the third probability, the second sub-value range includes the first sub-value range and the second sub-value range belongs to the value range.

[0015] In a possible implementation, for the first sub-processing strategy, the first probability is determined according to a preset sub-privacy budget of the user sub-data, a preset sub-relaxed privacy parameter, and the number of categories within a value range.

[0016] In a possible implementation, for the second sub-processing strategy, the first probability is determined according to a preset sub-privacy budget and a preset sub-relaxed privacy parameter of the user sub-data.

[0017] In one possible embodiment, the second processing strategy includes: for user sub-data of categorical attributes, determining the neighboring probability of the user sub-data based on the number of categories within the value range; for user sub-data of numerical attributes, determining the neighboring probability of the user sub-data based on the interval length of the first sub-value range; generating a total neighboring probability of the user data to be processed based on the neighboring probabilities of all user sub-data; generating a second probability of the user data to be processed based on the total neighboring probability and a preset privacy budget; generating perturbed user data based on the second probability, the perturbed user data including multiple perturbed user sub-data, the perturbed user sub-data being the ratio of the number of other data to the total number of perturbed user sub-data, which is the same as the fourth probability; the sum of the fourth probability and the second probability is 1.

[0018] In a possible implementation, the disturbed user data is shuffled by a shuffler to generate target user data; and the target user data is sent to a server.

[0019] In a second aspect, an embodiment of the present application provides a data processing device, including:

[0020] An acquisition module is used to acquire user data to be processed, where the user data to be processed includes user sub-data in multiple dimensions;

[0021] A determination module is used to determine a target processing strategy for the user data to be processed based on its data dimensions or a preset privacy budget. The target processing strategy is used to perturb the user data to be processed based on the attributes and value range of each user sub-data. The attributes of the user sub-data may include categorical attributes or numerical attributes.

[0022] The generation module is used to execute the corresponding target processing strategy on the user data to be processed and generate the disturbed user data.

[0023] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor;

[0024] Memory stores computer-executable instructions;

[0025] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0026] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.

[0027] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0028] The embodiments of the present application provide a data processing method, apparatus, electronic device, storage medium, and program product. First, after acquiring user data to be processed that includes user sub-data of multiple dimensions, corresponding target processing strategies are determined for user data to be processed with different data dimensions or different preset privacy budgets, thereby meeting the privacy requirements of different data. Furthermore, the target processing strategy is formulated based on the attributes and value range of each user sub-data, so as to maximize the availability of the retained data while meeting data privacy requirements. Finally, the corresponding target processing strategy is executed on the user data to be processed to generate disturbed user data. This improves the availability of user data after privacy protection. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0030] Figure 1 Schematic diagram of the data processing method provided in this embodiment Figure 1 ;

[0031] Figure 2 Schematic diagram of the data processing method provided in this embodiment Figure 2 ;

[0032] Figure 3 A schematic diagram of a scenario of a data processing method provided in an embodiment of the present application;

[0033] Figure 4 A schematic diagram of a flow chart of a first processing strategy provided in an embodiment of the present application;

[0034] Figure 5 A schematic flow chart of the second processing strategy provided in an embodiment of the present application;

[0035] Figure 6 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;

[0036] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0037] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0038] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0039] First, let’s explain the terms involved in this application:

[0040] Shuffler: This refers to an intermediary component that receives data and then rearranges or shuffles it. The purpose is to further obfuscate the data's origin, making it difficult for attackers to infer a specific user's original information by analyzing data patterns.

[0041] Spatial crowdsourcing: refers to the process of utilizing the spatial perception capabilities (such as location, image, and sound) of a large number of ordinary users to collaboratively complete tasks or data collection based on geographic location;

[0042] Federated learning: refers to a distributed machine learning method in which multiple participants collaborate to train models without sharing the original data to protect data privacy and achieve decentralized learning;

[0043] Hockey-stick divergence (Hockey-stick divergence): refers to the difference between two distributions, especially the difference in the tail, which makes it more effective when dealing with distributions with large tails. For example, suppose there are two probability distributions P and Q, their Hockey-stick divergence can be expressed as: where P and Q represent random variables and their probability density functions, and ∈ represents the local privacy budget.

[0044] If two variables P and Q satisfy Where δ represents the shuffler / shuffling program. Then the two variables P and Q are (∈, δ)-indistinguishable. For two datasets of the same size that differ only in a single user's data, they are called adjacent datasets. Differential privacy constrains the differences in query results on adjacent datasets (see the definition of differential privacy). Similarly, in the local scenario where a single individual's data is received as input, local (∈, δ)-differential privacy is introduced in the definition of local differential privacy. When δ = 0, this concept is simply referred to as ∈-LDP.

[0045] Differential privacy: refers to protecting individual privacy. A protocol Satisfies (∈,δ)-differential privacy if and only if for all adjacent datasets X,X′∈X n , and is (∈,δ)-indistinguishable. Where n represents the number of users, X represents the user data domain, Represents a randomization algorithm.

[0046] Local Differential Privacy (LDP): A Protocol satisfies (∈,δ)-local differential privacy if and only if for all x,x′∈X, and is (∈,δ)-indistinguishable. If δ = 0, it is denoted as ∈-LDP. Where Y represents the domain of the message after the shuffler obfuscation.

[0047] Data Processing Inequality (DPI): The Data Processing Inequality (DPI) describes the irreversible loss of information during data processing and is a key characteristic of distance measures used for data privacy, such as the Hockey-stick divergence. It asserts that further analysis of the mechanism's output does not weaken privacy guarantees. The DPI provides a theoretical basis for ensuring privacy protection, proving that the amount of sensitive information leaked during different data processing steps does not exceed the original amount.

[0048] A distance metric D on the probability distribution space: If the data processing inequality is satisfied, then and only if for For all distributions P and Q in and all (possibly random) functions g: D(g(P)||g(Q))≤D(P||Q).

[0049] Differential Privacy under the Shuffle Model: A Protocol Satisfies (∈,δ)-differential privacy in the shuffle model if and only if for all adjacent datasets X,X′∈X n , news after shuffling and is (∈,δ)-indistinguishable. Privacy amplification through shuffling,A key component of the shuffling model is the privacy amplification achieved through shuffling. Existing work shows that after the shuffling operation, the ∈-LDP messages from n users satisfy Differential privacy.

[0050] With the rapid development of information technology, data-driven development has become a vital force driving social progress and industrial upgrading. The collection, analysis, and application of massive amounts of data are profoundly changing the operational models of various industries. The analysis and utilization of user data can be driven by analyzing user behavior from a data-driven perspective, leading to a deeper understanding of user needs, preferences, and usage habits, thereby improving the user experience. However, while utilizing user data, the protection of user privacy is a key concern.

[0051] In existing technologies, privacy protection for user data is implemented locally and independently on terminal devices using a data privacy controller (DP). Using the DP's local model, a significant amount of noise is injected into each user's data before the data is sent to the DP's central model for publication. However, when user data involves multiple dimensions, the noise is injected into all of these dimensions, making the data ultimately sent to the central model particularly ineffective. This demonstrates that existing technologies suffer from the low usability of user data after privacy protection.

[0052] In this regard, the inventors believe that different processing strategies should be used to perturb user data with larger and smaller dimensions, thereby preventing the same processing strategy from reducing the usability of multi-dimensional data. For data with categorical or numerical attributes, the required privacy budget varies due to different privacy protection requirements, and privacy budget is also considered as a factor in the processing strategy. Furthermore, for each piece of user data, the amount of noise required for perturbation depends on the data's value range. Therefore, the processing strategy needs to perturb user data based on the attribute and value range. This improves the usability of user data after privacy protection.

[0053] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0054] Figure 1 Schematic diagram of the data processing method provided in this embodiment Figure 1 ,like Figure 1 As shown, the method includes:

[0055] S101: Obtain user data to be processed.

[0056] The user data to be processed refers to the raw data set collected from users. It includes user sub-data across multiple dimensions. Each dimension represents a different type of information, and the data within each dimension is referred to as user sub-data. The user sub-data across multiple dimensions is combined to form the user data to be processed.

[0057] It should be understood that the user data to be processed can include data from multiple fields, such as statistical analysis, machine learning, recommendation systems, spatial crowdsourcing, digital health, and smart cities. Because this data is extremely vulnerable to privacy leaks and abuse, threatening not only individual privacy security but also social and national security, it is necessary to find an ideal processing strategy to protect user data.

[0058] For example, the data of a certain user in the user data to be processed is used Indicates that i is used to represent the user and d is used to represent the dimension.

[0059] S102: Determine a target processing strategy corresponding to the user data to be processed based on the data dimension of the user data to be processed or a preset privacy budget.

[0060] The data dimension refers to the number of data types in the user data. For example, if the user data to be processed includes age, gender, and occupation, the corresponding data dimension is 3.

[0061] A privacy budget is a parameter that measures the strength of privacy protection, determining how much privacy is allowed to be leaked when perturbing data. A smaller privacy budget results in stronger privacy protection; a larger privacy budget results in more accurate data but weaker privacy protection.

[0062] The processing strategy refers to how data is perturbed. Balancing privacy budgets and data availability (i.e., data accuracy) is a key factor in selecting a processing strategy. The target processing strategy perturbs the user data to be processed based on the attributes and value range of each user's sub-data.

[0063] The attributes of user sub-data include categorical attributes and numerical attributes. Categorical attributes refer to data that represents a category or label, such as gender or blood type. Numerical attributes refer to data with a quantified value, such as age or income.

[0064] The value range is the set of all possible legal values ​​that the user sub-data can take. The value range can be discrete or continuous, depending on the attribute type.

[0065] For example, when the user sub-data is "gender", it is a categorical attribute with a value range of {male, female}, and the value range is discrete; when the user sub-data is "age", it is a numerical attribute with a value range of [0, 120], and the value range is continuous.

[0066] In order to enhance the usability of data after disturbance, different processing strategies can be determined according to the data dimension or the preset privacy budget. Therefore, in one possible implementation, when the data dimension of the user data to be processed is greater than the preset data dimension, or the preset privacy budget of the user data to be processed does not exceed the preset privacy budget threshold, the first processing strategy is determined as the target processing strategy.

[0067] If the data dimension of the user data to be processed does not exceed the preset data dimension, or the preset privacy budget of the user data to be processed is less than the preset privacy budget threshold, the second processing strategy is determined as the target processing strategy.

[0068] Among them, the preset data dimension refers to a parameter for measuring the complexity of user data. When the data dimension of user data is greater than the preset data dimension, it means that the complexity of user data is relatively high, and the first processing strategy is selected; when the data dimension of user data is less than or equal to (i.e., not exceeding) the preset data dimension, it means that the complexity of user data is relatively low, and the second processing strategy is selected.

[0069] The preset privacy budget threshold refers to a parameter that measures the size of the preset privacy budget. When the preset privacy budget of user data is less than or equal to (i.e., does not exceed) the preset privacy budget threshold, it means that the degree of privacy protection required for the user data to be processed is relatively strong, and the first processing strategy is selected. When the preset privacy budget of user data is greater than the preset privacy budget threshold, it means that the degree of privacy protection required for the user data to be processed is relatively weak, and the second processing strategy is selected.

[0070] Among them, the first processing strategy refers to a data processing strategy suitable for high-dimensional data or low privacy budget, and the second processing strategy refers to a data processing strategy suitable for low-dimensional data or high privacy budget.

[0071] The second treatment strategy is Figure 2 The details are described in detail in the embodiments of the present invention.

[0072] In practical applications, the first processing strategy is used to sequentially calculate a first probability for each user sub-data based on the attributes and corresponding value range of each user sub-data, and perform perturbation processing on the corresponding user sub-data in the user data to be processed based on the first probability. The second processing strategy is used to calculate a second probability for the user data to be processed based on the attributes and corresponding value range of each user sub-data, and perform perturbation processing on the user data to be processed based on the second probability.

[0073] Among them, when the attribute of the user sub-data is a categorical attribute, the first probability and the second probability refer to the probability that the user sub-data after disturbance is still the true value; when the attribute of the user sub-data is a numerical attribute, the first probability refers to the probability that the user sub-data after disturbance is close to the true value, and the second probability refers to the probability that the user sub-data after disturbance is still the true value.

[0074] In a possible implementation, the first processing strategy includes a first sub-processing strategy for processing user sub-data of categorical attributes and a second sub-processing strategy for processing user sub-data of numerical attributes.

[0075] It should be understood that different processing strategies are selected for different attributes of user sub-data.

[0076] For the first sub-processing strategy, a first probability is first determined based on the preset sub-privacy budget of the user sub-data and the number of categories within the value range. Then, the perturbed user sub-data is determined based on the first probability, the user sub-data, and the value range.

[0077] The probability that the disturbed user sub-data is user sub-data is a first probability, the probability that the disturbed user sub-data is other data within a value range is a third probability, and the sum of the first probability and the third probability is 1.

[0078] It should be understood that when the attribute of the user sub-data is a categorical attribute, the perturbed user sub-data may remain that data or may be another data type within the value range. For example, if the value range is {"teacher", "doctor", "engineer"}, and the user sub-data is "doctor", then the perturbed user sub-data may have a first probability of being "doctor" (i.e., still that data type) and a third probability of being "teacher" or "engineer" (i.e., another data type within the value range), thereby perturbing the data.

[0079] For example, for the first sub-processing strategy, that is, when the attribute of the user sub-data is a categorical attribute, the data can be perturbed using Generalized Random Response (GRR). In this case, the first probability refers to the probability of outputting the true value, which is expressed by the formula Determine the first probability, that is, at this time, for each data in the user sub-data, output the true value with the possibility of the first probability p. At the same time, with the third probability The possibility of outputting other values.

[0080] Among them, ∈ j It is used to represent the preset sub-privacy budget of the user's sub-data. The sum of the privacy budgets of all dimensions is the preset privacy budget ∈, which satisfies the formula ∈=∑ j∈[d] ∈ j , j is used to represent the dimension. Used to represent the sub-data of user i, It is used to indicate the number of categories of the sub-data of user i within the value range.

[0081] For example, when a user's sub-data is occupation "doctor", the occupation value range is set to {"teacher", "doctor", "engineer"}, that is, There are three types of categories. When the privacy budget of the current dimension j is set ∈ j When it is 1, the first probability 0.576, the third probability is obtained The user sub-data outputs the true value "doctor" with a probability of 57.6%, that is, the user sub-data after disturbance is still the user sub-data, and outputs other values ​​"teacher" with a probability of 21.2%, or outputs other values ​​"engineer" with a probability of 21.2%, that is, the user sub-data after disturbance is other data within the value range.

[0082] Accordingly, for the second sub-processing strategy, a first probability is first determined based on the preset sub-privacy budget of the user sub-data. Then, based on the user sub-data, a first sub-value range that includes the user sub-data is determined from within the value range. Finally, the perturbed user sub-data is determined based on the first probability, the user sub-data, the value range, and the first sub-value range.

[0083] Among them, the probability that the user sub-data after disturbance is data within the first sub-value range is the first probability, the probability that the user sub-data after disturbance is other data within the second sub-value range is the third probability, the second sub-value range includes the first sub-value range and the second sub-value range belongs to the value range.

[0084] It should be understood that when the attributes of user sub-data are numerical attributes, the perturbed user sub-data may be data within the first sub-value range or data within the second sub-value range. The first sub-value range is smaller than the second sub-value range, but both ranges fall within the value range. The difference is that when determining the perturbed user sub-data from the first sub-value range, the probability of selecting the true value is greater than when determining it from the second sub-value range. For example, when the value range is set to [0, 120] and the user sub-data represents age, the user sub-data is 30, representing the user's true age of 30. In this case, a number is randomly selected from [20, 50] with a first probability as the perturbed user sub-data, while a number is randomly selected from [10, 60] with a third probability as the perturbed user sub-data. It can be seen that the former is more likely to extract the true value than the latter, thereby protecting the data and preserving its availability.

[0085] In a possible implementation, a square wave mechanism may be used to convert the first sub-value range into an interval To determine, where l is the interval length, the setting of the interval length should satisfy the requirement of retaining data availability, that is, the interval length is set to be smaller so that the first sub-value range includes the range near the true value, so that the perturbed data determined from the first sub-value range can be close to the true value.

[0086] For example, if the user's sub-data represents age, the user's real age is 30 years old, the preset sub-privacy budget is set to 1, and the interval length l is set to 4, the formula can be used Calculate the first probability The third probability 1-p = 0.269. It can be seen that there is a probability of about 73.1% to output a value close to the true value, which is within the first sub-value range [28, 32]. There is a probability of 26.9% to output a value far from the true value, which is within the second sub-value range. The second sub-value range can be set to [25, 40].

[0087] In another possible implementation, a scale parameter ∈ is added to the user sub-data through the Laplace mechanism. j Noise, so that each dimension after perturbation satisfies ∈ j -LDP, through the basic combinatorial theorem, the entire perturbed data satisfies ∈-LDP.

[0088] In practical applications, to enhance the perturbation effect, a relaxed privacy parameter can be introduced to more accurately determine the first probability. In one possible implementation, for the first sub-processing strategy, the first probability is determined based on a preset sub-privacy budget for the user sub-data, a preset sub-relaxed privacy parameter, and the number of categories within a range of values.

[0089] For example, for the first sub-processing strategy, which is a processing strategy for processing user sub-data of classification attributes, a generalized random response is adopted, with a first probability Output true value With the third probability Randomly output other values, where δ j Used to represent the preset sub-relaxed privacy parameter, ∈ j Used to represent the preset sub-privacy budget, Used to indicate the number of categories within a value range.

[0090] It should be understood that the preset sub-relaxation privacy parameter δ j Satisfying δ=∑ j∈[d] δ j , where δ is the preset relaxed privacy parameter.

[0091] Correspondingly, for the second sub-processing strategy, determining the first probability according to the preset sub-privacy budget of the user sub-data includes: for the second sub-processing strategy, determining the first probability according to the preset sub-privacy budget of the user sub-data and the preset sub-relaxation privacy parameter.

[0092] For example, the second sub-processing strategy, which is a processing strategy for processing user sub-data of numerical attributes, adopts a Laplace mechanism or a square wave mechanism. The square wave mechanism is in the interval (l is the interval length) with the first probability Uniform sampling is used, otherwise random selection is made within a larger range with a third probability 1-p.

[0093] S103: Execute a corresponding target processing strategy on the user data to be processed to generate disturbed user data.

[0094] Among them, the disturbed user data refers to the data processed by the privacy protection mechanism (i.e., the target processing strategy). Its value is close to the original true value under a certain probability, or it may be the original true value, and deviate from the original value under another probability, thereby ensuring that data privacy is not leaked while preserving the data availability as much as possible.

[0095] In practical applications, the disturbed user data can be To indicate that, Used to represent a user's data. Each perturbed data point may remain as the pre-perturbed data point, or may be a value close to the pre-perturbed data point, or may be a value far from the true value within the range of values. This protects the privacy of individual user data and avoids the risk of data leakage.

[0096] In a possible implementation, the disturbed user data is shuffled by a shuffler to generate target user data, and then the target user data is sent to the server.

[0097] It should be understood that after the perturbed user data passes through the shuffler, the data's order is disrupted, making it impossible for an attacker to infer which user a record belongs to by observing the data sequence. Furthermore, even if some records have low noise, it is difficult for an attacker to associate them with a specific user because the positions of all records are randomized. Furthermore, although the data order is disrupted, the statistical characteristics of the overall dataset remain unchanged, making it still useful for histogram estimation, federated learning, and other applications.

[0098] In actual applications, target user data is sent to the server so that the server can aggregate and analyze large amounts of user data to extract useful statistical information or build a machine learning model.

[0099] The present application provides a data processing method that first obtains user data to be processed, including user sub-data of multiple dimensions. Based on the data dimensions or a preset privacy budget, a target processing strategy is determined for each user sub-data. This allows for targeted selection of processing strategies for the user data. The target processing strategy varies depending on the attributes of the user sub-data. The data is perturbed using the strategy to generate perturbed user data. This method selects different perturbation methods based on the data dimensions, privacy budget, and data attributes, thereby improving the usability of the privacy-protected user data.

[0100] In practical applications, both processing strategies of this application can achieve multi-dimensional user data publishing to ensure that each dimension can retain information after data randomization (i.e., no dimensional sampling is performed) to support the goal of various applications that require cross-attribute information. Among them, the first processing strategy divides the overall local budget ∈ into d parts and uses a randomizer with different budgets to process each dimension independently. The second processing strategy publishes all attributes jointly and uses the entire local budget ∈, which does not necessarily require cross-dimensional independence. Both strategies ensure that after shuffling, (∈ c ,δ)-DP.

[0101] Figure 2 Schematic diagram of the data processing method provided in this embodiment Figure 2 ,like Figure 2 As shown, this embodiment Figure 1 Based on the embodiment, the second processing strategy in the data processing method is described in detail. The second processing strategy includes:

[0102] S201. For user sub-data of categorical attributes, determine the proximity probability of the user sub-data according to the number of categories within a value range;

[0103] S202: For user sub-data of numerical attributes, determine the proximity probability of the user sub-data according to the interval length of the first sub-value range;

[0104] S203. Generate a total neighboring probability of the user data to be processed based on the neighboring probabilities of all user sub-data;

[0105] S204: Generate a second probability of the user data to be processed according to the total neighboring probability and the preset privacy budget.

[0106] The number of categories refers to the number of possible values ​​within a category. For example, for the categorical attribute "gender," the categories are "male" and "female."

[0107] The neighborhood probability refers to the probability that the user data after perturbation is still the true value or close to the true value.

[0108] The interval length refers to the size of the neighborhood around the true value.

[0109] For example, by the formula: Calculate the neighboring probability of the user sub-data corresponding to each attribute nearby j ,in, Used to represent user sub-data, It is used to indicate the number of categories, and l is used to indicate the length of the interval.

[0110] Then, through the formula: The total neighboring probability nearby of the user data to be processed is calculated, where d is used to represent the dimension of the user sub-data.

[0111] Then, by the formula: The second probability of the user data to be processed is calculated, where ∈ is used to represent the preset privacy budget.

[0112] In another possible implementation, for user sub-data with numerical attributes, the proximity probability is set to 0.5 when using the Laplace mechanism. Based on the proximity probabilities of all user sub-data and the preset sub-privacy budget, the total proximity probability of the user data to be processed is generated, and β is calculated as the second probability. It should be understood that the entire record is considered "contiguous" only when all attributes are "contiguous."

[0113] S205: Generate disturbed user data according to the second probability.

[0114] The disturbed user data includes multiple disturbed user sub-data, and the disturbed user sub-data is the ratio of the number of other data to the total number of disturbed user sub-data, which is the same as the fourth probability; the sum of the fourth probability and the second probability is 1.

[0115] It should be understood that the disturbed user data comes from the user data outputting the true value with the second probability β, that is, the disturbed user sub-data is still the user sub-data, and outputting other values ​​with the fourth probability 1-β, that is, the disturbed user sub-data is other data within the value range.

[0116] For example, when the user data to be processed includes "age", "height", "weight" and "income", it can be seen that they are all numerical attributes. When using the Laplace mechanism, the neighboring probabilities of the user sub-data are all 0.5, then the total neighboring probability of the user data to be processed is nearby = 0.5*0.5*0.5*0.5 = 0.0625.

[0117] Next, based on the known preset privacy budget of 1, according to the formula Obtain the second probability That is, the true value is output according to the 10.05% probability as the disturbed user sub-data; according to the fourth probability 1-β=0.8995, that is, other values ​​are output according to the 89.95% probability, that is, the disturbed user sub-data is other data within the value range.

[0118] An embodiment of the present application provides a data processing method that first determines the neighborhood probability of user sub-data with categorical attributes based on the number of categories; and for user sub-data with numerical attributes, determines the neighborhood probability based on the interval length. The neighborhood probabilities of the user sub-data across all dimensions are then combined to form a total neighborhood probability, which is then combined with a preset privacy budget to generate a second probability. Finally, the second probability is used to output a value close to the true value, and the fourth probability is used to output other values. This method perturbs user data by attribute without splitting the privacy budget, and perturbs user sub-data across multiple dimensions simultaneously, preserving the usability of the privacy-protected user data.

[0119] Figure 3 A schematic diagram of a data processing method according to an embodiment of the present invention is shown in FIG. Figure 3 As shown in the figure, the data from multiple users are perturbed by a local randomizer and then shuffled by a shuffler. This method can be applied to scenarios such as histogram estimation, spatial crowdsourcing, federated learning, and causal reasoning.

[0120] Figure 4 A flow chart of the first processing strategy provided in the embodiment of the present application is shown as follows: Figure 4 Shown, including:

[0121] S401, obtain the original data input by the user end, use Indicates that i is used to represent the user and d is used to represent the dimension;

[0122] S402: For user sub-data with numerical attributes, perturbation is performed using a square wave mechanism or a Laplace mechanism; for user sub-data with categorical attributes, perturbation is performed using a generalized random response mechanism;

[0123] S403, combine the disturbed data, and use To express;

[0124] S404, sending the data combination to the shuffler;

[0125] S405. Send the data output by the shuffler to the server, so that the server applies single attribute statistics and joint distribution statistics and outputs statistical results.

[0126] In practical applications, for the first processing strategy, an independent randomizer is designed, and the jth attribute will pass through the randomizer R j :X j →Yj Processed independently and using budget∈ j All these randomizers are combined into a ∈-LDP randomizer It maps data X to Y, where ∈ = ∑ j∈[d] ∈ j Each user then passes the processed attributes Sent as a message to the shuffler.

[0127] It should be understood that independent randomization is suitable for high-dimensional or numerical data. For example, for large-scale, high-dimensional datasets, independent randomization combined with the shuffle model reduces the estimation error of numerical attributes (such as age) by 99%, significantly outperforming the local LDP model.

[0128] Figure 5 A flow chart of the second processing strategy provided in the embodiment of the present application is shown as follows: Figure 5 Shown, including:

[0129] S501, obtain the original data input by the user end, use Indicates that i is used to represent the user and d is used to represent the dimension;

[0130] S502: Calculate the neighboring probabilities of the user sub-data for numerical attributes and categorical attributes, and then determine the total neighboring probability of the user data to be processed;

[0131] S503: Output the true value of the original data input by the user terminal with a second probability β, and output other values ​​with a fourth probability of 1-β;

[0132] S504, combining the disturbed data and sending it to the shuffler;

[0133] S505: Send the data output by the shuffler to the server, so that the server applies single attribute statistics and joint distribution statistics and outputs statistical results.

[0134] In practical applications, the second processing strategy is to de-identify the user data as a whole. Specifically, for the jth dimension / attribute, the real value x i,j Define a nearby domain near j (x i,j ), x i,j That is to say For example, in the category data field X j In , the near domain can be defined as a set containing only true values: j (x i,j )={x i,j}. In the numeric data field X jThe neighborhood can be defined as a relatively small range between the true value x i,j The distance r j Within: near j (x i,j )={y|y∈Rand|yx i,j |≤r j}, y is a real number. For other types of attributes, the nearby area can be defined as needed. For simplicity, assume that near j (v) for all possible v∈X j have the same measurements. j To represent the measure of the neighborhood of the j-th dimension (for example, Lebesgue measure or other related measures).

[0135] For the second processing strategy, a joint randomizer is designed, which has a probability λ = [0.0, 1.0] from the joint neighborhood area Output a value uniformly randomly from the joint output domain Y = ∏ with probability 1-λ j∈[d] Y j Choose a value uniformly randomly from . Represents the union of all possible neighborhood areas, with v j To represent Y j The measure of , that is, the size of the output domain in that dimension. Formally, the joint randomizer works as follows:

[0136]

[0137] Among them, uniform(near(x i )) means selecting a value uniformly at random from the neighborhood of the i-th user, and uniform(Y) means selecting a value uniformly at random from the joint output domain Y.

[0138] It should be understood that joint randomization preserves the correlation between attributes through overall perturbation and performs well in low-dimensional categorical data. For example, in a small, low-dimensional dataset, joint randomization reduces single-attribute classification error by 95% and joint distribution error by 99.5%, far exceeding independent methods.

[0139] Figure 6 A structural diagram of a data processing device provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the data processing device 60 provided in this embodiment includes:

[0140] An acquisition module 601 is configured to acquire user data to be processed, where the user data to be processed includes user sub-data in multiple dimensions.

[0141] Determination module 602 is used to determine a target processing strategy for the user data to be processed based on the data dimension or the preset privacy budget of the user data to be processed. The target processing strategy is used to perturb the user data to be processed based on the attributes and value range of each user sub-data. The attributes of the user sub-data may include categorical attributes or numerical attributes.

[0142] The generation module 603 is used to execute the corresponding target processing strategy on the user data to be processed to generate disturbed user data.

[0143] In one possible implementation, the determination module 602 is further configured to determine a first processing strategy as a target processing strategy when the data dimension of the user data to be processed is greater than a preset data dimension, or the preset privacy budget of the user data to be processed does not exceed a preset privacy budget threshold. The first processing strategy is configured to calculate a first probability of each user sub-data in sequence based on the attributes of each user sub-data and the corresponding value range, and perform perturbation processing on the corresponding user sub-data in the user data to be processed based on the first probability.

[0144] In one possible implementation, the determination module 602 is further configured to determine the second processing strategy as the target processing strategy when the data dimension of the user data to be processed does not exceed a preset data dimension, or the preset privacy budget of the user data to be processed is less than a preset privacy budget threshold; the second processing strategy is configured to calculate a second probability of the user data to be processed based on the attributes of each user sub-data and the corresponding value range, and perform perturbation processing on the user data to be processed based on the second probability.

[0145] In one possible implementation, the determination module 602 is further configured to determine, for the first sub-processing strategy, a first probability based on a preset sub-privacy budget of the user sub-data and the number of categories within a value range; determine the perturbed user sub-data based on the first probability, the user sub-data, and the value range, where the probability that the perturbed user sub-data is the user sub-data is the first probability, the probability that the perturbed user sub-data is other data within the value range is the third probability, and the sum of the first probability and the third probability is 1;

[0146] In one possible implementation, the determination module 602 is further used to determine, for the second sub-processing strategy, a first probability based on a preset sub-privacy budget of the user sub-data; determine, from the value range, a first sub-value range containing the user sub-data based on the user sub-data; determine the disturbed user sub-data based on the first probability, the user sub-data, the value range, and the first sub-value range, the probability that the disturbed user sub-data is data within the first sub-value range is the first probability, the probability that the disturbed user sub-data is other data within the second sub-value range is the third probability, the second sub-value range includes the first sub-value range, and the second sub-value range belongs to the value range.

[0147] In a possible implementation, the determination module 602 is further configured to determine, for the first sub-processing strategy, a first probability according to a preset sub-privacy budget of the user sub-data, a preset sub-relaxed privacy parameter, and the number of categories within a value range.

[0148] In a possible implementation, the determination module 602 is further configured to determine, for the second sub-processing strategy, the first probability according to a preset sub-privacy budget and a preset sub-relaxed privacy parameter of the user sub-data.

[0149] In one possible implementation, the determination module 602 is further used to determine, for user sub-data of categorical attributes, the neighboring probability of the user sub-data based on the number of categories within the value range; for user sub-data of numerical attributes, determine the neighboring probability of the user sub-data based on the interval length of the first sub-value range; generate a total neighboring probability of the user data to be processed based on the neighboring probabilities of all user sub-data; generate a second probability of the user data to be processed based on the total neighboring probability and a preset privacy budget; generate perturbed user data based on the second probability, the perturbed user data including multiple perturbed user sub-data, the perturbed user sub-data being the ratio of the number of data of other data to the total number of perturbed user sub-data, which is the same as the fourth probability; the sum of the fourth probability and the second probability is 1.

[0150] In a possible implementation, the generation module 603 is further configured to shuffle the disturbed user data through a shuffler to generate target user data; and send the target user data to the server.

[0151] The data processing device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0152] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes: at least one processor 701 and a memory 702. Optionally, the electronic device 70 further includes a communication component 703. The processor 701, the memory 702 and the communication component 703 are connected via a bus 704.

[0153] During the specific implementation process, at least one processor 701 executes the computer-executable instructions stored in the memory 702, so that the at least one processor 701 performs the above method.

[0154] The specific implementation process of the processor 701 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0155] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules in the processor.

[0156] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0157] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be classified into address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0158] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0159] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0160] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0161] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0162] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.

[0163] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0164] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0165] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0166] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0167] Finally, it should be noted that those skilled in the art will readily identify other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed herein. The present invention is not limited to the precise structure described above and illustrated in the accompanying drawings, and various modifications and variations may be made without departing from the scope thereof. The scope of the present invention is limited solely by the appended claims.

Claims

1. A data processing method, characterized in that: include: Acquire user data to be processed, where the user data to be processed includes user sub-data in multiple dimensions; Determining a target processing strategy corresponding to the user data to be processed based on the data dimension or the preset privacy budget of the user data to be processed, wherein the target processing strategy is used to perform perturbation processing on the user data to be processed based on the attributes and value range of each user sub-data, where the attributes of the user sub-data include categorical attributes or numerical attributes; A corresponding target processing strategy is executed on the user data to be processed to generate disturbed user data.

2. The method according to claim 1, characterized in that Determining a target processing strategy corresponding to the user data to be processed based on the data dimension of the user data to be processed or the preset privacy budget includes: If the data dimension of the user data to be processed is greater than the preset data dimension, or the preset privacy budget of the user data to be processed does not exceed the preset privacy budget threshold, a first processing strategy is determined as the target processing strategy, and the first processing strategy is used to calculate a first probability of each user sub-data in sequence according to the attribute of each user sub-data and the corresponding value range, and perform perturbation processing on the corresponding user sub-data in the user data to be processed according to the first probability; If the data dimension of the user data to be processed does not exceed the preset data dimension, or the preset privacy budget of the user data to be processed is less than the preset privacy budget threshold, the second processing strategy is determined as the target processing strategy; the second processing strategy is used to calculate a second probability of the user data to be processed based on the attributes of each user sub-data and the corresponding value range, and perform perturbation processing on the user data to be processed based on the second probability.

3. The method according to claim 2, characterized in that The first processing strategy includes a first sub-processing strategy for processing the user sub-data of the categorical attribute and a second sub-processing strategy for processing the user sub-data of the numerical attribute; For the first sub-processing strategy, determining the first probability according to a preset sub-privacy budget of the user sub-data and the number of categories within the value range; determining the perturbed user sub-data based on the first probability, the user sub-data, and the value range, where the probability that the perturbed user sub-data is the user sub-data is the first probability, and the probability that the perturbed user sub-data is other data within the value range is a third probability, and the sum of the first probability and the third probability is 1; For the second sub-processing strategy, determining the first probability according to a preset sub-privacy budget of the user sub-data; determining, according to the user sub-data, a first sub-value range from the value range that includes the user sub-data; The disturbed user sub-data is determined based on the first probability, the user sub-data, the value range, and the first sub-value range. The first probability is the probability that the disturbed user sub-data is data within the first sub-value range. The third probability is the probability that the disturbed user sub-data is other data within the second sub-value range. The second sub-value range includes the first sub-value range and the second sub-value range belongs to the value range.

4. The method according to claim 3, characterized in that The first sub-processing strategy, determining the first probability according to a preset sub-privacy budget of the user sub-data and the number of categories within the value range, includes: For the first sub-processing strategy, determining the first probability according to a preset sub-privacy budget of the user sub-data, a preset sub-relaxed privacy parameter, and the number of categories within the value range; Accordingly, for the second sub-processing strategy, determining the first probability according to the preset sub-privacy budget of the user sub-data includes: For the second sub-processing strategy, the first probability is determined according to the preset sub-privacy budget of the user sub-data and the preset sub-relaxed privacy parameter.

5. The method according to claim 3 or 4, characterized in that The second processing strategy includes: For the user sub-data of the classification attribute, determining the proximity probability of the user sub-data according to the number of categories within the value range; For the user sub-data of the numerical attribute, determining the proximity probability of the user sub-data according to the interval length of the first sub-value range; Generating a total neighboring probability of the user data to be processed according to the neighboring probabilities of all user sub-data; generating a second probability of the user data to be processed according to the total neighboring probability and a preset privacy budget; Perturbed user data is generated based on the second probability, where the perturbed user data includes multiple perturbed user sub-data. The perturbed user sub-data is the ratio of the number of the other data to the total number of the perturbed user sub-data, which is the same as the fourth probability; and the sum of the fourth probability and the second probability is 1.

6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: The disturbed user data is shuffled by a shuffler to generate target user data; The target user data is sent to the server.

7. A data processing device, characterized in that: include: An acquisition module, configured to acquire user data to be processed, wherein the user data to be processed includes user sub-data of multiple dimensions; a determination module, configured to determine a target processing strategy corresponding to the user data to be processed based on the data dimension or a preset privacy budget of the user data to be processed, wherein the target processing strategy is configured to perform perturbation processing on the user data to be processed based on the attributes and value range of each user sub-data, wherein the attributes of the user sub-data include categorical attributes or numerical attributes; The generating module is used to execute the corresponding target processing strategy on the user data to be processed to generate disturbed user data.

8. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when the computer program is executed by a processor.