Privacy protection method and device, equipment, storage medium and program product

By adopting local differential privacy mechanisms and deprived projection in user data privacy protection, noise perturbation data is converted into deprived characterization data, and combined with adversarial training public hypothesis detection model and private hypothesis detection model, the problems of high computing overhead, poor real-time performance and difficulty in weighing data availability and privacy security in the existing technology are solved, and efficient and real-time privacy protection is achieved.

CN120234833AInactive Publication Date: 2025-07-01CHINA MOBILE M2M +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510678783.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The privacy protection schemes for user data in the prior art have large computing overhead and poor real-time performance, and it is difficult to weigh data availability and privacy security.

Method used

The original observed data is disturbed based on the local differential privacy mechanism, and the noise perturbation data is converted into deprived characterization data through deprived projection, and sent to the fusion center. The Fusion Center performs hypothesis detection model on user requests by public hypothesis detection model and private hypothesis detection model. The public hypothesis detection model and private hypothesis detection model are obtained through adversarial training.

Benefits of technology

On the basis of ensuring data security, improve data availability and reduce computing overhead, and are suitable for real-time data processing scenarios, and optimize the privacy protection mechanism to weigh the privacy protection security and data availability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234833A_ABST
    Figure CN120234833A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data security, and provides a privacy protection method and device, equipment, a storage medium and a program product, and the method comprises the steps: carrying out the disturbance of collected original observation data based on a local differential privacy mechanism, and obtaining noise disturbance data; converting the noise disturbance data into privacy-removed representation data through privacy-removed projection, and sending the privacy-removed representation data to a fusion center for completing hypothesis detection of a user request; and the fusion center carries out hypothesis detection on privacy-removed representation data through a public hypothesis detection model and a private hypothesis detection model obtained through adversarial training. Through local differential privacy and privacy-removing projection, a confrontation training mechanism fusing a center public hypothesis detection model and a privacy hypothesis detection model is combined, so that a privacy protection mechanism can be optimized, privacy protection and data availability are balanced, the calculation overhead is reduced, and the real-time performance of data processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security technology, and in particular, to a privacy protection method, device, equipment, storage medium, and program product. Background Art

[0002] In existing protection schemes for user privacy data, there are mainly two categories: centralized schemes and decentralized schemes. Among them, the centralized scheme relies on a central server to process and store data, and ensures the privacy of data through encryption and access control. There is a risk of single-point failure, and the trust in the central server is the bottleneck of privacy protection. The decentralized scheme reduces the risk of single-point failure by dispersing data processing tasks among multiple nodes. Currently, existing decentralized privacy protection technologies mainly include differential privacy, homomorphic encryption, and multi-party secure computing, etc. Data privacy protection is achieved through adding noise, encrypted computing, and distributed computing. There are problems such as large computational overhead and poor real-time performance. Moreover, it is difficult to balance data availability and privacy security. Summary of the Invention

[0003] This application provides a privacy protection method, device, equipment, storage medium, and program product to solve the defects in the existing privacy protection scheme for user data, such as large computational overhead, poor real-time performance, and difficulty in balancing data availability and privacy security.

[0004] This application provides a privacy protection method, including: Based on the local differential privacy mechanism, perturb the collected original observation data to obtain noise-perturbed data; Convert the noise-perturbed data into de-privatized representation data through de-privatization projection, and send the de-privatized representation data to a fusion center to complete the hypothesis detection of a user request; The fusion center performs public hypothesis detection on the de-privatized representation data through a public hypothesis detection model to identify the public information in the de-privatized representation data, and performs private hypothesis detection on the de-privatized representation data through a private hypothesis detection model to identify the privacy information in the de-privatized representation data; Wherein, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimization goal of the best detection performance of the public hypothesis detection model and the weakest inference ability of the private hypothesis detection model.

[0005] In one embodiment, the converting the noise-perturbed data into de-privatized representation data through de-privatization projection includes: Obtain a predefined projection matrix; Based on the projection matrix, perform a linear transformation on the noise perturbation data to obtain de-privatized representation data; the projection matrix is used to reduce the dimension of the noise perturbation data through de-privatized projection.

[0006] In one embodiment, the process of adversarial training of the public hypothesis detection model and the private hypothesis detection model includes: Initialize the projection matrix, the first model parameters of the public hypothesis detection model, and the second model parameters of the private hypothesis detection model; Based on the projection matrix, the first model parameters, and the second model parameters, construct an optimization objective; Based on the optimization objective, perform iterative adversarial training on the public hypothesis detection model and the private hypothesis detection model; In each round of the iterative process, alternately optimize the projection matrix, the public hypothesis detection model, and the private hypothesis detection model based on the optimization objective until a preset iterative termination condition is met.

[0007] In one embodiment, the alternately optimizing the projection matrix, the public hypothesis detection model, and the private hypothesis detection model based on the optimization objective includes: Given the projection matrix and the first model parameters of the public hypothesis detection model, optimize the private hypothesis detection model by maximizing the loss value of the first loss function in the optimization objective to obtain the third model parameters of the private hypothesis detection model; Given the projection matrix, based on the third model parameters, optimize the public hypothesis detection model by minimizing the loss value of the second loss function in the optimization objective to obtain the fourth model parameters of the public hypothesis detection model; Based on the third model parameters and the fourth model parameters, optimize the projection matrix by maximizing the loss value of the first loss function and minimizing the loss value of the second loss function.

[0008] In one embodiment, the constructing an optimization objective based on the projection matrix, the first model parameters, and the second model parameters includes: Construct a first loss function based on the projection matrix and the second model parameters; Construct a second loss function based on the projection matrix, the first model parameters, and a first penalty factor; the first penalty factor is used to control the loss weight of the public hypothesis detection model in the optimization objective; Construct a third loss function according to the similarity constraint function of the projection matrix and the second penalty factor; the second penalty factor is used to control the loss weight of the similarity constraint function in the optimization objective; the similarity constraint function is used to represent the similarity distance of the data subset, and the similarity distance includes the intra-class distance and the inter-class distance of the data subset, and the data subset is obtained by projecting through the projection matrix; With the goal of maximizing the loss value of the first loss function and minimizing the loss value of the second loss function, construct an optimization objective based on the first loss function, the second loss function, and the third loss function.

[0009] In one embodiment, the perturbing the collected original observation data based on the local differential privacy mechanism to obtain noise-perturbed data includes: Perform normalization processing on the collected original observation data to obtain normalized data; Generate noise perturbation based on the local differential privacy mechanism; the noise perturbation follows a preset probability distribution function; Add the noise perturbation to the normalized data to obtain noise-perturbed data.

[0010] This application also provides a privacy protection device, including the following modules: A noise perturbation module, configured to perturb the collected original observation data based on the local differential privacy mechanism to obtain noise-perturbed data; A de-privatized projection module, configured to convert the noise-perturbed data into de-privatized characterization data through de-privatized projection, and send the de-privatized characterization data to a fusion center to complete hypothesis detection of a user request; The fusion center performs public hypothesis detection on the de-privatized characterization data through a public hypothesis detection model to identify public information in the de-privatized characterization data, and performs private hypothesis detection on the de-privatized characterization data through a private hypothesis detection model to identify privacy information in the de-privatized characterization data; Wherein, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimization objective of the best detection performance of the public hypothesis detection model and the weakest inference ability of the private hypothesis detection model.

[0011] This application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the privacy protection method as described in any one of the above.

[0012] The present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the privacy protection method described in any one of the above is implemented.

[0013] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the privacy protection method described in any one of the above is implemented.

[0014] The privacy protection method, device, equipment, storage medium and program product provided by the present application use the local differential privacy mechanism to perturb the original observation data. On this basis, the noisy perturbed data is converted into de-privatized characterization data through de-privatization projection and sent to the fusion center. The fusion center performs hypothesis detection on the user request through the public hypothesis detection model and the private hypothesis detection model to complete the user request. The public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training. Through local differential privacy and de-privatization projection, combined with the adversarial training mechanism of the public hypothesis detection model and the private hypothesis detection model in the fusion center, the privacy protection mechanism can be optimized, a trade-off can be made between privacy protection security and data availability, and the computational overhead can be reduced, improving the real-time performance of data processing. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic flowchart of the privacy protection method provided by the embodiments of the present application.

[0017] Figure 2 It is a schematic flowchart of the privacy protection process provided by the embodiments of the present application.

[0018] Figure 3 It is a schematic structural diagram of the privacy protection device provided by the present invention.

[0019] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application. Detailed Embodiments

[0020] To make the objectives, technical solutions, and advantages of this application more clear, the following will clearly and completely describe the technical solutions in this application in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0021] In related privacy protection solutions, differential privacy is a technique that ensures that individual data in a statistical database does not significantly affect query results. By adding noise to query results, differential privacy can prevent attackers from inferring individual data through multiple queries. This method is widely used in the fields of data publishing and analysis. However, due to the addition of noise, the accuracy and availability of data will decrease, especially in high-sensitivity data scenarios, and the computational overhead is large, making it unsuitable for real-time data processing.

[0022] Homomorphic encryption technology allows direct computation on encrypted data without decryption, enabling various operations while maintaining the encrypted state of the data, thereby protecting data privacy. The main application scenarios of homomorphic encryption include cloud computing and distributed computing. However, its computational overhead is relatively high, restricting the feasibility of applications. Moreover, its encryption and decryption processes are complex, consuming a large amount of computing resources, and the implementation is complex, making it difficult to achieve large-scale applications.

[0023] The method of secure multi-party computation (MPC) allows multiple participating parties to jointly compute a function without revealing their respective inputs. This method relies on encryption protocols and computational models and can achieve privacy protection among multiple untrusted participating parties. However, the secure multi-party computation protocol is complex, with high computational and communication overheads and poor real-time performance, making it unsuitable for high-frequency data exchange scenarios.

[0024] In order to improve data security, the above methods generally increase the computational complexity, further increasing the computational overhead, but the data availability will decrease accordingly, making it difficult to achieve a balance between data availability and security.

[0025] Based on this, the embodiments of this application provide a privacy protection method. Through an adversarial training mechanism, the privacy protection mechanism is optimized to improve the accuracy of public hypothesis detection of data, while reducing the detection ability of private hypotheses, and reducing the computational overhead, making it suitable for real-time data processing scenarios. Specifically, Figure 1 is a schematic flowchart of the privacy protection method provided by the embodiments of this application, as Figure 1 shown. This method includes the following steps: Step 100: Based on the local differential privacy mechanism, perturb the collected original observation data to obtain noise-perturbed data; Step 200: Convert the noise-perturbed data into de-privatized representation data through de-privatized projection, and send the de-privatized representation data to the fusion center to complete the hypothesis detection of the user request; The fusion center performs public hypothesis detection on the de-privatized representation data through a public hypothesis detection model to identify the public information in the de-privatized representation data, and performs private hypothesis detection on the de-privatized representation data through a private hypothesis detection model to identify the private information in the de-privatized representation data; Among them, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimization goal of the best detection performance of the public hypothesis detection model and the weakest inference ability of the private hypothesis detection model.

[0026] The privacy protection method provided in the embodiment of the present application is applied to a data processing node, which can be a sensor for collecting original observation data. Specifically, each sensor, as a data processing node, can independently collect original observation data, and these original observation data may contain sensitive information of users, and privacy protection needs to be carried out on it.

[0027] Specifically, the privacy protection mechanism first perturbs the collected original observation data based on the local differential privacy mechanism to obtain noise-perturbed data. Among them, local differential privacy (LDP) is a variant of differential privacy, which realizes privacy protection by adding noise at the source of data collection. The local differential privacy LDP technology is applicable to decentralized data collection scenarios, such as mobile devices and Internet of Things devices. Each data provider independently perturbs its data and then sends the perturbed data to the central server, so as to achieve privacy protection without trusting the central server.

[0028] Convert the noise-perturbed data into de-privatized representation data through de-privatized projection, and send the de-privatized representation data to the fusion center to complete the hypothesis detection of the user request. De-privatized data projection is a technology that maps the original data to a low-dimensional privacy protection space through linear transformation. It combines differential privacy and dimensionality reduction methods, and can retain a certain degree of data availability, which can reduce the computational overhead of transmission and storage while protecting data privacy.

[0029] When the noise intensity is relatively high in local differential privacy, the data availability will be significantly reduced. Moreover, in the process of de-privatizing data projection, useful information may be lost, affecting the subsequent data availability and the accuracy of data analysis. In this embodiment, based on the local differential privacy mechanism, combined with de-privatizing data projection, it is possible to improve data availability and reduce computational overhead while ensuring data security.

[0030] Furthermore, the fusion center performs public hypothesis detection on the de-privatized representation data through a public hypothesis detection model to identify the public information in the de-privatized representation data, and performs private hypothesis detection on the privatized representation data through a private hypothesis detection model, thereby identifying the private information in the de-privatized representation data. Among them, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training, and during the adversarial training process, the optimization goal is that the detection performance of the public hypothesis detection model is optimal while the inference ability of the private hypothesis detection model is the weakest.

[0031] The hypothesis detection of the user request by the fusion center is realized through the hypothesis detection model obtained by the adversarial training mechanism, and the hypothesis detection model includes a public hypothesis detection model and a private hypothesis detection model. Through adversarial training, the accuracy of the public hypothesis detection of the data is improved, while the detection ability of the private hypothesis is reduced, the data security of the private information is improved, and the privacy protection mechanism is optimized.

[0032] In this embodiment, through the local differential privacy mechanism, the original observation data is perturbed, and on this basis, the noise-perturbed data is converted into de-privatized representation data through de-privatizing projection and sent to the fusion center. The fusion center performs hypothesis detection on the user request through the public hypothesis detection model and the private hypothesis detection model, so as to complete the user request. The public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training. Through local differential privacy and de-privatizing projection, combined with the adversarial training mechanism of the public hypothesis detection model and the private hypothesis detection model of the fusion center, a balance is made between privacy protection security and data availability, and the computational overhead is reduced, improving the real-time performance of data processing.

[0033] In one embodiment, referring to Figure 2 the privacy protection process shown, for the original observation data collected by the data processing node, first, noise perturbation based on the local differential privacy mechanism is added to the original observation data of the sensor, and then, through de-privatizing projection, the noise-perturbed data x is converted into de-privatized representation data z. Operations such as noise perturbation and de-privatizing projection in the privacy protection mechanism are independently completed locally by each data processing node, and the de-privatized representation data z is sent to the fusion center to complete the hypothesis detection of the user request.

[0034] Furthermore, the fusion center separately performs public hypothesis detection and private hypothesis detection on the de-privatized representation data z. Among them, the public hypothesis detection is implemented based on a public hypothesis detection model, which is used to identify the public information of the de-privatized representation data z and belongs to an availability classifier. The private hypothesis detection is implemented based on a private hypothesis detection model, which belongs to an inference privacy classifier.

[0035] Optionally, locally at the data processing node, noise perturbation is added to the original observation data. First, local differential privacy is applied, and the added noise perturbation conforms to a corresponding probability distribution. Based on this, step 100 may further include: Step 101, performing normalization processing on the collected original observation data to obtain normalized data; Step 102, generating noise perturbation based on the local differential privacy mechanism; the noise perturbation follows a preset probability distribution function; Step 103, adding the noise perturbation to the normalized data to obtain noise-perturbed data.

[0036] Before performing the privacy protection transformation, first preprocess the original observation data, and then perturb the preprocessed data to ensure the privacy and accuracy of the data. To determine the main parameters of the probability distribution function followed by the added noise, first, the original observation data needs to be normalized so that the value range of the observed values of the original observation data is between [0, 1].

[0037] Based on the local differential privacy mechanism, generate noise perturbation, which follows a preset probability distribution function, and add the generated noise perturbation to the normalized data to obtain noise-perturbed data.

[0038] Among them, the original observation data can be normalized in the manner shown in the following formula 1: ; (1) ; (1) Add noise b to the normalized observed data to achieve local privacy differentiation. The noise-perturbed data after adding the noise perturbation can be expressed as: .

[0039] Optionally, represents the observed value of the th data point in the noise-perturbed data. In one embodiment, the noise b follows a Laplace distribution , and its corresponding probability density function can be expressed as: ; (2) is a scale parameter, which is used to determine the shape of the Laplace distribution. Through local differential privacy, it ensures that data processing nodes complete privacy protection locally, reducing the trust dependence on the central server.

[0040] The de-privatization projection of the noise-perturbed data is implemented based on a projection matrix. Based on this, in step 200, the noise-perturbed data is converted into de-privatized characterization data through de-privatization projection, and it may further include: Step 201, obtain a predefined projection matrix; Step 202, based on the projection matrix, perform a linear transformation on the noise-perturbed data to obtain de-privatized characterization data; the projection matrix is used to reduce the dimension of the noise-perturbed data through de-privatization projection.

[0041] Obtain a predefined projection matrix. Based on this projection matrix, perform a linear transformation on the noise-perturbed data to obtain de-privatized characterization data, where the projection matrix is used to reduce the dimension of the noise-perturbed data through de-privatization projection.

[0042] Perform a linear transformation on the perturbed data through de-privatization projection, thereby achieving a balance between data privacy protection and data availability. During the privacy protection process, adding noise perturbation can achieve privacy protection, but it will lead to a reduction in data availability. Through de-privatization projection, the data with added noise perturbation is mapped to a low-dimensional privacy protection space, which can enhance the privacy of the data while retaining the key features of the data.

[0043] For the predefined projection matrix, it is used to map the high-dimensional noise-perturbed data to a low-dimensional space. Exemplarily, if it is assumed that has a dimension of ( represents the number of data points in the observed data), the projection matrix has a dimension of , where . Through the projection matrix perform a linear transformation on the perturbed noise-perturbed data to obtain de-privatized characterization data , then there is .

[0044] Compress the high-dimensional de-privatized characterization data to a low-dimensional space through linear transformation. The projection matrix It has the following characteristics: in terms of privacy protection, the projected data will not disclose the sensitive information of the original observed data; in terms of data availability, the projected data can retain the main features of the original observed data so that the fusion center can perform effective hypothesis detection; in terms of computational efficiency, the projection process has high computational performance and is suitable for the real-time processing requirements of data.

[0045] Each data processing node sends the generated de-privatized characterization data to the fusion center. During this process, since the data dimension of the de-privatized characterization data is smaller than that of the original observed data, therefore, transmitting the de-privatized characterization data can reduce the data transmission volume between the data processing node and the central server of the fusion center, thereby reducing the communication cost. The fusion center only receives the de-privatized characterization data and does not need to receive the original observed data or the noise-perturbed data , thus enhancing the privacy protection effect.

[0046] The fusion center has two detection mechanisms for the de-privatized characterization data, public hypothesis detection and private hypothesis detection. The two hypothesis detection mechanisms use the de-privatized characterization data to achieve different detection tasks, ensuring the privacy and availability of the data.

[0047] The two hypothesis detection mechanisms are respectively implemented based on different detection models. Among them, public hypothesis detection is implemented based on the public hypothesis detection model, and private hypothesis detection is implemented based on the private hypothesis detection model. The public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training.

[0048] Optionally, the fusion center receives the de-privatized characterization data sent by each data processing node and summarizes and integrates the received de-privatized characterization data to form a comprehensive de-privatized characterization matrix .

[0049] Public hypothesis detection is used to identify the public information in the de-privatized characterization data, ensuring the correct detection and utilization of public hypothesis data under the premise of protecting privacy. Specifically, the fusion center performs public hypothesis detection on the de-privatized characterization matrix through the public hypothesis detection model, and the detection method can be expressed as: . The public hypothesis detection model is constructed based on a pre-trained machine learning model , and the machine learning model is responsible for performing public hypothesis detection according to the input de-privatized characterization matrix to obtain the prediction status of the public hypothesis . The detection result of the public hypothesis detection model Indicates the predicted state of the public hypothesis, such as whether there is an anomaly, the probability of an event occurring, etc. The fusion center will feed back the detection result to the service requester to complete the hypothesis detection of the user request.

[0050] In this embodiment, the public hypothesis detection has advantages such as real-time and high efficiency. For real-time, based on the low dimension of the de-identified representation data, the fusion center can quickly process and detect the de-identified representation data, providing real-time public hypothesis detection results; for high efficiency, through de-identified projection, the complexity of data transmission and processing is reduced, improving the overall efficiency of hypothesis detection.

[0051] Furthermore, the private hypothesis detection is used to identify and protect the privacy information in the original observation data, preventing untrusted fusion centers or malicious parties from inferring the user's privacy information by analyzing the de-identified representation data. The private hypothesis detection model is used to infer and detect the true state of the private hypothesis, and the inference method can be expressed as: . The private hypothesis detection model is also based on a pre-trained machine learning model constructed, and the machine learning model is used to analyze the de-identified representation matrix and attempt to infer the true state of the private hypothesis .

[0052] To prevent an untrusted fusion center from inferring the true state of the private hypothesis through the private hypothesis detection model, an adversarial training mechanism is adopted. The goal of adversarial training is to optimize the de-identified projection matrix and the private hypothesis detection model, so that the detection performance of the public hypothesis detection model is optimal, while the inference ability of the private hypothesis detection model is the weakest.

[0053] Optionally, the process of adversarial training of the public hypothesis detection model and the private hypothesis detection model specifically includes the following steps: Step 301, initialize the projection matrix, the first model parameters of the public hypothesis detection model, and the second model parameters of the private hypothesis detection model; Step 302, construct an optimization objective based on the projection matrix, the first model parameters, and the second model parameters; Step 303, perform iterative adversarial training on the public hypothesis detection model and the private hypothesis detection model based on the optimization objective; in each round of iteration, alternately optimize the projection matrix, the public hypothesis detection model, and the private hypothesis detection model based on the optimization objective until the preset iteration termination condition is met.

[0054] First, initialize the projection matrix and the model parameters of the hypothesis detection model. The model parameters include the first model parameters of the public hypothesis detection model and the second model parameters of the private hypothesis detection model. The initialization of the projection matrix and the model parameters can be obtained by random initialization or using a pre-trained model.

[0055] Further, based on the initialized projection matrix, the first model parameters, and the second model parameters, construct an optimization objective for adversarial training. Based on the constructed optimization objective, perform iterative adversarial training on the public hypothesis detection model and the private hypothesis detection model until a preset iterative termination condition is met. In each round of the iterative process, alternately optimize the projection matrix, the first model parameters, and the second model parameters based on the constructed optimization objective. The iterative termination condition can be reaching a preset number of iterations or the model converging.

[0056] In one embodiment, denote the first model parameters of the public hypothesis detection model as and denote the second model parameters of the private hypothesis detection model as . The optimization objective of the adversarial training is to minimize the loss of the public hypothesis detection model and maximize the loss of the private hypothesis detection model. Based on this, step 302 may further include: Step 312, construct a first loss function based on the projection matrix and the second model parameters; Step 322, construct a second loss function based on the projection matrix, the first model parameters, and a first penalty factor; the first penalty factor is used to control the loss weight of the public hypothesis detection model in the optimization objective; Step 332, construct a third loss function according to the similarity constraint function of the projection matrix and a second penalty factor; the second penalty factor is used to control the loss weight of the similarity constraint function in the optimization objective; the similarity constraint function is used to represent the similarity distance of the data subset, and the similarity distance includes the within-class distance and the between-class distance of the data subset, and the data subset is obtained by projecting through the projection matrix; Step 342, with the goal of maximizing the loss value of the first loss function and minimizing the loss value of the second loss function, construct an optimization objective based on the first loss function, the second loss function, and the third loss function.

[0057] When constructing the optimization objective, based on the projection matrix and the second model parameters of the private hypothesis detection model , construct the first loss function of the private hypothesis detection model, and construct the second loss function based on the projection matrix, the first model parameters of the public hypothesis detection model, and a preset first penalty factor, where the first penalty factor is used to control the loss weight of the public hypothesis detection model in the optimization objective.

[0058] Furthermore, construct the third loss function according to the similarity constraint function of the projection matrix and a second penalty factor, where the second penalty factor is used to control the loss weight of the similarity constraint function in the optimization objective. The similarity constraint function is used to represent the similarity distance of the data subset, and the similarity distance includes the within-class distance and between-class distance of the data subset. The data subset is obtained by projecting through the projection matrix.

[0059] With the goal of maximizing the loss value of the first loss function and minimizing the loss value of the second loss function, construct an optimization objective based on the first loss function, the second loss function, and the third loss function. Optionally, the constructed optimization objective is as shown in formula (3) below: ; (3) where, is the loss function of the private hypothesis detection model, is the loss function of the public hypothesis detection model, and are penalty factors, is the similarity constraint function of the projection matrix.

[0060] The goal of adversarial training is to optimize both the public hypothesis detection model and the private hypothesis detection model simultaneously, maximizing the accuracy of public hypothesis detection while minimizing the inference ability of private hypothesis detection. That is, the goal of the adversarial training mechanism is to minimize the empirical loss function of public hypothesis detection and the regularized empirical risk loss function of private hypothesis detection. The optimization objective can be expressed as an adversarial optimization problem, by adjusting the projection matrix and the parameters of the hypothesis detection model and , to achieve the optimization objective shown in formula (3). Among them, is the empirical loss function of public hypothesis detection, measuring the detection performance of the public hypothesis detection model under the given projection matrix . is the regularized empirical risk loss function of private hypothesis detection, measuring the performance of the private hypothesis detection model under the given projection matrix. is the similarity constraint of the projection matrix, ensuring that the projected data can meet certain privacy protection and data availability requirements. and are penalty factors, used to control the weights of the public hypothesis detection loss and the similarity constraint in the overall optimization objective.

[0061] Based on the constructed optimization objective, an iterative training method is adopted. In each round of iteration, the projection matrix, the public hypothesis detection model, and the private hypothesis detection model are alternately optimized. Specifically, in step 303, the alternate optimization based on the projection matrix, the public hypothesis detection model, and the private hypothesis detection model may include: Step 313, given the projection matrix and the first model parameters of the public hypothesis detection model, optimize the private hypothesis detection model by maximizing the loss value of the first loss function of the private hypothesis detection model in the optimization objective, to obtain the third model parameters of the private hypothesis detection model; Step 323, given the projection matrix, based on the third model parameters, optimize the public hypothesis detection model by minimizing the loss value of the second loss function of the public hypothesis detection model in the optimization objective, to obtain the fourth model parameters of the public hypothesis detection model; Step 333, based on the third model parameters and the fourth model parameters, optimize the projection matrix by maximizing the loss value of the first loss function and minimizing the loss value of the second loss function.

[0062] The alternate optimization of the projection matrix, the public hypothesis detection model, and the private hypothesis detection model is carried out in the order of the private hypothesis detection model, the public hypothesis detection model, and the projection matrix.

[0063] Specifically, based on the initialized projection matrix, the first model parameters of the public hypothesis detection model, and the second model parameters of the private hypothesis detection model, given the projection matrix, optimize the private hypothesis detection model by maximizing the loss value of the first loss function of the private hypothesis detection model in the optimization objective, to obtain the third model parameters of the private hypothesis detection model.

[0064] That is, given the projection matrix and the first model parameters of the public hypothesis detection model , optimize the private hypothesis detection model, specifically optimize the second model parameters of the private hypothesis detection model , by maximizing the loss function of the private hypothesis detection model to obtain the optimized third model parameters and improve the inference ability of the private hypothesis detection model .

[0065] Furthermore, after optimizing the private hypothesis detection model, fix the projection matrix and the third model parameters of the private hypothesis detection model, and optimize the public hypothesis detection model. Specifically, by minimizing the function value of the second loss function of the public hypothesis detection model in the optimization objective, optimize the first model parameters of the public hypothesis detection model , obtain the fourth model parameter of the public hypothesis detection model.

[0066] By minimizing the loss function of the public hypothesis detection model , improve the accuracy of the public hypothesis detection model .

[0067] After optimizing the public hypothesis detection model and the private hypothesis detection model, fix and , optimize the projection matrix , and optimize the projection matrix by maximizing the loss value of the first loss function and minimizing the loss value of the second loss function.

[0068] That is, based on the third model parameter and the fourth model parameter obtained after optimization, by maximizing the loss value of the first loss function and minimizing the loss value of the second loss function, while considering the similarity constraint function of the projection matrix , achieve the optimization goal of adversarial training: ; (4) Iterate this alternating optimization process until the model converges or reaches the preset number of iterations, gradually improving the effect of adversarial training.

[0069] Among them, the similarity constraint function of the projection matrix is used to ensure that the projected data meets the requirements of privacy protection and data availability. The similarity constraint function can measure the within-class and between-class distances of the projected data, making the data points of different classes as separated as possible in the projection space, while the data points of the same class are as close as possible.

[0070] Exemplarily, using , , and to represent data subsets of different classes, represents two different data subsets and between the distances, then there are: ; (5) The penalty factors and are used to control the loss weights in the optimization goal. By adjusting the values of the penalty factors and , a balance can be achieved between the public hypothesis detection performance and the private hypothesis detection protection, ensuring that the projection matrix maximizes the data availability to the greatest extent under the premise of meeting privacy protection.

[0071] Through the adversarial training mechanism, a dynamic balance between the privacy protection mechanism and hypothesis detection is achieved. By optimizing the projection matrix and the model parameters of the hypothesis detection model, the accuracy of the public hypothesis detection model is maximized, and the inference ability of the private hypothesis detection model is minimized. Thus, while protecting user privacy, data detection and utilization can be efficiently carried out.

[0072] Through adversarial training, it is ensured that the de-privatized representation data remains efficient in public hypothesis detection, while minimizing the risk of private hypotheses being inferred. The optimized de-privatized projection matrix and hypothesis detection model can better cope with different privacy protection requirements and attack risks, improving the robustness and security of the system.

[0073] The combination of public hypothesis detection and private hypothesis detection can not only effectively utilize the de-privatized representation data to detect public hypotheses, but also protect private hypotheses through the adversarial training mechanism, preventing untrusted fusion centers from inferring users' sensitive information. The two complement each other to ensure maximum protection of user privacy while providing efficient data utilization.

[0074] In this embodiment, by combining the adversarial training mechanism with local differential privacy and de-privatized projection, a trade-off between privacy protection and data availability is achieved. By optimizing the privacy protection mechanism and detection performance, not only the accuracy of public hypothesis detection of data is improved, but also the computational overhead is significantly reduced, enhancing the data protection ability of the fusion center in the face of malicious nodes.

[0075] Next, the privacy protection device provided by the embodiments of the present application will be described. The privacy protection device described below can be mutually corresponding and referred to the privacy protection method described above.

[0076] Referring to Figure 3 , the privacy protection device provided by the embodiments of the present application includes: A noise perturbation module 10, configured to perturb the collected original observation data based on the local differential privacy mechanism to obtain noise-perturbed data; A de-privatized projection module 20, configured to convert the noise-perturbed data into de-privatized representation data through de-privatized projection, and send the de-privatized representation data to a fusion center to complete hypothesis detection of user requests; The fusion center performs public hypothesis detection on the de-privatized representation data through a public hypothesis detection model to identify public information in the de-privatized representation data, and performs private hypothesis detection on the de-privatized representation data through a private hypothesis detection model to identify privacy information in the de-privatized representation data; Among them, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimization goal of maximizing the detection performance of the public hypothesis detection model and minimizing the inference ability of the private hypothesis detection model.

[0077] In one embodiment, the de-privatization projection module 20 is further configured to: Obtain a predefined projection matrix; Based on the projection matrix, perform a linear transformation on the noise-perturbed data to obtain de-privatized characterization data; the projection matrix is used to reduce the dimension of the noise-perturbed data through de-privatization projection.

[0078] In one embodiment, the noise perturbation module 10 is further configured to: Normalize the collected original observation data to obtain normalized data; Generate noise perturbation based on the local differential privacy mechanism; the noise perturbation follows a preset probability distribution function; Add the noise perturbation to the normalized data to obtain noise-perturbed data.

[0079] In one embodiment, the privacy protection device further includes an adversarial training module, which is configured to: Initialize the projection matrix, the first model parameters of the public hypothesis detection model, and the second model parameters of the private hypothesis detection model; Based on the projection matrix, the first model parameters, and the second model parameters, construct an optimization goal; Based on the optimization goal, perform iterative adversarial training on the public hypothesis detection model and the private hypothesis detection model; In each round of iteration, alternately optimize the projection matrix, the public hypothesis detection model, and the private hypothesis detection model based on the optimization goal until a preset iteration termination condition is met.

[0080] In one embodiment, the adversarial training module is further configured to: Given the projection matrix and the first model parameters of the public hypothesis detection model, optimize the private hypothesis detection model by maximizing the loss value of the first loss function of the private hypothesis detection model in the optimization goal to obtain the third model parameters of the private hypothesis detection model; Given the projection matrix, based on the third model parameters, optimize the public hypothesis detection model by minimizing the loss value of the second loss function of the public hypothesis detection model in the optimization goal to obtain the fourth model parameters of the public hypothesis detection model; Based on the third model parameter and the fourth model parameter, optimize the projection matrix by maximizing the loss value of the first loss function and minimizing the loss value of the second loss function.

[0081] In one embodiment, the adversarial training module is further configured to: Construct a first loss function based on the projection matrix and the second model parameter; Construct a second loss function based on the projection matrix, the first model parameter, and a first penalty factor; the first penalty factor is used to control the loss weight of the public hypothesis detection model in the optimization objective; Construct a third loss function according to the similarity constraint function of the projection matrix and a second penalty factor; the second penalty factor is used to control the loss weight of the similarity constraint function in the optimization objective; the similarity constraint function is used to represent the similarity distance of the data subset, and the similarity distance includes the intra-class distance and the inter-class distance of the data subset, and the data subset is obtained by projecting through the projection matrix; With the goal of maximizing the loss value of the first loss function and minimizing the loss value of the second loss function, construct an optimization objective based on the first loss function, the second loss function, and the third loss function.

[0082] Figure 4 An example of the physical structure diagram of an electronic device is shown as Figure 4 shown. The electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete mutual communication through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the privacy protection method, and the method includes: Based on the local differential privacy mechanism, perturb the collected original observation data to obtain noise-perturbed data; Convert the noise-perturbed data into de-privatized characterization data through de-privatized projection, and send the de-privatized characterization data to the fusion center to complete the hypothesis detection of the user request; The fusion center performs public hypothesis detection on the de-privatized characterization data through a public hypothesis detection model to identify the public information in the de-privatized characterization data, and performs private hypothesis detection on the de-privatized characterization data through a private hypothesis detection model to identify the privacy information in the de-privatized characterization data; Among them, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimization goal of maximizing the detection performance of the public hypothesis detection model and minimizing the inference ability of the private hypothesis detection model.

[0083] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0084] On the other hand, an embodiment of this application also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the privacy protection method provided by the above-mentioned various methods. The method includes: Based on the local differential privacy mechanism, perturb the collected original observation data to obtain noise-perturbed data; Convert the noise-perturbed data into de-privatized characterization data through de-privatized projection, and send the de-privatized characterization data to the fusion center to complete the hypothesis detection of the user request; The fusion center performs public hypothesis detection on the de-privatized characterization data through a public hypothesis detection model to identify the public information in the de-privatized characterization data, and performs private hypothesis detection on the de-privatized characterization data through a private hypothesis detection model to identify the privacy information in the de-privatized characterization data; Among them, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimization goal of maximizing the detection performance of the public hypothesis detection model and minimizing the inference ability of the private hypothesis detection model.

[0085] On yet another aspect, an embodiment of this application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the privacy protection method provided by the above-mentioned various methods. The method includes: Based on the local differential privacy mechanism, the original observed data collected is perturbed to obtain noise-perturbed data; The noise-perturbed data is converted into de-privatized characterization data through de-privatization projection, and the de-privatized characterization data is sent to the fusion center to complete the hypothesis detection of the user request; The fusion center performs public hypothesis detection on the de-privatized characterization data through a public hypothesis detection model to identify the public information in the de-privatized characterization data, and performs private hypothesis detection on the de-privatized characterization data through a private hypothesis detection model to identify the private information in the de-privatized characterization data; Among them, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimization goal of the best detection performance of the public hypothesis detection model and the weakest inference ability of the private hypothesis detection model.

[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0088] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A privacy protection method, characterized in that, Including: Based on the local differential privacy mechanism, perturb the collected original observation data to obtain noise-perturbed data; Convert the noise-perturbed data into de-privatized representation data through de-privatization projection, and send the de-privatized representation data to the fusion center to complete the hypothesis detection of the user request; The fusion center performs public hypothesis detection on the de-privatized representation data through a public hypothesis detection model to identify the public information in the de-privatized representation data, and performs private hypothesis detection on the de-privatized representation data through a private hypothesis detection model to identify the private information in the de-privatized representation data; Among them, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimization goal of the best detection performance of the public hypothesis detection model and the weakest inference ability of the private hypothesis detection model.

2. The privacy protection method according to claim 1, wherein The converting the noise-perturbed data into de-privatized representation data through de-privatization projection includes: Obtain a predefined projection matrix; Based on the projection matrix, perform a linear transformation on the noise-perturbed data to obtain de-privatized representation data; the projection matrix is used to reduce the dimension of the noise-perturbed data through de-privatization projection.

3. The privacy protection method according to claim 2, wherein The process of adversarial training of the public hypothesis detection model and the private hypothesis detection model includes: Initialize the projection matrix, the first model parameters of the public hypothesis detection model, and the second model parameters of the private hypothesis detection model; Based on the projection matrix, the first model parameters, and the second model parameters, construct an optimization goal; Based on the optimization goal, perform iterative adversarial training on the public hypothesis detection model and the private hypothesis detection model; In each round of iteration, alternately optimize the projection matrix, the public hypothesis detection model, and the private hypothesis detection model based on the optimization goal until the preset iteration termination condition is met.

4. The privacy protection method according to claim 3, wherein The alternately optimizing the projection matrix, the public hypothesis detection model, and the private hypothesis detection model based on the optimization goal includes: Given the projection matrix and the first model parameters of the public hypothesis detection model, optimize the private hypothesis detection model by maximizing the loss value of the first loss function of the private hypothesis detection model in the optimization goal to obtain the third model parameters of the private hypothesis detection model; Given the projection matrix, based on the third model parameters, optimize the public hypothesis detection model by minimizing the loss value of the second loss function of the public hypothesis detection model in the optimization goal to obtain the fourth model parameters of the public hypothesis detection model; Based on the third model parameters and the fourth model parameters, optimize the projection matrix by maximizing the loss value of the first loss function and minimizing the loss value of the second loss function.

5. The privacy protection method according to claim 3, wherein The constructing an optimization goal based on the projection matrix, the first model parameters, and the second model parameters includes: Construct a first loss function based on the projection matrix and the second model parameters; Construct a second loss function based on the projection matrix, the first model parameter, and the first penalty factor; the first penalty factor is used to control the loss weight of the public hypothesis detection model in the optimization objective; Construct a third loss function according to the similarity constraint function of the projection matrix and the second penalty factor; the second penalty factor is used to control the loss weight of the similarity constraint function in the optimization objective; the similarity constraint function is used to represent the similarity distance of the data subset, and the similarity distance includes the within-class distance and the between-class distance of the data subset, and the data subset is obtained by projecting through the projection matrix; With the goal of maximizing the loss value of the first loss function and minimizing the loss value of the second loss function, construct an optimization objective based on the first loss function, the second loss function, and the third loss function.

6. The privacy protection method according to claim 1, wherein The perturbing the collected original observation data based on the local differential privacy mechanism to obtain noise-perturbed data includes: Performing normalization processing on the collected original observation data to obtain normalized data; Generating noise perturbation based on the local differential privacy mechanism; the noise perturbation follows a preset probability distribution function; Adding the noise perturbation to the normalized data to obtain noise-perturbed data.

7. A privacy protection device, characterized in that, Includes: A noise perturbation module, configured to perturb the collected original observation data based on the local differential privacy mechanism to obtain noise-perturbed data; A de-privatized projection module, configured to convert the noise-perturbed data into de-privatized characterization data through de-privatized projection, and send the de-privatized characterization data to the fusion center to complete the hypothesis detection of the user request; The fusion center performs public hypothesis detection on the de-privatized characterization data through a public hypothesis detection model to identify the public information in the de-privatized characterization data, and performs private hypothesis detection on the de-privatized characterization data through a private hypothesis detection model to identify the private information in the de-privatized characterization data; Wherein, the public hypothesis detection model and the private hypothesis detection model are obtained through adversarial training with the optimal detection performance of the public hypothesis detection model and the weakest inference ability of the private hypothesis detection model as the optimization objective.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the privacy protection method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the privacy protection method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the privacy protection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Distributed zero-order dual average online optimization method and system capable of protecting privacy

    CN117118723A

  • System and method for DNN-based cyber-security using federated learning-based generative adversarial network

    US20230308465A1

  • Systems, Methods, And Devices to Curate and Present Content and Physical Elements Based on Personal Biometric Identifier Information

    US20240361827A1