A personalized k-anonymity optimization method based on game theory

By employing a personalized k-anonymity optimization method based on game theory, this paper addresses the issues of inappropriate k-value selection and user assistance selection in location privacy protection within the K-anonymity algorithm, achieving optimal service quality and privacy protection under optimal attacker attack conditions.

CN114844665BActive Publication Date: 2026-03-27GUIZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-12
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing K-anonymity algorithms, when implementing location privacy protection, may lead to location information leakage or poor service quality due to improper selection of the k value and assisting users, making it difficult to achieve a balance between service quality and privacy protection.

Method used

A personalized k-anonymity optimization method based on game theory is adopted. By selecting the optimal k value through location privacy measurement, service quality measurement and game theory, the user reward function is optimized to achieve the best service quality and personalized privacy protection.

Benefits of technology

In cases where attackers launch optimal attacks based on prior knowledge, optimal service quality and personalized privacy protection are achieved, reducing the risk of privacy leaks and improving service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114844665B_ABST
    Figure CN114844665B_ABST
Patent Text Reader

Abstract

The application designs a personalized k-anonymity optimization method based on game theory, aiming to realize personalized location privacy protection while ensuring users to obtain higher service quality. Compared with the traditional location privacy protection method, the method can realize the optimal service quality and personalized privacy protection in the case that an attacker initiates an optimal attack based on prior knowledge. The technical key points are that the location privacy metric and the service quality metric are proposed based on the personalized k-anonymity concept and the attack model of malicious users, and on this basis, the k-anonymity algorithm of optimal privacy and service quality is realized through game, so as to solve the optimal k value. According to the scene where the user is located and the anonymous mechanism used, the current optimal k value is given, and the best privacy protection and the highest utility are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical fields:

[0001] This invention belongs to the field of information security technology and involves knowledge of information entropy and game theory; Background technology:

[0002] The widespread use of smart mobile devices with continuous internet connectivity has spurred the development of various location-based services (LBSs). LBSs are services built around geographic location data. They are provided by mobile terminals using wireless communication networks (or satellite positioning systems) and spatial databases to obtain users' geographic location coordinates and integrate them with other information to provide users with location-related value-added services. In 1996, the US FCC issued an executive order, E911, which enabled the tracking of users' locations via wireless signals at any time and place, with an accuracy of 50-300 meters. This opened up the commercial application of LBSs. Subsequently, countries around the world deployed commercial LBSs and applied them to various aspects of life, such as social networking, ride-hailing, and dining. After decades of development, LBSs have experienced remarkable growth and continue to this day. The emergence of LBSs has provided users with great convenience, but the services obtained by users come at the cost of sacrificing user privacy.

[0003] Much work has focused on location privacy protection mechanisms (LPPMs), allowing users to access services through LBS while protecting sensitive information. These mechanisms increase adversaries' uncertainty about a user's true location by hiding their location from LBS or sending distorted or fake locations. Existing methods for protecting query content privacy can be broadly categorized into distortion-based, encryption-based, differential privacy-based, and anonymity-based techniques. Anonymity-based techniques offer advantages such as reduced privacy risks due to anonymous users, user-defined privacy levels, and strong algorithm portability. K-anonymity, proposed by Samarati and Sweeney in 1998, ensures that each individual record in the published dataset is indistinguishable from the other K-1 individuals for sensitive attributes. To quantify the effectiveness of several location privacy protection mechanisms, virtual location and differential privacy have proposed the expected distance error method, which has been the de facto standard for measuring location privacy. However, Shokri et al. asserted that the expected estimation error, lacking accuracy, affects privacy levels, and therefore proposed... A systematic approach to quantify user location privacy is proposed. This approach models location-based applications and privacy protection mechanisms while considering attacker models. To better align with real-world applications, some studies have made assumptions about the attacker's background knowledge. For example, Kido et al. considered factors such as universality, congestion, and consistency when generating fake locations, attempting to make these virtual locations as realistic as possible. Similarly, Xue et al. used side-channel information from mobile users to construct anonymized stealth regions. As location privacy gains increasing attention, the K-anonymity algorithm is being used more widely. However, current K-anonymity algorithms still suffer from problems such as location information leakage and low service quality due to improper selection of the k-value and assisting users. When the k-value is too high, too many users participate in constructing the anonymity domain, reducing the service quality for the requesting user. When the k-value is too low, too few assisting users increase the risk of privacy information leakage for the requesting user. Furthermore, the assisting users involved in constructing the anonymity domain also affect the service quality and privacy protection for the requesting user. Therefore, achieving a balance between service quality and privacy protection, calculating the optimal k-value, and selecting appropriate assisting users are crucial issues. Summary of the Invention:

[0004] To address the aforementioned problems in location privacy protection through k-anonymity, this invention proposes a personalized k-anonymity optimization method based on game theory. This method aims to achieve optimal service quality and personalized privacy protection even when an attacker launches an optimal attack based on prior knowledge. The method includes steps such as location privacy measurement, service quality measurement, and game theory selection of the optimal k value. The specific process is as follows:

[0005] 1) The mathematical model for location privacy measurement, privacy sink Z′, and background knowledge Z is as follows:

[0006]

[0007]

[0008] The privacy source entropy H(Z′) is defined as follows:

[0009]

[0010] Using privacy conditional entropy H(z′) r |Z) represents the inference of the requesting user's location z′ given that the attacker has obtained background knowledge Z and the private destination Z′. r The degree of uncertainty;

[0011]

[0012] Using I(z′) r Z) represents the degree of privacy leakage when an attacker uses background knowledge to carry out an optimal attack.

[0013]

[0014] 2) Service quality measurement: Based on a given anonymous set Z′ and anonymous machine mechanism f, an attacker can use the standard Bayesian formula to calculate the posterior probability of each position in Z′ and launch an inference attack. The inferred position obtained by the attacker is the position with the highest posterior probability.

[0015]

[0016] The attacker's expected error is as follows:

[0017] EE(f)=1-maxPr(z′) i |Z′)

[0018] The service quality is calculated as follows:

[0019]

[0020] Step 3: Game theory to select the optimal value of k, given the user's payoff function as E. u Based on the principle of maximizing benefits, users will maximize their benefit function;

[0021]

[0022]

[0023] The following conditions must be met for this expression to be valid:

[0024] η+θ=1

[0025]

[0026]

[0027] Therefore, we derive the following formula for the r value that provides users with optimal service quality and location privacy:

[0028]

[0029] To demonstrate the effectiveness of this invention, security analysis and simulation experiments show that this invention can provide a k-value that optimizes service quality and location privacy based on different user scenarios and different anonymity mechanisms used. Furthermore, this invention can reduce privacy leakage to a low level when the scenario is favorable and the anonymity mechanism is well protected. Attached Figure Description

[0030] Figure 1 The values ​​of r for different privacy requirements in different scenarios are described in detail.

[0031] Figure 2 The relationship between privacy requirements and r-values ​​under different anonymity mechanisms is described in detail;

[0032] Figure 3 It describes in detail how the degree of privacy leakage decreases as the value of r increases in different scenarios;

[0033] Figure 4 The study details how the degree of privacy leakage decreases as the value of r increases under different anonymization mechanisms. Detailed implementation method:

[0034] The technical solution of this invention will be clearly and completely described below with reference to the accompanying drawings. This invention provides a personalized k-anonymous optimization method based on game theory, and the specific steps are as follows:

[0035] Step 1: Location privacy measurement. Assume the requesting user is the information sender, the attacker is the receiver, and the privacy leakage channel is the communication channel. The information set possessed by the sender is called the privacy source, referring to the anonymous domain constructed by the requesting user. The information set obtained by the receiver is called the privacy sink, referring to the anonymous domain observed by the attacker, which is the same as the privacy source, i.e., the anonymous domain Z′={z′1,z′2,...,z′...}. r}, where z′ j (j = 1, 2, ..., r) represents a user's location privacy information, and the attacker's background knowledge is the user's location region Z = {z1, z2, ..., zr}. m}, where z i (i = 1, 2, ..., m) represents the privacy messages of the basic events. The mathematical models for the privacy sink Z′ and background knowledge Z are as follows:

[0036]

[0037]

[0038] In this case, the privacy source entropy H(Z′) is defined as follows: the privacy source entropy H(Z′) is the average amount of privacy information in the privacy source Z′, representing the degree of privacy uncertainty of Z′. The larger H(Z′) is, the lower the probability of privacy source Z′ being leaked.

[0039]

[0040] Using privacy conditional entropy H(z′) r |Z) represents the inference of the requesting user's location z′ given that the attacker has obtained background knowledge Z and the private destination Z′. r The degree of uncertainty;

[0041]

[0042] Using I(z′) r Z) represents the degree of privacy leakage when the attacker uses background knowledge to carry out the optimal attack. As can be seen from the following formula, the degree of privacy leakage of a user is inversely proportional to the number of users r in the anonymous domain. The larger the value of r, the smaller the degree of privacy leakage.

[0043]

[0044] Step 2: Service Quality Measurement. Based on a given anonymity set Z′ and anonymity machine mechanism f, an attacker can use the standard Bayesian formula to calculate the posterior probability of each position in Z′ and launch an inference attack. The standard Bayesian formula is as follows:

[0045]

[0046] Based on the calculation of the posterior distribution probability of all positions in the anonymous domain, the attacker obtains the inferred position with the highest posterior probability.

[0047] z″ r =arg maxPr(z′) i |Z′)

[0048] The expected error of the attacker's optimal prediction against the anonymity mechanism f is used to measure the quality of service received by the user. The attacker's expected error is as follows:

[0049] EE(f)=1-maxPr(z′) i |Z′)

[0050] The quality of service for a user is determined by the anonymity mechanism f, the anonymity domain generated by the user's chosen privacy level r, and the attacker's predicted error. The privacy level r is inversely proportional to the quality of service; the higher the privacy level r, the more severe the anonymity generalization, and the worse the service quality the user receives. The predicted error EE(f) is also inversely proportional to the quality of service; the higher the attacker's predicted error EE(f), the better the privacy protection effect generated by the anonymity mechanism f, and the worse the service quality. The calculation method for the service quality is as follows:

[0051]

[0052] Step 3: Game theory selects the optimal value of k, given the user's payoff function as E. u Based on the principle of maximizing benefits, users will maximize their benefit function, which is inversely proportional to the degree of privacy leakage. The greater the degree of privacy leakage, the greater the loss of benefits for users. It is also directly proportional to the quality of service provided to users. The better the service quality, the greater the benefits for users.

[0053]

[0054]

[0055] The following conditions must be met for this expression to be valid:

[0056] η+θ=1

[0057]

[0058]

[0059] Therefore, we derive the following formula for the r value that provides users with optimal service quality and location privacy:

[0060]

[0061] As can be seen from the above, this value will select different optimal values ​​depending on the specific scenario of the requesting user and the anonymity protection mechanism used, so that K-anonymous users can select the number of assisting users that best suits their own conditions in different states.

Claims

1. A personalized k-anonymous optimization method based on game theory, the specific steps of which are as follows: Step 1: Location Privacy Measurement. Assume the requesting user is the information sender, the attacker is the receiver, and the privacy leakage channel is the communication channel. The information set possessed by the sender is called the privacy source, referring to the anonymous domain constructed by the requesting user. The information set obtained by the receiver is called the privacy sink, referring to the anonymous domain observed by the attacker, which is the same as the privacy source, i.e., the anonymous domain Z' = {z'1, z'2, ..., z'}. r },in, z j '(j=1,2,...,r) represents a user's location privacy information, and the attacker's background knowledge is the user's location region Z={z1,z2,...,zr}. m }, where z i (i = 1, 2, ..., m) represents the privacy messages of the basic events. The mathematical models for the privacy sink Z' and background knowledge Z are as follows: In this case, the privacy source entropy H(Z') is defined as follows: the privacy source entropy H(Z') is the average amount of privacy information in the privacy source Z', representing the degree of privacy uncertainty of Z'. The larger H(Z') is, the lower the probability of privacy source Z' being leaked. Using privacy conditional entropy H(z') r |Z) represents the inference of the requesting user's location z' given that the attacker has background knowledge Z and the private destination Z'. r The degree of uncertainty; Using I(z') r Z) represents the degree of privacy leakage when the attacker uses background knowledge to carry out the optimal attack. As can be seen from the following formula, the degree of privacy leakage of a user is inversely proportional to the number of users r in the anonymous domain. The larger the value of r, the smaller the degree of privacy leakage. Step 2: Service Quality Measurement. Based on the given anonymity domain Z' and anonymity machine mechanism f, attackers can use the standard Bayesian formula to calculate the posterior probability distribution of each position in Z' and launch an inference attack. The standard Bayesian formula is as follows: The attacker obtains the inferred position based on the posterior probability of all positions in the anonymous domain, which is the position with the highest posterior probability. With" r =arg maxPr(z i |Z) The expected error of the attacker's optimal prediction against the anonymity mechanism f is used to measure the quality of service received by the user. The attacker's expected error is as follows: EE(f)=1-maxPr(z′ i |Z′) The service quality is calculated as follows: Step 3: Game theory to select the optimal value of k, given the user's payoff function as E. u Based on the principle of maximizing benefits, users will maximize their benefit function, which is inversely proportional to the degree of privacy leakage. The greater the degree of privacy leakage, the greater the loss of benefits for users. It is also directly proportional to the quality of service provided to users. The better the service quality, the greater the benefits for users. The following conditions must be met for this expression to be valid: η+θ=1 Therefore, we derive the following formula for the r value that provides users with optimal service quality and location privacy:

Citation Information

Patent Citations

  • Demand privacy protection method based on differential privacy and association rules

    CN108520182A

  • Method and system for achieving anonymity in location based services

    KR1020160066661A