A combined recommendation method based on k-anonymity privacy protection

By adopting a k-anonymized privacy protection method and clustering generalization technology in the recommendation model, combining personalized protection of user sensitive attributes, and using a combined recommendation algorithm based on content and knowledge, the shortcomings of traditional recommendation models in terms of privacy protection are solved and more efficient privacy protection is achieved.

CN114386084BActive Publication Date: 2025-05-23GUIZHOU POWER GRID CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110807462.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-16
Publication Date
2025-05-23
Estimated Expiration
2041-07-16

AI Technical Summary

Technical Problem

Traditional recommendation models have the risk of security privacy leakage when dealing with user privacy protection, especially in the case of link attacks, and the prior art is difficult to effectively protect user privacy.

Method used

A privacy protection method based on k-anonymity is adopted to reduce information loss by building generalization trees and clustering generalization, and combined with personalized protection of user sensitive attributes, the protection of sensitive attributes is achieved. At the same time, a combined recommendation algorithm based on content and knowledge is adopted to improve privacy protection.

Benefits of technology

It effectively solves the shortcomings of traditional recommendation models in terms of privacy protection. Through personalized sensitive attribute protection and combined recommendation algorithms, the intensity of privacy protection is improved and the security of user privacy is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114386084B_ABST
    Figure CN114386084B_ABST
Patent Text Reader

Abstract

The invention discloses a combined recommendation method based on k-anonymity privacy protection, which is applied to the recommendation algorithm through an improved k-anonymity data rotation method, presents different k values ​​on the same data table according to the privacy requirements of the user, so as to meet the personalized k-anonymity; the generalized query is transmitted to the server in a rotation manner, the server runs the combined recommendation algorithm and builds a resource space storage return result of a high-response scheduling, builds a token request response based on the principle of locality, and the combined result returned by the server is returned to the user by polling tokens; the traditional recommendation algorithm solves the problems of personal privacy attack or third-party server privacy leakage caused by feature extraction, data feature extraction and calculation. By applying the improved k-anonymity privacy protection method to different recommendation methods and combining the results calculated by different recommendation methods, the existing privacy-protected combined recommendation results are obtained, which solves the problems of poor recommendation effect and privacy leakage of traditional recommendation models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a combination recommendation method based on k-anonymity privacy protection. Background Art

[0002] Traditional recommendation methods mainly include collaborative filtering and content-based recommendation methods. Among them, the most classic algorithm is collaborative filtering, such as matrix factorization, which uses the direct interaction information between users and items to generate recommendations for users. Collaborative filtering is currently the most widely used recommendation algorithm.

[0003] The collaborative filtering algorithm is a recommendation process that filters massive amounts of information by collaborating with everyone's feedback, evaluations, and opinions to select information that may be of interest to target customers.

[0004] Content-based recommendation methods make recommendations based on item familiarity, the degree of match between user attributes and item features, and the effectiveness of recommendations strongly depends on the quality of feature engineering.

[0005] The shallow model used by the collaborative filtering algorithm cannot learn the deep features of users and items. Content-based recommendation methods use existing behaviors to recommend items with similar attributes. This method requires effective feature extraction. Traditional models rely on manually designed features, and their effectiveness and scalability are very limited, which restricts the performance of content-based recommendation methods.

[0006] As more and more data on the Internet can be perceived and obtained, most recommendation systems ignore the protection of user privacy and do not consider the possible leakage of user privacy in the recommendation process.

[0007] To address the privacy leakage problem caused by link attacks, the k-anonymity algorithm generalizes the attributes of quasi-identifiers so that each published tuple has at least k-1 other tuples with the same attributes. Since the quasi-identifier attribute values ​​are the same after generalization, the attacker cannot identify the individual to whom the information belongs based on the background knowledge he has, but the quasi-identifier attributes have not yet been personalized and need further study.

[0008] In order to improve the security mechanism of the recommendation system and make privacy protection more valuable, a combined recommendation method based on k-anonymity privacy protection is now needed to solve the above problems. Summary of the invention

[0009] In view of this, the first aspect of the present invention aims to provide a combined recommendation method based on k-anonymity privacy protection, which solves the risk of security and privacy leakage of traditional recommendation models.

[0010] The purpose of the first aspect of the present invention is achieved through the following technical solutions:

[0011] A combined recommendation method based on k-anonymity privacy protection, comprising:

[0012] Step S1: extract anonymous tuples according to the data table to be anonymized, and construct a generalization tree for each quasi-identifier;

[0013] Step S2: Use a clustering-based method to generalize tuples to reduce information loss caused by over-generalization;

[0014] Step S3: Evaluate the personalized anonymization results of the data table, anonymously check whether the tuples meet the anonymity requirements set for each, and integrate and generalize the tuples that do not meet the requirements;

[0015] Step S4: semantically analyze the keywords in the request information sent by the user in the tuple generalization, generalize the sensitive attributes, reduce the accuracy of the sensitive attributes, and forward them to other nodes in the tuple by polling. Users who meet the forwarding termination conditions are sent to the server;

[0016] Step S5: After receiving the data packet, the server performs resource matching, retrieves candidate data that meets the information required by the user from the database, and determines at least two recommendation algorithms according to the number of explicit features and the number of implicit features contained in the user information, and obtains the recommendation result of the recommendation algorithm;

[0017] Step S6: The server constructs a high-response token and data storage structure based on the principle of locality, and performs token verification with the user in a data rotation manner on the result returned by the combined recommendation method, and finally returns the recommended data to the user.

[0018] Furthermore, in step S2, a generalization idea based on clustering is adopted to divide the tuple data into clusters according to anonymity requirements.

[0019] Furthermore, in step S3, the personalized anonymous results of the data table are evaluated to check whether the tuples meet the generalization requirements.

[0020] Further, in step S4, the generalized processing is subjected to k-anonymization processing, and the user requests that meet the data rotation termination condition enter the server.

[0021] Furthermore, in step S5, for different numbers of explicit features and implicit features, privacy combination recommendations using different recommendation methods are performed.

[0022] Further, in step S6, the server stores the data request token and the recommendation result in a scheduling storage resource with a high response ratio priority to build a local storage, and performs token verification with the user in a data rotation manner on the result returned by the combined recommendation method.

[0023] The second aspect of the present invention aims to provide an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the program, the steps of the personalized recommendation method based on k-anonymity privacy protection as described above are implemented.

[0024] The third aspect of the present invention aims to provide a non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is processed and executed, the steps of the combined recommendation method based on k-anonymity privacy protection as described above are implemented.

[0025] The beneficial effects of the present invention are:

[0026] The present invention uses an improved k-anonymity privacy protection method, constructs an anonymous set for sensitive attributes based on a clustering generalization method, defines personalized sensitive attribute levels according to user sensitive attributes, and uses a sensitive attribute generalization tree to achieve personalized protection of sensitive attributes, thereby meeting the user's personalized protection needs for sensitive attributes. At the same time, a combined recommendation algorithm based on content recommendation and knowledge recommendation is used to solve the problems of artificial feature extraction, sparse and high-dimensional data features, and lack of attention to user privacy in traditional recommendation algorithms. An improved k-anonymity model is used to combine user sensitive attributes and communicate with the server in a polling manner to obtain recommendation results, thereby improving the privacy protection of traditional recommendation models.

[0027] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description and the preceding claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0029] Figure 1 A block diagram of a recommended method of the present invention;

[0030] Figure 2 This is a schematic diagram of network attack and defense;

[0031] Figure 3 This is a diagram of data polling. DETAILED DESCRIPTION

[0032] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the preferred embodiments are only for illustrating the present invention, rather than for limiting the protection scope of the present invention.

[0033] As Figure 1 shown, the present invention provides a combined recommendation method based on k-anonymous privacy protection, including:

[0034] Extract anonymous tuples from the data table to be anonymized, and construct a generalization tree for each quasi-identifier; after anonymization, if the attributes of the tuple remain unchanged, there will be no information loss; if the attribute value is replaced by the value of a node in the attribute generalization tree, a certain amount of information loss will occur. The closer the replaced attribute value is to the root node of the generalization tree, the greater the information loss. The information loss of an attribute value is defined as the generalization height of the attribute value.

[0035] For the data table T, the anonymous requirement for n tuples is defined as the set S = {k 1 , k 2 , …, k m} (m < n)m, and the elements in S are not repeated. k i < k j ∈ S, if i < j, then k i < k j . After anonymization, from a global perspective, if the anonymous requirement is k a (a ∈ [1, m])k a anonymous, then it is said that T satisfies global personalized k-anonymity.

[0036] Then, evaluate the personalized anonymous result of the data table, check whether the tuples meet the anonymous requirements, and integrate and generalize the non-compliant tuples. Taking the education level as an example, the generalization is shown in Table 1.

[0037] Table 1 k-anonymous equivalence classes after generalization

[0038]

[0039]

[0040] After the anonymized data set is integrated, perform semantic analysis on the keywords in the query request sent by the user and then perform sensitive attribute generalization processing to reduce the accuracy of sensitive attributes. Forward the query request that has been generalized to meet the data packet forwarding, randomly poll and forward it to other nodes, increment the count by 1 for each forwarding, and adjust the forwarding termination to submit the user's query request to the server through token verification; the high-response unit of the server stores the request token and the return result, accesses according to the locality principle, and the server performs resource matching according to the combined recommendation algorithm, and returns the matching result to the user in the same way of token polling verification.

[0041] Specifically, it is divided into two parts: sending, combined recommendation and receiving

[0042] (1) Sending

[0043] First, the user constructs a query request and generates and saves a timestamp that uniquely identifies this set of query requests.

[0044] Subsequently, the user generates a random number i, where 1 < i < k, to mark the number of forwarding times, and randomly sends the constructed request together with the timestamp and the random number to other individuals in the same equivalence class. The user who receives this request records the number of this request and forwards the request to the users in the same group. To prevent the first forwarding user from inferring that the person before him is the real request initiator, the number of forwarding times initially submitted by the real request initiator is not 0, but a random number in the range of [1, (k - 1)].

[0045] After other hosts receive the query request, they check whether i is equal to k. If not, they continue to forward randomly. If equal, they submit it to the server, and the IP address of the host that sends this request is recorded by the timestamp.

[0046] (2) Combined recommendation

[0047] The server determines at least two recommendation algorithms according to the quantities of explicit features and implicit features included in the user request information; when the quantity of explicit features does not exceed the first preset threshold and the quantity of implicit features does not exceed the second preset threshold, the explicit features and implicit features are complemented; when the explicit features do not exceed the first preset value and the implicit features exceed the second preset value, a combined recommendation algorithm based on knowledge and content is determined; when the explicit features exceed the first preset value and the implicit features do not exceed the second preset value, a combined recommendation algorithm based on collaborative filtering and social network is adopted; when the quantities of both explicit features and implicit features exceed the preset value, at least two of the above three combined recommendation algorithms are adopted.

[0048] (3) Receiving

[0049] The server returns the combined recommendation result to the host that uploads the query request to the server. This host first checks the timestamp to determine whether it is the request constructed by itself. If not, after receiving the result, it queries the previously recorded IP address and then sends the received result back to the previously recorded IP address; if so, it retains it. The next host repeats the above process, and the sending-back process is completed in turn, so that the received result finally returns to the hands of the real request submitter. In this process, only the users who participated in the forwarding process during data sending can receive the returned result, and the other users who did not participate in the sending process will not appear in the receiving process.

[0050] Generally speaking, the method of this embodiment can be refined into the following specific steps:

[0051] Step S1: Extract anonymous tuples according to the data table to be anonymized and construct a generalization tree for each quasi-identifier;

[0052] Step S2: Divide all the extracted tuple data into clusters according to their respective anonymity requirements. The anonymity requirement of each tuple is the anonymity requirement of its corresponding cluster.

[0053] Step S3: Perform a tuple generalization process for each cluster, find an equivalent class in the cluster whose number of tuples is less than the anonymity requirement of the cluster, and find the two equivalent classes with the highest similarity in the cluster and merge them;

[0054] Step S4: generalize the quasi-identifier attribute values ​​of all tuples in the equivalence class to the same attribute value to ensure minimum generalization, thereby reducing information loss caused by over-generalization. When all clusters are generalized, the tuple generalization ends.

[0055] Step S5: perform an anonymous requirement check on all tuples in the data table and calculate the number of tuples with the same quasi-identifier attribute values;

[0056] Step S6: If the number of tuples with the same quasi-identifier attribute value is greater than or equal to the anonymity requirement of the tuple, the tuple is deemed to have met the anonymity requirement set for itself and can be directly output; if the number is less than the anonymity requirement of the tuple, the tuple is deemed to have failed to meet the anonymity requirement of the tuple, and it is determined whether the tuple has been processed by steps S1-S4. If so, step S7 is executed; otherwise, step S8 is executed;

[0057] Step S7: The original records in the data table corresponding to the tuples processed by S1-S4 are mixed with other tuples that do not meet the anonymity requirements to be re-anonymized, so as to reduce the data loss caused by the repetition of the generalization process;

[0058] Step S8: Execute S1-S4 to generalize the tuples that do not meet the tuple anonymity requirement. When the proportion of the number of tuples that do not meet the anonymity requirement to the total number of tuples in the data table reaches the privacy protection threshold, steps S5-S8 end.

[0059] Step S9: forwarding the query request packets of each user among users in the equivalence class, and uploading data packets of users that meet the termination condition of the data rotation process (the cumulative number of forwarding times is equal to the personalized k value);

[0060] Step S10: After receiving the data packet, the server performs resource matching, retrieves candidate data that meets the information required by the user from the database, and determines at least two recommendation algorithms according to the number of explicit features and the number of implicit features contained in the user information, obtains the recommendation results of each recommendation algorithm, and combines and packages all the obtained recommendation results;

[0061] Step S11: The server stores the data request token and the packaged result of step S10 in a high response ratio priority scheduling storage resource to build a local storage, and performs token verification with the user in a data rotation manner on the result returned by the combined recommendation method, and finally returns the recommended data to the user.

[0062] It should be appreciated that embodiments of the present invention may be implemented or enforced by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method may be implemented in a computer program using standard programming techniques - including a non-transitory computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program may be implemented in a high-level procedural or object-oriented programming language to communicate with a computer system. However, if desired, the program may be implemented in an assembly or machine language. In any case, the language may be a compiled or interpreted language. In addition, the program may be run on a programmed ASIC for this purpose.

[0063] Furthermore, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that may be executed by one or more processors.

[0064] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, a RAM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention also includes the computer itself.

[0065] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents physical and tangible objects, including specific visual depictions of physical and tangible objects produced on the display.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A combined recommendation method based on k-anonymity privacy protection, Features: include: Step S1: extract anonymous tuples according to the data table to be anonymized, and construct a generalization tree for each quasi-identifier; Step S2: Use a clustering-based method to generalize tuples to reduce information loss caused by over-generalization; Step S3: Evaluate the personalized anonymization results of the data table, anonymously check whether the tuples meet the anonymity requirements set for each, and integrate and generalize the tuples that do not meet the requirements; Step S4: semantically analyze the keywords in the request information sent by the user in the tuple generalization, generalize the sensitive attributes, reduce the accuracy of the sensitive attributes, and forward them to other nodes in the tuple by polling. Users who meet the forwarding termination conditions are sent to the server; Step S5: After receiving the data packet, the server performs resource matching, retrieves candidate data that meets the information required by the user from the database, and determines at least two recommendation algorithms according to the number of explicit features and the number of implicit features contained in the user information, and obtains the recommendation result of the recommendation algorithm; Step S6: The server constructs a high-response token and data storage structure based on the principle of locality, and performs token verification with the user in a data rotation manner on the result returned by the combined recommendation method, and finally returns the recommended data to the user.

2. According to claim 1, a combined recommendation method based on k-anonymity privacy protection, Features: In step S2, a generalization idea based on clustering is adopted to divide the tuple data into clusters according to the anonymity requirement.

3. According to claim 1, a combined recommendation method based on k-anonymity privacy protection, Features: In step S3, the personalized anonymous results of the data table are evaluated to check whether the tuples meet the generalization requirements.

4. According to claim 1, a combined recommendation method based on k-anonymity privacy protection, Features: In step S4, the generalized processing is k-anonymized, and the user requests that meet the data rotation termination condition enter the server.

5. According to claim 1, a combined recommendation method based on k-anonymity privacy protection, Features: In step S5, for different numbers of explicit features and implicit features, privacy combination recommendations are made using different recommendation methods.

6. A combined recommendation method based on k-anonymity privacy protection according to claim 1, Features: In step S6, the server stores the data request token and the recommendation result in a scheduling storage resource with a high response ratio priority to build a local storage, and performs token verification with the user in a data rotation manner on the result returned by the combined recommendation method.

7. An electronic device comprising a memory, a processor and a computer program stored in the memory and operable to run on the processor, It is characterized in that When the processor executes the program, the steps of the combined recommendation method based on k-anonymity privacy protection according to any one of claims 1 to 6 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is processed and executed, the steps of the combined recommendation method based on k-anonymity privacy protection as claimed in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • KR20240070155A