Generative recommendation system-oriented member reasoning attack method

By training two recommendation models of non-coinciding data sets in the generative recommendation system and performing member inference based on the overlap of the recommendation list, the problems of personalized recommendation and privacy protection of non-member users are solved, and accurate membership judgment is achieved.

CN119939657APending Publication Date: 2025-05-06CHONGQING UNIV OF POSTS & TELECOMM
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510021407.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

It is difficult for the prior art to personalize recommendations to non-member users in recommendation systems, and at the same time, member reasoning attacks threaten user privacy, and traditional methods are difficult to apply in large-scale recommendation systems.

Method used

A member reasoning attack method for a generative recommendation system is proposed, by training the recommendation model of two completely non-coinciding data sets, and membership inference is made based on the overlap of the recommendation list output by the two models.

Benefits of technology

It realizes personalized recommendations for non-member users, enhances the privacy protection capabilities of user data, and can accurately determine whether the user is a member of the model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939657A_ABST
    Figure CN119939657A_ABST
Patent Text Reader

Abstract

The invention relates to a member reasoning attack method for a generative recommendation system, and belongs to the field of computer recommendation systems and the field of information security. The method specifically comprises the following steps: converting historical interaction data of a target user into an input format required by a target model, merging the data into another shadow data set datashadow which is completely not overlapped with a target data set datatarget, training an auxiliary model modeleauxiary which is the same as the target model in structure by using the same parameters, respectively inputting the data of the target user into the target model modularget and the auxiliary model modularget to obtain recommendation lists output by the two models, and calculating the overlap ratio epsilon of the recommendation lists as the feature of the target user; and inputting the information into an inference model, and finally outputting the member relationship of the target user. The method has good effectiveness in deducing the accuracy of the membership.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of computer recommendation systems and information security, and relates to a member reasoning attack method based on a generative recommendation system. Background Art

[0002] Recommendation systems play an important role in modern Internet applications. Their main task is to provide users with personalized content recommendations by analyzing their historical behaviors and preferences, thereby improving user experience and promoting the commercial benefits of the platform. Traditional recommendation systems (such as models based on collaborative filtering or matrix decomposition) usually rely on user-item interaction matrices and cannot make effective personalized recommendations for non-member users who have not participated in model training. However, with the development of deep learning and generative models, more and more recommendation systems are able to make personalized recommendations for non-member users by learning their behavior patterns.

[0003] In the field of security and privacy, membership inference attack (MIA) is a method of inferring whether specific data participates in model training through the output of the model. The essence of this attack method is to challenge the privacy protection ability of the recommendation system because it directly threatens the security of the user's private data. Through membership inference attacks, attackers can not only infer whether a certain user participates in the training of the recommendation model, but also further infer the user's behavior patterns, interests and hobbies, and even more sensitive personal information (such as location, age, gender, etc.), and use this sensitive information to commit fraud or steal their accounts. According to the General Data Protection Regulation, users have the right to be forgotten, which means that membership inference attacks will be an effective way to determine whether a user's data is used to train the recommendation system.

[0004] Current research on membership inference attacks focuses on traditional classifier scenarios, especially when only labels (i.e., decisions) are used for reasoning. In these studies, attackers usually rely on the output labels of the target model and use shadow datasets or adversarial examples to attack. Although these methods are effective in some scenarios, they are often difficult to apply in real-world recommender systems for two main reasons: first, recommender systems usually deal with large-scale and complex datasets, which makes it difficult for traditional membership inference attack techniques to be effectively applied. In recommender systems, the sparsity of the user-item matrix makes it difficult for attackers to obtain enough information to infer whether a specific data point is used to train the model; second, unlike classical classifiers, the output of recommender systems is a ranked list of items other than unordered labels. In this case, sequential information plays an important role and can greatly facilitate user preference prediction. Therefore, it is necessary for our attack model to capture sequential information from the recommended items, which is still ignored by previous membership inference attack methods.

[0005] However, current membership inference attacks on recommendation systems all use different recommendation methods based on member samples and non-member samples (i.e., member samples can be personalized after model training, while non-member samples are recommended based on popularity). This does not meet the needs of personalized recommendations for non-member samples (i.e., cold-start users). Therefore, it is of great value to design a recommendation model membership inference attack that can perform personalized recommendations on non-member users. Summary of the invention

[0006] In view of this, the purpose of the present invention is to provide a membership inference attack method for a generative recommendation system, aiming to enable users to more accurately understand whether their data is used to train a recommendation model.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] A membership inference attack method for a generative recommendation system (taking CFGAN as an example) specifically includes the following steps:

[0009] S1: First use the target training set dataset tttttttttttt The data of train a target recommendation model (cfgan), here denoted as model tttttttttttt , and the dataset tttttttttttt The data in is divided into the training set and test set of the attack model, denoted as member tttttttttt ,member tttttttt , where member tttttttttt Used to train the inference model, membertttttttt Used to verify the accuracy of the attack.

[0010] S2: Target user data ttttsssssstt The historical interaction data is integrated into another dataset tttttttttttt Completely independent new dataset dataset tt h ttaaaaaa Next, we use these two datasets to jointly train an auxiliary model model ttaaaattssttttttaa ,This model uses the same structure and parameter training method as the target model;

[0011] S3: Target user data ttttsssssstt Convert it into the data type required by the model input, input it into the two trained recommendation models respectively, and obtain the two recommendation lists output by the model tttttttttttt ,list tt h ttaaaaaa , where list tttttttttttt is the Top-k recommendation list output by the target model, list tt h ttaaaaaa It is the Top-k recommendation list output by the auxiliary model.

[0012] S4: Obtain the target user's recommendation list in the two recommendation models and input the recommendation list into the inference model model ttttiittttttttiitt The inference model here is a binary classification discriminator, based on the overlap between the two recommendation lists:

[0013] Where n iiaatt Indicates the number of overlapping items in the two recommendation lists, and k indicates the total number of recommendation lists, i.e., Top-k. The membership relationship is determined based on the relationship between ε and θ. If ε>θ, the target user is determined as a member, and if ε<θ, the target user is determined as a non-member. θ indicates a judgment threshold.

[0014] Furthermore, in step S1, the target recommendation model adopts a generative recommendation model of a generative adversarial network framework, such as cfgan, whose core components are a generator (G) and a discriminator. The generator inputs the user's feature vector and the item feature matrix Output the recommendation probability distribution of the user: in is the preference prediction for the item, θ GG are the generator parameters; the discriminator Input is real interactive data and the generator output Output the probability of distinguishing true and false data: where θ DD is the discriminator parameter. The loss function based on the generator and the discriminator is:

[0015]

[0016] The training process is achieved by alternating optimization of the generator and the discriminator.

[0017] Where η is the learning rate

[0018] Further, in step S2, the dataset tttttttttttt Completely non-overlapping datasets tt h ttaaaaaa It is the core of auxiliary model training. Specific data can be extracted from other public data sets in the same field to ensure that the data properties (such as user behavior characteristics, project types) are similar to the target training set. Or, when public data is insufficient, shadow data can be constructed through generation methods: simulate the probability distribution P(u,v) of user behavior and generate interaction records from it. ttaaaattssttttttaa Adoption and target model tttttttttttt The same architecture and parameter initialization method are used, and then the historical interaction data of the target user is combined to repeat step S1 to train the auxiliary model model ttaaaattssttttttaa .

[0019] Furthermore, in step S3, the focus is on generating recommendation lists through the target model and the auxiliary model, and extracting relevant features of these lists to provide input for the subsequent inference model. First, the target user's interaction data ttttsssssstt The input format accepted by the otaku model is not yet accepted. If an explicit recommendation model is used, the user's historical interaction behavior is converted into a coefficient vector V aa =[r aa1 ,r aa2 ,……,r aatt ],r aatt ∈{0,5},r aatt represents the target user's rating of item i; if an implicit recommendation model is used, r aatt Converted into data that is either 0 or 1, that is V aa The target model and the auxiliary model are input for recommendation respectively. The target model and the auxiliary model output the recommendation list respectively. The target model and the auxiliary model generate the recommendation score according to the input. Then sort and select the first k items Similarly, the same method can be used to obtain the recommendation list of auxiliary models. tt httaaaaaa

[0020] Furthermore, in step S4, the user features input into the inference model are obtained based on the number of overlapping items in the two recommendation lists in step S3. iiaatt =|list tttttttttttt ∩list tttttttttttt |, then calculate the overlap The inference model used here is a simple binary classification neural network MLP, which is trained to determine whether the target user is a member. The input features of the inference model come from the output list of the target recommendation model and the auxiliary recommendation model obtained in step S3 (i.e., list tttttttttttt ,list tt h ttaaaaaa The training process of the inference model specifically uses the subset member of the target dataset tttttttttt As the member training set, and then randomly generate some user samples as nonmember tttttttttt The training data set (x, y) of the inference model is constructed together, where y=1 indicates that the target user is a member, and y=0 indicates that the target user is a non-member. Based on this, the judgment threshold parameter θ of the inference model can be trained. After that, inference can be made based on the size relationship between the user's features and the judgment threshold parameter θ.

[0021] The beneficial effect of the present invention is that: the present invention aims at the privacy leakage problem in the generative recommendation system scenario, and proposes a member reasoning attack method for the generative recommendation system. The present invention simultaneously inputs the historical interaction data of the target user into two recommendation models trained by completely non-overlapping data sets, and infers the membership relationship based on the overlap of the two recommendation lists. Since members and non-members use the same recommendation model, the user characteristics of the two will be very similar, so here we consider a white box attack method.

[0022] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0024] Figure 1 A schematic diagram for inferring membership relationships of target users;

[0025] Figure 2 This is a model diagram of a generative recommendation system (taking cfgan as an example); DETAILED DESCRIPTION

[0026] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0027] Among them, the drawings are only used for illustrative explanations, and they only represent schematic diagrams rather than actual pictures, and should not be understood as limitations on the present invention. In order to better illustrate the embodiments of the present invention, some parts of the drawings may be omitted, enlarged or reduced, and do not represent the size of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0028] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right", "front", "rear", etc. indicate the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0029] See also Figure 1-2 , specifically including the following steps:

[0030] Step 1: Use the target training set dataset tttttttttttt The data of train a target recommendation model (cfgan), here denoted as model tttttttttttt , whose core component is the generator and the discriminator The generator inputs the user's feature vector and the item feature matrix Output the recommendation probability distribution of the user: in is the preference prediction for the item, θ GG are the generator parameters; the discriminator Input is real interactive data and the generator output Output the probability of distinguishing true and false data: where θ DD is the discriminator parameter. The loss function based on the generator and the discriminator is:

[0031]

[0032] The training process is achieved by alternating optimization of the generator and the discriminator.

[0033] Where η is the learning rate and the dataset tttttttttttt The data in is divided into the training set and test set of the attack model, denoted as member tttttttttt ,member tttttttt , where member tttttttttt Used to train the inference model, member tttttttt Used to verify the accuracy of the attack.

[0034] Step 2: Target user data ttttsssssstt The historical interaction data is merged into another dataset tttttttttttt Completely non-overlapping shadow datasets tt h ttaaaaaa , dataset tt h ttaaaaaa It is the core of auxiliary model training. Specific data can be extracted from other public data sets in the same field to ensure that the data properties (such as user behavior characteristics, project types) are similar to the target training set. Or, in the case of insufficient public data, shadow data can be constructed through generation methods: simulate the probability distribution P(u,v) of user behavior and generate interaction records from it. Then use these data to jointly train an auxiliary model model with the same structure and parameter training as the target model. ttaaaattssttttttaa The specific training loss function is the same as step 1.

[0035] Step 3: Target user data ttttsssssstt Convert it into the data type required by the model input. The specific conversion format depends on whether the recommendation model uses explicit or implicit data. Input the interaction data into the trained target model model. tttttttttttt and auxiliary model model ttaaaattssttttttaa In , we obtain the Top-k recommendation lists output by the two recommendation models respectively.

[0036] Step 4: Add the two recommendation lists obtained in step 3 tttttttttttt , list tt h ttaaaaaaThen input them into the trained inference model model ttttiittttttttiitt The inference model used here is a simple binary classification neural network MLP. The input features of the inference model come from the output list of the target recommendation model and the auxiliary recommendation model obtained in step S3 (i.e., list tttttttttttt ,list tt h ttaaaaaa ) of the overlap degree ε( Where n iiaatt The training process of the inference model specifically uses the subset member of the target dataset. tttttttttt As the member training set, and then randomly generate some user samples as nonmember tttttttttt The training data set (x, y) of the inference model is constructed together, where y = 1 means the target user is a member, and y = 0 means the target user is a non-member. After training, the judgment threshold θ can be obtained, and then the membership relationship can be inferred based on the size relationship between the input user features and θ.

[0037] Example

[0038] The member reasoning attack method for a generative recommendation system described in the present invention specifically comprises the following steps:

[0039] Step 1: Target user data ttttsssssstt The historical interaction data is merged into another dataset tttttttttttt Completely non-overlapping shadow datasets tt h ttaaaaaa , and then train an auxiliary model with the same structure as the target model with the same parameters ttaaaattssttttttaa , the training framework of the specific model is as follows Figure 2 shown.

[0040] Step 2: Calculate the number of overlapping items n in the recommendation lists output by the two recommendation models for the target user iiaatt =|list tttttttttttt ∩list tttttttttttt |, then calculate the overlap And use it as the user feature of the target user and input it into the inference model model ttttiittttttttiitt , the inference model outputs the membership relationship of the target user (0 indicates non-member, 1 indicates member). The specific inference process is as follows Figure 1 shown.

[0041] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A membership inference attack method for a generative recommender system, characterized in that: The method specifically comprises the following steps: S1: First use the target training set dataset target The data of train a target recommendation model (cfgan), here denoted as model target , and the dataset target The data in is divided into the training set and test set of the attack model, denoted as member train ,member test , where member train Used to train the inference model, member test Used to verify the accuracy of the attack; S2: Target user data sample The historical interaction data is integrated into another dataset target Completely independent new dataset dataset shadow Next, we use these two datasets to jointly train an auxiliary model model auxiliary ,This model uses the same structure and parameter training method as the target model; S3: Target user data sample The data is converted into the data type required by the model input, and input into the two trained recommendation models respectively to obtain the two recommendation lists output by the model. target ,list shadow , where list target is the Top-k recommendation list output by the target model, list shadow It is the Top-k recommendation list output by the auxiliary model; S4: Obtain the target user's recommendation list in the two recommendation models and input the recommendation list into the inference model model inference The inference model here is a binary classification discriminator, based on the overlap between the two recommendation lists: Recommendation, where n con represents the number of overlapping items in the two recommendation lists, k represents the total number of recommendation lists, i.e., Top-k. The membership relationship is judged according to the relationship between ε and θ. If ε>θ, the target user is judged as a member, and if ε<θ, the target user is judged as a non-member, where θ represents a judgment threshold.

2. The member reasoning attack method for a generative recommendation system according to claim 1, characterized in that: In step S1, the target recommendation model adopts a generative recommendation model of a generative adversarial network framework, such as cfgan, whose core components are a generator (G) and a discriminator. The generator inputs the user's feature vector and the item feature matrix Output the recommendation probability distribution of the user: in is the preference prediction for the item, θ G are the generator parameters; Discriminator Enter interaction data and the generator output Output the probability of distinguishing true and false data: where θ D is the discriminator parameter. The loss function based on the generator and the discriminator is: The training process is achieved by alternating optimization of the generator and the discriminator. Where η is the learning rate.

3. The member reasoning attack method for a generative recommendation system according to claim 2, characterized in that: In step S2, the dataset target Completely non-overlapping datasets shadow It is the core of auxiliary model training. Specific data can be extracted from other public data sets in the same field to ensure that the data properties (such as user behavior characteristics, project types) are similar to the target training set. Or, when public data is insufficient, shadow data can be constructed through generation methods: simulating the probability distribution P(u,v) of user behavior and generating interaction records from it. auxiliary Adoption and target model target The same architecture and parameter initialization method are used, and then the historical interaction data of the target user is combined to repeat step S1 to train the auxiliary model model auxiliary。 4. The member reasoning attack method for a generative recommendation system according to claim 3, characterized in that: In step S3, the focus is on generating recommendation lists through the target model and the auxiliary model, and extracting relevant features of these lists to provide input for the subsequent inference model. First, the target user's interaction data sample Convert to the input format accepted by the model. If an explicit recommendation model is used, convert the user's historical interaction behavior into a coefficient vector V u =[r u1 ,r u2 ,……,r un ],r ui ∈{0,5},r ui represents the target user's rating of item i; If an implicit recommendation model is used, r ui Converted into data that is either 0 or 1, that is V u The target model and the auxiliary model are input for recommendation respectively. The target model and the auxiliary model output the recommendation list respectively. The target model and the auxiliary model generate the recommendation score according to the input. Then sort and select the first k items Similarly, the recommendation list of the auxiliary model is obtained in the same way shadow。 5. The member reasoning attack method for a generative recommendation system according to claim 4, characterized in that: The user features input to the inference model in step S4 are obtained based on the number of overlapping items in the two recommendation lists in step S3. con =|list target ∩list shadow |, then calculate the overlap The inference model used here is a simple binary classification neural network MLP, which is trained to determine whether the target user is a member. The input features of the inference model come from the output list of the target recommendation model and the auxiliary recommendation model obtained in step S3 (i.e., list target , list shadow ), the training process of the inference model specifically uses the subset member of the target dataset train As the member training set, and then randomly generate some user samples as nonmember train The training data set (x, y) of the inference model is jointly constructed, where y=1 indicates that the target user is a member, and y=0 indicates that the target user is a non-member. Based on this, the judgment threshold parameter θ of the inference model can be trained, and then inference can be performed based on the size relationship between the user's characteristics and the judgment threshold parameter θ.

Citation Information

Patent Citations

  • Federal learning member inference method based on prediction confidence sequence

    CN113850399A

  • User privacy protection method and system for recommendation system, equipment and medium

    CN116186693A

  • Anti-robust cross-domain recommendation model and training method

    CN118364894A

  • Unlearning of recommendation models

    US20240070525A1