Recommendation system transferable black box attack method and device based on gradient integration and medium
By integrating the gradients of different proxy models and generating adversarial samples, the problem of insufficient effectiveness and migration ability of a single proxy model in the prior art is solved, and a more efficient and robust black box attack is achieved.
Patent Information
- Application Number
- CN202510300819.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The black box adversarial attack method of the existing recommendation system relies on a single proxy model and cannot integrate multi-model decisions, resulting in insufficient effectiveness of adversarial samples and cross-model migration capabilities, and is susceptible to local extreme situations of the model, affecting the robustness of the attack.
A gradient integration-based method is adopted to integrate the gradients of different proxy models, and through standardization and weighting integration, global gradients are obtained, the update direction of fake user data is dynamically adjusted, and adversarial samples are generated.
The attack performance of the recommendation system against samples in different models is improved, the migrationability of the counter samples and the effectiveness of the attack is enhanced, the overfitting problem is avoided, and the robustness of the attack is ensured.
Smart Images

Figure CN120217364A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of black-box attack strategies for improving recommendation systems, and particularly to a transferable black-box attack method, device, and medium for a recommendation system based on gradient integration. Background Art
[0002] In the current information age, recommendation systems are widely used in various Internet platforms to help users screen personalized content, products, or services. However, with the popularization of recommendation systems, user privacy and system security are facing more and more challenges. Especially in the black-box attack scenario, an attacker attempts to influence the recommendation results by generating adversarial samples (i.e., fake users) without knowledge of the internal structure of the target model, thereby damaging the user experience or deviating from the original intention of the recommendation system. To achieve this goal, an attacker usually selects a known model as a surrogate model and uses this surrogate model to approximate the behavioral characteristics of the target model, thereby generating adversarial samples.
[0003] In a recommendation system, a transferable black-box attack refers to an attacker, without any knowledge of the internal structure, training parameters, defense methods, etc. of the recommendation system, using existing knowledge, data, or models to transfer this information to the target recommendation system to infer the behavior of the target system and conduct an attack. Most existing black-box adversarial attack methods for recommendation systems rely on a single surrogate model to generate transferable adversarial samples, which has many drawbacks. On the one hand, it determines the adversarial perturbation direction only based on the decision of a single model, unable to integrate multiple model decisions, difficult to share attack information among different surrogate models, and at the same time, it ignores the difference in gradients of the input in each model. Since the model gradient reflects its sensitivity to the input, this will cause the adversarial samples generated by a single surrogate model to be sensitive only to the important features of a single model and less sensitive to the features of other target models with different architectures, thus limiting the attack effectiveness and cross-model transfer ability. On the other hand, the method based on a single surrogate model is vulnerable to local extreme situations of the model and will produce overfitting, greatly affecting the robustness of the attack.
[0004] In contrast, integrating multiple surrogate models has significant advantages. It can combine multiple model decisions to accurately determine the adversarial perturbation direction, achieve the mutual coordination of attack information among models, make full use of the gradients of the input in different models, enhance the capture of key features by adversarial samples, and greatly improve the attack effectiveness and cross-model transfer ability. At the same time, it can avoid the overfitting problem caused by local extremes of a single model, effectively ensuring the robustness of the attack. However, how to effectively integrate the information of multiple surrogate models to improve the attack success rate and transfer ability of adversarial samples remains a challenging problem.
[0005] The invention application with the application number 202211518012.4 discloses a method, system and electronic device for generating transferable black-box adversarial attack samples. The application solution can effectively improve the success rate of migrating the adversarial attack samples generated for the surrogate white-box model to the black-box attack. However, the solution also has the problem of mainly relying on the decision of a single model.
[0006] Therefore, a new black-box attack strategy based on gradient integration is needed. The gradients of multiple surrogate models can be utilized, and the sensitivity of each model to the input can be regarded as a key feature to generate adversarial samples at a finer-grained level. Summary of the Invention
[0007] Aiming at the above problems, the purpose of the present invention is to provide a method, device and medium for transferable black-box attack of a recommendation system based on gradient integration, which can improve the attack performance of adversarial samples of the recommendation system on different models and enhance the transferability of adversarial samples.
[0008] Embodiments of the present invention provide a method, device and medium for transferable black-box attack of a recommendation system based on gradient integration.
[0009] First aspect: A method for transferable black-box attack of a recommendation system based on gradient integration, comprising:
[0010] S1. Randomly sample the original user-item interaction data to generate initial fake user data, and select recommendation models with different model architectures and hyperparameter configurations as surrogate models;
[0011] S2. Input the fake user data into different surrogate models, update the parameters by backpropagation according to the loss function trained by the surrogate models, and obtain the optimal model parameters;
[0012] S3. Combine the adversarial loss functions of the surrogate models, calculate the gradients of the surrogate models with respect to the optimal model parameters, and standardize them. Use similarity to calculate the importance distribution of the gradients of each surrogate model, and obtain the global gradient by weighting;
[0013] S4. Update the fake user data according to the global gradient, and iterate steps S2 to S3 to obtain the final user adversarial samples;
[0014] S5. Use the target model to verify the attack effect and transfer ability of the user adversarial samples.
[0015] Further, in S1, the generation of the initial fake user data is represented by the formula:
[0016] X′ = RandomSample(X, k) (1)
[0017] Among them, X is the set of original users, X′ is the false user data generated by random sampling, k is the number of sampled items, and a part of the items are randomly sampled as the target item set.
[0018] Furthermore, in the step S2, according to the loss function of the proxy model, backpropagation is used for updating to obtain the optimal model parameters, and the formula is expressed as:
[0019]
[0020] Among them, X ∈ R n×m is the original user rating matrix, n is the number of real users, and m is the number of all items; X′ ∈ R n ′×m is the rating matrix of the false users generated by the proxy model, where n′ represents the number of false users, and θ R is the optimal model parameter of the recommendation model, and θ' R represents the parameter of the recommendation model trained by using the original users, and θ” R represents the parameter of the recommendation model trained by using the false users.
[0021] Furthermore, calculate the gradient of the proxy model with respect to the false users, and the formula is expressed as:
[0022]
[0023] Among them, L attack_i is the adversarial loss function of the proxy model i, and θ R is the optimal model parameter.
[0024] Furthermore, the adversarial loss function is expressed by the formula:
[0025]
[0026] Among them, z is the target item of concern, and x uz is the predicted rating or preference score of user u for the target item z, and x ui is the predicted rating or preference score of user u for item i.
[0027] Furthermore, in the step S3, standardization is performed to obtain the global gradient, and the formula is expressed as:
[0028]
[0029] Among them, w i is the allocation weight of the proxy model i, is the gradient of the proxy model i.
[0030] Furthermore, the allocation weight w i of the proxy model i is expressed by the formula:
[0031]
[0032] Among them, s i is the importance coefficient of the gradient of the false user output by the i-th surrogate model.
[0033] Furthermore, in S4, the false user data is updated according to the global gradient, and the formula is expressed as:
[0034]
[0035] Among them, β is the update step size, and Proj Λ (·) is the projected gradient operator. Through continuous iterative update processes, the false user data is made to approach the original distribution, is not easily detectable, and is updated in the direction that can successfully interfere with the recommendation model.
[0036] Second aspect: An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method provided in the first aspect are implemented.
[0037] Third aspect: A non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in the first aspect are implemented.
[0038] Advantages of the present invention:
[0039] 1. The transferable black-box attack method for the recommendation system based on gradient integration in the present invention integrates the gradients of different surrogate models, dynamically adjusts the optimization and update direction of false users, enables the model to better utilize diverse information for updating, alleviates the problem of excessive dependence of false users on the network structure of the surrogate model, improves the attack performance of adversarial samples of the recommendation system on other target models, enhances the transferability of adversarial samples of the recommendation system, and improves the effectiveness of the attack and the transfer ability of adversarial samples between different models.
[0040] 2. By integrating multiple surrogate models in the present invention, the decision-making basis is more comprehensive, the perturbation direction can be accurately determined, and the attack information can be smoothly transmitted between different models, improving the attack synergy.
[0041] 3. The present invention regards the gradient information of multiple surrogate models as the core element, fully utilizes this information to capture key features, and the generated adversarial samples can be widely adapted to different architecture models, enhancing the attack effectiveness and cross-model transferability, and making up for the insufficient utilization of gradient information.
[0042] 4. With the integration of the multi - surrogate model, the present invention comprehensively considers the characteristics of different models, balances extreme factors, and ensures the smooth and reliable attack process. Even in a complex and changeable recommendation system environment, it can stably output an efficient attack, overcoming the problem of unstable attacks caused by overfitting. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic flowchart of the transferable black - box attack method for a recommendation system based on gradient integration according to the present invention;
[0044] Figure 2 It is a principle flowchart of the transferable black - box attack method for a recommendation system based on gradient integration according to the present invention;
[0045] Figure 3 It is a flowchart of the method for obtaining the global gradient by standardizing different surrogate models according to the present invention;
[0046] Figure 4 It is a distribution diagram of transferable false users generated by the method of the present invention among real users;
[0047] Figure 5 It is a schematic diagram of the structure of the electronic device according to the present invention.
[0048] Appendix Figure 2 English - Chinese Glossary:
[0049] Surrogate Model: Surrogate Model;
[0050] Original User Interaction Matrix: Original User Interaction Matrix; Gradient Correlation Framework: Gradient Correlation Framework;
[0051] global: global;
[0052] update: update;
[0053] Hacker: Hacker;
[0054] Target Model: Target Model;
[0055] Target Items: Target Items;
[0056] User: User;
[0057] Recommended list before attack: Recommended list before attack; Recommended list after attack: Recommended list after attack; Appendix Figure 3 English - Chinese Glossary:
[0058] Integrated gradients; Integrated gradients;
[0059] Multiplication; Multiplication operation;
[0060] weight; Weight;
[0061] Cosine Similarity; Cosine similarity;
[0062] Appendix; Attached; Figure 4 English-Chinese glossary:
[0063] fake; False;
[0064] real; True; Detailed implementation manners; Specific implementation manners;
[0065] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0066] Most of the existing black-box adversarial attack methods for recommendation systems rely on a single surrogate model to generate transferable adversarial samples. On the one hand, it only determines the direction of adversarial perturbation based on the decision of a single model, unable to comprehensively consider the decisions of multiple models, making it difficult to share attack information among different surrogate models. At the same time, it also ignores the gradients of the input in each model. Since the model gradient reflects its sensitivity to the input, this will result in the adversarial samples generated by a single surrogate model being only sensitive to the important features of a single model and having poor sensitivity to the features of other target models with different architectures, thus limiting the attack effectiveness and cross-model transfer ability. On the other hand, the method based on a single surrogate model is vulnerable to local extreme situations of the model and will cause overfitting, greatly affecting the robustness of the attack.
[0067] To address the above problems, the present invention provides a transferable black-box attack method for recommendation systems based on gradient integration. Figure 1 It is a schematic flowchart of the method provided in the embodiments of the present invention.
[0068] Figure 2 It shows the overall principle process of the transferable black-box attack method for recommendation systems based on gradient integration. Starting from initializing fake users and surrogate models, it sequentially goes through gradient calculation, global gradient adjustment and integration, and optimization of fake user data, and finally generates adversarial perturbations and verifies the attack effect in the target model, clearly presenting the logical sequence and data flow between each step. The method includes:
[0069] S1. Randomly sample the original user-item interaction data to generate initial fake user data; at the same time, select recommendation models with different model architectures and hyperparameter configurations as surrogate models.
[0070] Randomly sample the interaction data of the original user-items to generate initial fake user data, which is expressed by the formula:
[0071] X′ = RandomSample(X, k) (1)
[0072] Where the original user set is X, and the fake user data X′ is randomly sampled from X, and the sampling quantity is k.
[0073] The sampled items of the user form the target item set. As Figure 1 shown, 5 items can be randomly sampled from all items as the target item set.
[0074] Select different recommendation models as surrogate models. These surrogate models have different architectures and hyperparameter configurations. For example: select N surrogate models with different architectures and hyperparameter configurations from the recommendation model set D proxy and denote them as:
[0075] M i (i ∈ {1, 2, …, N}) (2)
[0076] These surrogate models will be used to simulate different target attack behaviors.
[0077] S2. Input the fake user data into different surrogate models, design the loss function for training the surrogate models, update the parameters through backpropagation, and obtain the optimal model parameters.
[0078] Design different loss functions between surrogate models Merge the fake user data X′ with the original data X and input it into the surrogate model M i to calculate the loss of the internal training of different surrogate models, perform internal backpropagation of different surrogate models, and iteratively update the model parameters to continuously reduce the loss of the model on the training data and gradually approach the optimal solution to obtain the optimal model parameters, laying a basic framework for subsequent updating of fake users. The formula is expressed as:
[0079]
[0080] Where X ∈ R n×m is the original user rating matrix, n is the number of real users, and m is the number of all items; X′ ∈ R n ′×m is the rating matrix of the fake users generated by the surrogate model, where n′ represents the number of fake users, and θ Ris the optimal model parameter of the recommendation model, θ' R represents the parameter, θ”, for training the recommendation model using the original users R represents the parameter for training the recommendation model using fake users represents the loss function for internal training of the proxy model
[0081] S3. Combine the proxy model adversarial loss function, calculate the gradient of the proxy model with respect to the optimal model parameter, and standardize it. Use similarity to calculate the importance distribution of the gradients of each proxy model, and obtain the global gradient by weighting
[0082] Such as Figure 3 shown Figure 3 is a schematic flow diagram of the gradient integration method for different proxy models designed in the present invention. It details the process of the gradient integration method for different proxy models. It highlights the process from calculating gradients for each proxy model, evaluating the importance of gradients, assigning weights to integrating into the global gradient, enabling readers to deeply understand the core mechanism and operation details of gradient integration
[0083] To promote the target item z to more ordinary users, the cross-entropy loss can be selected as the adversarial loss function, and the formula is expressed as
[0084]
[0085] where z is the target item of concern, and x uz is the predicted score or preference score of user u for the target item z, and x ui is the predicted score or preference score of user u for item i. This means that if the rating of the target item z by the user is greater than other items, the target item will appear in the top-k recommendation list of the user; specifically, the smaller the adversarial target loss, the higher the ranking of the target item in the user's top-k list
[0086] Through the updated optimal model parameter θ R , calculate the gradient of each proxy model with respect to the fake user X′, and the formula is expressed as
[0087]
[0088] The gradient reflects the sensitivity of the proxy model to each feature under the current input, and is the key basis for subsequent gradient integration and adversarial perturbation generation
[0089] Then, regard the gradient information of the fake user calculated by each proxy model as the key feature, and at the same time standardize the gradients of each proxy model regarding the fake user. Here, the standardization is to solve the imbalance problem caused by different magnitudes of the gradient of the task loss function of different proxy models
[0090] To reasonably integrate these gradients, calculate the importance distribution of the gradients of each surrogate model, and let s i represent the importance coefficient of the gradient of the output of the i-th surrogate model with respect to the fake user. Determine the importance coefficient s by calculating the correlation index of the gradients of different surrogate models i .
[0091] Here, the cosine similarity is used to calculate the similarity between the gradients of each surrogate model to measure the common attention area of the surrogate models for the input features.
[0092] Based on the calculated importance coefficient s i , assign a weight w to each surrogate model i . A normalized weight assignment method can be adopted, and the formula is expressed as:
[0093]
[0094] where the sum of the weights of all surrogate models is 1.
[0095] In this way, the contribution ratio of each surrogate model in the global gradient can be dynamically adjusted according to the importance of its gradient. This method not only retains the sensitivity of the surrogate model to important features but also enhances the robustness of adversarial samples by leveraging the diversity between different models.
[0096] Utilize the global weight assignment mechanism to dynamically adjust the weights of the gradients of different surrogate models, integrate the weighted gradient information, and obtain the global gradient The formula is:
[0097]
[0098] This global gradient synthesizes the sensitivity information of multiple surrogate models to input features and can more comprehensively and accurately reflect the influence trend of different features on the recommendation results from different model perspectives, providing strong guidance for generating effective fake users.
[0099] S4. Update the fake user data according to the global gradient, iterate steps S2 - S3, and obtain the final user adversarial sample.
[0100] Then, adjust the fake user X′ according to the global gradient , and the formula is as follows:
[0101]
[0102] where β is the update step size, and Proj Λ (·) is the projection gradient operator, which projects each updated fake user into the feasible region using the threshold ρ (i.e., x ∈ {0, 1}).
[0103] One iteration process includes optimizing the parameters of the internal iteration model and updating the fake users externally. By continuously iterating this update process, the fake user data gradually deviates from the original distribution and updates in the direction that can successfully interfere with the target model.
[0104] By setting the number of iterations, the global gradient obtained through iterative calculation Generate the final user adversarial sample X′ final , which is expressed by the formula:
[0105]
[0106] Among them, β is the update step size, which controls the size of the fake user update. Its value needs to be finely adjusted according to the characteristics of the target model and the attack target. An overly large β may cause the generated adversarial sample to deviate too much from the normal data distribution, and the adversarial sample is easily detected by the target model; an overly small β may make the adversarial perturbation insufficient to affect the recommendation result and unable to achieve the attack purpose. In practical applications, the appropriate β value can be determined by conducting multiple experiments within a certain range and observing the attack effect and the concealment of the sample.
[0107] S5. Use the target model to verify the attack effect and transfer ability of the user adversarial sample.
[0108] Mix the generated fake user data X′ with the original user data X and inject it into the target model M target .
[0109] By evaluating the difference between the recommendation results of different target models under the input of the adversarial sample and the normal situation, calculate the attack success rate, and indirectly verify the transferability of the adversarial sample. The attack success rate can be defined as using the hit rate of the target item as an index to evaluate the attack effect among a certain number of test samples. If one of the items appears in the ranking list, it will be regarded as a popular item. The more target items, the more significant the attack performance and the higher the stability. Here, it is necessary to measure HR@k on the target item set (here k takes 50, and HR@k is to measure the average score of the target items that appear in the top 50 recommendation lists of normal users), so that more target items appear in the target model recommendation list. The transfer ability is equivalent to the attack effect tested on the target model with a different structure from the proxy model.
[0110] Using the method of the present invention, the effect is as Figure 4 shown, Figure 4This is a distribution map of transferable fake users among real users provided by an embodiment of the present invention, presenting the distribution of transferable fake users among real users. In a visual way, it helps to intuitively observe the positions and distribution characteristics of fake users in the overall real user group, providing an intuitive reference basis for analyzing the attack effect and transferability.
[0111] As Figure 1 shown, the method of the present invention for obtaining user adversarial samples has an overall structure of a double-layer optimization problem of an internal problem and an external problem.
[0112] The attacker first learns transferable fake users from the surrogate model, then injects the learned fake users and real users into the target model, and finally outputs a recommended list of user adversarial samples. The formula can be generally expressed as:
[0113]
[0114] Among them, formula (11) is the expression formula for optimizing the external problem, which can be considered as a general expression of formulas (5), (8) and (9). and respectively represent the loss function for internal training of the surrogate model and the adversarial loss function outside the surrogate model in the double-layer optimization.
[0115] The external problem of the double-layer optimization, that is, formula (11), is to learn the fake user X′ under the optimal model parameter θ R to optimize the attack of the objective function against the attack. The external problem needs to be calculated through the optimal solution of the internal problem.
[0116] The internal problem, that is, formula (12), is to retrain the surrogate model injected with the fake user data X′ to obtain the optimal model parameter θ R .
[0117] The double-layer optimization problem of the internal problem and the external problem makes the parameters of one optimization problem be constrained by the optimal solution and the optimal solution of another optimization problem. This double-layer optimization structure helps the attacker to optimize the situation of the adversarial target by learning the behavior of fake users.
[0118] The present invention also provides an electronic device. Figure 5 This is a schematic structural diagram of the electronic device provided by an embodiment of the present invention. As Figure 5As shown, the electronic device may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. The processor may call the logical instructions in the memory to execute the following method, for example:
[0119] S1. Randomly sample the original user-item interaction data to generate initial fake user data, and select recommendation models with different model architectures and hyperparameter configurations as surrogate models;
[0120] S2. Input the fake user data into different surrogate models, and update it by backpropagation according to the loss function of the surrogate model to obtain the optimal model parameters;
[0121] S3. Combine the adversarial loss function of the surrogate model, calculate the gradient of the surrogate model with respect to the optimal model parameters, and standardize it. Use similarity to calculate the importance distribution of the gradients of each surrogate model, and obtain the global gradient by weighting;
[0122] S4. Update the fake user data according to the global gradient, and iterate steps S2 - S3 to obtain the final user adversarial samples;
[0123] S5. Use the target model to verify the attack effect and transfer ability of the user adversarial samples.
[0124] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0125] The embodiments of the present invention also provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the methods provided in the above-mentioned embodiments, for example, including:
[0126] S1. Randomly sample the original user-item interaction data, generate initial fake user data, and select recommendation models with different model architectures and hyperparameter configurations as surrogate models.
[0127] S2. Input the fake user data into different surrogate models, and update them by backpropagation according to the loss function of the surrogate models to obtain the optimal model parameters.
[0128] S3. Combine the adversarial loss function of the surrogate models, calculate the gradient of the surrogate models with respect to the optimal model parameters, and standardize it. Use similarity to calculate the importance distribution of the gradients of each surrogate model, and obtain the global gradient by weighting.
[0129] S4. Update the fake user data according to the global gradient, and iterate steps S2 - S3 to obtain the final user adversarial samples.
[0130] S5. Use the target model to verify the attack effect and transfer ability of the user adversarial samples.
[0131] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0132] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A gradient ensemble-based transferable black-box attack method for recommendation systems, characterized in that: include: S1, randomly sample the original user-item interaction data, generate initial false user data, and select recommendation models with different model architectures and hyperparameter configurations as proxy models; S2. Input the fake user data into different proxy models, and update them through back propagation according to the loss function of the proxy model training to obtain the optimal model parameters of each proxy model. S3. Combine the proxy model with the adversarial loss function, calculate the gradient of the proxy model with respect to the optimal model parameters, and standardize it. Use the similarity to calculate the importance distribution of the gradients of each proxy model, and obtain the global gradient by weight. S4. Update the fake user data according to the global gradient, iterate steps S2 to S3, and obtain the final user adversarial sample; S5. Use the target model to verify the attack effect and migration ability of the user's adversarial samples.
2. According to claim 1, a gradient ensemble-based recommendation system transferable black-box attack method is characterized in that: The initial false user data is generated in S1, and the formula is expressed as: X′=RandomSample(X,k) (1) Among them, X is the original user set, X′ is the fake user set generated by random sampling, k is the number of sampled items, and the sampled items are used as the target item set.
3. According to claim 1, a gradient ensemble-based recommendation system transferable black-box attack method is characterized in that: In S2, the loss function of the proxy model training is used to back-propagate and update the parameters to obtain the optimal model parameters. The formula is: Where X∈R n×m is the original user rating matrix, n is the number of real users, and m is the number of all items; X′∈R n′×m is the rating matrix of fake users generated by the proxy model, where n′ represents the number of fake users, θ R is the optimal model parameter of the recommended model, θ' R represents the parameters of the recommendation model trained using the original users, θ” R Represents the parameters of the recommendation model trained using fake users.
4. According to claim 3, a gradient ensemble-based recommendation system transferable black-box attack method is characterized in that: Calculate the gradient of the proxy model with respect to the fake user, the formula is expressed as: Among them, L attack_i is the adversarial loss function of the proxy model i, θ R are the optimal model parameters.
5. According to claim 4, a gradient ensemble-based recommendation system transferable black box attack method is characterized in that: The adversarial attack loss function is expressed as: Among them, z is the target project of interest, x uz is the predicted rating or preference score of user u for the target item z, x ui Predict a rating or preference score for item i for user u.
6. According to claim 4, a gradient ensemble-based recommendation system transferable black box attack method is characterized in that: In S3, normalization is performed and the global gradient is obtained by similarity weighting. The formula is expressed as: Among them, w i is the assigned weight of the agent model i, is the gradient of the proxy model i.
7. The method for migratable black-box attack on recommendation system based on gradient integration according to claim 6 is characterized in that: The assigned weight w of the proxy model i i , the formula is: Among them, s i is the importance coefficient of the gradient of the i-th proxy model output with respect to the fake user.
8. The method for migratable black-box attack on recommendation system based on gradient integration according to claim 1 is characterized in that: In S4, the false user data is updated according to the global gradient, and the formula is expressed as: Among them, β is the update step size, Proj Λ (·) is the projected gradient operator. Through the continuous iterative update process, the fake user data is made close to the original user data distribution, which is not easy to be detected and is updated in the direction that can successfully interfere with the recommendation model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of a gradient integration-based migration black-box attack method for a recommendation system are implemented as claimed in any one of claims 1 to 8.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a gradient integration-based recommendation system transferable black-box attack method as claimed in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Transferable black box adversarial attack sample generation method and system and electronic equipment
CN115544499A