Model stealing attack method aiming at recommendation system and system thereof

The LLM Ranker with MC and PS modules generates high-quality synthetic data to address data distribution mismatch and position bias, enhancing model extraction attacks in sequence recommendation systems.

CN120316348APending Publication Date: 2025-07-15INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510451777.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing model stealing attack methods have problems with data distribution mismatch and position deviation in the sequence recommendation system, which leads to the inability to effectively simulate real user behavior, affecting the efficiency and fidelity of model stealing attacks.

Method used

A large language model (LLM) sorter is used, combining memory compression module and preference stability module to generate synthetic data that conforms to user behavior patterns, and train alternative models through knowledge distillation to eliminate position deviations and improve data quality.

Benefits of technology

It significantly improves the efficiency and fidelity of model theft attack, and the generated synthetic data is closer to real user behavior, enhancing the effect of model theft.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316348A_ABST
    Figure CN120316348A_ABST
Patent Text Reader

Abstract

The invention discloses a model stealing attack method and system for a recommendation system, and the method comprises the steps: simulating a real user behavior through a large language model LLM sequencer, and generating synthetic data which accords with a user behavior mode; wherein the LLM sequencer comprises a memory compression MC module and a preference stabilization PS module, and the memory compression MC module selectively retains effective historical interaction stored in the LLM; a preference stabilization (PS) module extracts a summary of user preferences from historical interactions stored in the LLM; and training an alternative model based on the synthetic data, and realizing model stealing by adopting a knowledge distillation method. According to the method, synthetic data with higher representativeness and wider coverage can be generated, so that an attacker can steal the target model more efficiently, and meanwhile, the generated data can reflect preferences and behavior patterns of real users better, so that the attack effect and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a model stealing attack method based on large language models (LLMs), and particularly to a black-box model extraction and stealing attack method and system for a sequential recommendation system. Background Art

[0002] In recent years, model extraction attacks have become a major threat to recommendation systems, especially in fields such as e-commerce and social media. When providing personalized recommendations, recommendation systems may face risks from attackers who, through interaction with the target recommendation system, obtain training data or use this data to train substitute models, thereby replicating the functions of the target system and even launching further attacks through the substitute models. In actual attack scenarios, since the original training data is usually unavailable, attackers typically use public data or synthetic data to organize training data, which is used to train substitute models, threatening the security and privacy of the system.

[0003] The method of using public datasets for model extraction attacks has achieved certain results in tasks such as image and text. However, in sequential recommendation systems, since the items in the recommendation system are usually represented by IDs, these public datasets do not match the items of the target recommendation system and cannot provide effective input data. This makes it impractical to use public data for model extraction attacks on sequential recommendation systems. In addition, some methods use synthetic data for attacks, but their core decision-making mechanisms are all random sampling, failing to consider the consistency of sequential patterns and user preferences. Synthetic data often lacks the characteristics of real user interactions, resulting in a significant distribution difference between synthetic data and real data.

[0004] In addition, with the powerful capabilities of large language models (LLMs) demonstrated in multiple fields, especially in simulating human behavior and understanding complex contexts, in recent years, some research has begun to attempt to apply LLMs to recommendation systems. These studies have shown that LLMs can effectively simulate the interaction behavior between users and recommendation systems, maintain consistent user preferences, and make reasonable decisions based on recommendation results. However, despite the great potential of LLMs in generating high-quality synthetic data, they also face some problems, such as how to effectively utilize long historical interaction data and how to maintain stable preferences.

[0005] Based on the above analysis, existing model extraction attack methods can generally be divided into model extraction based on public data and model extraction based on synthetic data;

[0006] Model Stealing Based on Public Data: In practical applications, using public data as substitute data for model stealing attacks is usually restricted. Especially when the items in the target recommendation system are not related to the items in the public data, the public data cannot provide effective input for model stealing attacks. This is particularly prominent in recommendation systems because the items in recommendation systems are often represented by unique IDs and cannot be directly mapped between different platforms. Therefore, public data sets often fail to meet the attack requirements when generating model stealing attacks against recommendation systems.

[0007] Model Stealing Based on Synthetic Data: Although synthetic data alleviates the problem of lack of data to a certain extent, it lacks continuous attention to sequential patterns and preferences. For example, the random sampling mechanism ignores the internal structure of the sequence and the stability of user preferences, resulting in a large distribution difference between synthetic data and real data, which limits the performance of model stealing attacks.

[0008] In addition, the autoregressive sampling based on the DFME method leads to the problems of overexposure or underexposure, which further affects the quality of the generated data and the effect of the attack.

[0009] Based on the above analysis, there is an urgent need to research and develop a technology that can generate high-quality synthetic data by simulating real user behaviors, improving the stealing efficiency and model fidelity. Summary of the Invention

[0010] To solve the following problems existing in the existing data generation methods for model stealing attacks without data:

[0011] 1) Data distribution mismatch: The synthetic data generated by existing methods cannot effectively simulate real user behaviors, resulting in low performance of substitute models;

[0012] 2) Exposure bias and position bias: The autoregressive generation framework overly relies on the recommendation results of the target model, causing uneven exposure of items; The LLM Ranker has obvious preference differences for items in different positions in the recommendation list, which is called position bias;

[0013] The present invention proposes a model stealing method for black-box sequential recommendation systems based on large language models, which generates high-quality synthetic data by simulating real user behaviors and debiasing operations, significantly improving the attack efficiency and model fidelity.

[0014] In a first aspect, an embodiment of the present application provides a model stealing attack method for a recommendation system, and the method includes:

[0015] Data generation steps based on large language models: Simulate real user behavior through the LLM sorter to generate synthetic data that conforms to the user behavior pattern; among them, the LLM sorter includes: Memory Compression (MC) module: used to selectively retain the effective historical interactions stored in the LLM; Preference Stability (PS) module: used to extract the summary of user preferences from the historical interactions stored in the LLM.

[0016] Model stealing steps: Based on the synthetic data, train a surrogate model and use the knowledge distillation method to achieve model stealing.

[0017] In a specific embodiment of the present invention, the above model stealing attack method for a recommendation system further includes:

[0018] Debiasing processing steps: After randomly shuffling the recommendation list to be input into the LLM sorter, it is used in combination with the LLM sorter so that the result selected by the sorter is not affected by the position of the item in the recommendation list.

[0019] In a specific embodiment of the present invention, the above data generation steps based on large language models further include:

[0020] Memory compression steps: When the number of historical interactions exceeds the predetermined memory size, the MC module retains the earliest and latest interaction items in the historical interactions through a selective retention strategy;

[0021] Preference stability steps: Generate a summary of user preferences by analyzing historical interaction information, so that the LLM sorter follows a consistent preference pattern when selecting items.

[0022] In a specific embodiment of the present invention, the above debiasing processing steps further include:

[0023] Calculate the expected value using the number of sampling times of the LLM sorter to measure whether the coverage rate of the target item in the preset historical interactions is achieved;

[0024] Randomly shuffle the recommendation list to be input into the LLM sorter so that the result selected by the sorter is not affected by the position of the target item in the recommendation list.

[0025] In a specific embodiment of the present invention, the above memory compression steps further include: The MC module selects the most representative interaction items according to the sequence of input historical interactions and the maximum memory length after compression;

[0026] When the number of historical interaction sequences is less than or equal to the maximum memory length, the MC module returns the complete interaction history;

[0027] When the number of the historical interaction sequences is greater than the maximum memory length, the MC module retains the first half of the items of the previous maximum memory length and the last half of the items of the latest maximum memory length through a selective retention strategy, and discards the intermediate items.

[0028] In a specific embodiment of the present invention, the above preference stabilization step further includes:

[0029] When the historical interaction reaches a preset length, the PS module is activated;

[0030] The PS module generates a preference summary from the historical interaction to capture key behavioral patterns, where the behavioral patterns include the types of frequently selected objects and context dependence.

[0031] In a specific embodiment of the present invention, the above model stealing step includes:

[0032] Minimizing the difference between the target model and the surrogate model through a knowledge distillation method;

[0033] The distillation loss is optimized by comparing the output of the target model with the output of the surrogate model, so that the ranking of the surrogate model is as close as possible to the ranking of the target model, realizing model stealing.

[0034] In a second aspect, an embodiment of the present application provides a model stealing attack system for a recommendation system, adopting the model stealing attack method for a recommendation system as described above. The system includes:

[0035] A data generation module based on a large language model: simulating real user behavior through an LLM sorter to generate synthetic data that conforms to the user behavior pattern; where the LLM sorter includes: a memory compression MC module: used to selectively retain valid historical interactions stored in the LLM; a preference stabilization PS module: used to extract a summary of user preferences from the historical interactions stored in the LLM;

[0036] A model stealing module: training a surrogate model based on the synthetic data and realizing model stealing by using a knowledge distillation method;

[0037] A debiasing processing module: after randomly sampling and shuffling the recommendation list to be input into the LLM sorter, it is used in combination with the LLM sorter to make the result selected by the sorter not affected by the position of the item in the recommendation list.

[0038] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the model stealing attack method for a recommendation system as described above are implemented.

[0039] Fourthly, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the model stealing attack method for the recommendation system as described above are implemented.

[0040] Compared with the related prior art, it has the following outstanding beneficial effects:

[0041] 1) The method of the present invention uses an LLM Ranker to simulate user behavior and selects the next interaction item from the recommendation results given by the recommendation system in combination with the user's historical interactions. The LLM Ranker is the core of the present invention. It selects the next item to interact with the user from the recommendation list of the target recommendation system and makes decisions based on historical interactions and user preferences. Specifically, it includes a memory compression module, a preference stability module, and corresponding debiasing techniques:

[0042] 2) The method of the present invention proposes a Memory Compression module (MC): dynamically compresses long interaction sequences, retains representative items (the first and last halves), reduces computational overhead, and maintains long-term and short-term behavior patterns. Technical effect: The Memory Compression module helps the ranker effectively process overly long historical records by selectively retaining the most representative historical interactions, thus avoiding excessive computational overhead.

[0043] 3) The method of the present invention proposes a Preference Stability module (PS): extracts a summary of user preferences from historical behaviors, uses an LLM to process historical interactions and generates a preference summary to capture key behavior patterns, including frequently selected item types and context dependencies. By integrating this stable preference information into the decision-making process, the ranker can simulate consistent behavior patterns, ensuring that its actions are consistent with the long-term trends of user behavior rather than being overly influenced by short-term fluctuations. Technical effect: The Preference Stability module extracts a summary of user preferences from historical interactions, ensuring that the ranker can maintain consistent preferences when selecting items and avoiding preference drift.

[0044] 4) The method of the present invention proposes a debiasing technique: Since the autoregressive framework may lead to exposure bias, the data generated by the ranker may have overexposure or underexposure of certain items. To mitigate these biases, the present invention combines a random sampler to generate items that do not appear in the recommendation results, covering 90% of the item space. The method of the present invention alleviates exposure bias; for the positional bias that exists when the LLM processes information, the present invention shuffles the order of the recommendation list and then inputs it into the LLM to eliminate the influence of the LLM positional bias. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0046] Figure 1 It is a schematic diagram of the model stealing attack method for the recommendation system according to the present invention;

[0047] Figure 2 It is a schematic diagram of the principle of the method framework of the embodiment of the present invention;

[0048] Figure 3 It is a schematic diagram of the Prompt template used in the LLM Ranker of the embodiment of the present invention;

[0049] Figure 4 It is a schematic diagram of the working process and the used Prompt template of the preference stability (PS) module of the embodiment of the present invention;

[0050] Figure 5 It is a schematic diagram of the model stealing attack system for the recommendation system according to the present invention;

[0051] Figure 6 It is a schematic diagram of the computer hardware according to the present invention. Detailed implementation manners

[0052] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (s) or plural items (s). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.

[0053] It should also be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship. Specifically, it can be understood by referring to the context before and after.

[0054] It should also be understood that in various embodiments of the present invention, the magnitude of the serial numbers of the above processes does not mean the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0055] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0056] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0057] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0058] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0059] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are given below and are described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments including the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.

[0060] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0061] When the method of the present invention conducts research on model stealing attacks on sequence recommendation systems, it is found that there are significant challenges in data generation and model stealing efficiency improvement in the existing technology. Especially in the black-box setting, the core mechanisms of existing methods often rely on simple random strategies. For example, the core selection strategy in the autoregressive data generation framework uses a random algorithm. This results in a large gap between the generated data and real user behaviors, making the stolen model have a large functional difference from the target recommendation system, and thus affecting the attack effect and model fidelity. Therefore, it is difficult for existing methods to be practically applied in generating high-quality data for model stealing. After analysis, the method of the present invention finds that the above defects mainly stem from the fact that existing data generation methods fail to effectively consider the complexity of user behaviors when simulating user behaviors. Under random generation or generation strategies based on simple rules, the generated data cannot effectively capture the preference changes and sequence patterns of real users when interacting with the recommendation system, resulting in a significant difference in distribution between the generated data and real data, thereby affecting the success rate of model stealing attacks.

[0062] The present invention aims to simulate the interaction behavior between real users and a recommendation system by using a large language model (LLM)-driven ranker, which can significantly improve the quality of synthetic data and narrow the gap with the training data of the target recommendation system. Specifically, the present invention proposes a framework that combines a Memory Compression (MC) module and a Preference Stabilization (PS) module, and uses the LLM-driven ranker to simulate users with stable preferences to generate more realistic simulation data. First, by introducing the memory compression module, a simple strategy is used to process the long interaction history of users into shorter sequences, which can effectively extract important items in the historical interaction, thereby reducing the computational overhead and enhancing the efficiency of the LLM in processing long historical data. Secondly, the preference stabilization module is introduced to summarize and solidify the preferences of users using the LLM at the initial stage of the interaction, ensuring that the ranker maintains consistent preferences throughout the interaction process, avoiding the problem of preference drift, and thus generating more stable and realistic interaction data that conforms to the behavior of real users. Finally, in order to further optimize the quality of the generated data, the present invention also designs a debiasing mechanism, which increases the diversity of the generated data by introducing random items, alleviates the inherent exposure bias problem in the autoregressive generation framework; by eliminating the sequential relationship of the item list input to the LLM, it avoids the potential position bias effect in the large language model. This mechanism increases the diversity of the generated data, further improves the data coverage, makes the generated data closer to real data, and thus enhances the effect of the model stealing attack.

[0063] In summary, by combining the advantages of the LLM to simulate user behavior, the present invention proposes a new model stealing attack method, effectively solving the deficiencies of the prior art in data generation and attack capabilities, being able to more accurately simulate the behavior of real users, generate high-quality synthetic data, and significantly improve the fidelity of model stealing.

[0064] The following will describe the method of the embodiments of the present application in detail with specific embodiments:

[0065] Embodiment 1

[0066] As Figure 1 and Figure 2 shown, Figure 1 is a schematic diagram of the model stealing attack method for the recommendation system of the present invention, Figure 2 is the schematic diagram of the framework principle of the specific embodiment of the present invention;

[0067] The framework of the present invention generates sequences autoregressively from random items. For each expansion, it queries the target model using the current sequence and shuffles the recommendation results to eliminate position bias. The LLM Ranker processes these recommendations through the Memory Compression (MC) and Preference Stabilization (PS) modules to achieve historical information compression and preference consistency, and then selects the most suitable item. The last step mitigates the exposure bias by introducing random items. Specifically, the embodiments of the present application provide a model stealing attack method for a recommendation system, and the method includes:

[0068] Data generation step 101 based on the large language model: Simulate real user behavior through the large language model LLM ranker to generate synthetic data that conforms to the user behavior pattern; wherein, the LLM ranker includes: Memory Compression MC module: used to selectively retain the effective historical interactions stored in the LLM; Preference Stabilization PS module: used to extract the summary of user preferences from the historical interactions stored in the LLM;

[0069] Debiasing processing step 102: After shuffling the recommendation list to be input into the LLM ranker through a random sampler, it is used in combination with the LLM ranker to make the result selected by the ranker not affected by the position of the item in the recommendation list.

[0070] Model stealing step 103: Based on the synthetic data, train an alternative model and use the knowledge distillation method to achieve model stealing.

[0071] In the specific embodiments of the present invention, the above data generation step 101 based on the large language model further includes:

[0072] Memory compression step: When the number of historical interactions exceeds the predetermined memory size, the MC module retains the earliest and latest interaction items in the historical interactions through a selective retention strategy;

[0073] Preference stabilization step: By analyzing the historical interaction information, generate a preference summary of the user, so that the LLM ranker follows a consistent preference pattern when selecting items.

[0074] In the specific embodiments of the present invention, the above memory compression step further includes: The MC module selects the most representative interaction items according to the sequence of the input historical interactions and the maximum memory length after compression; when the number of the historical interaction sequences is less than or equal to the maximum memory length, the MC module returns the complete interaction history; when the number of the historical interaction sequences is greater than the maximum memory length, the MC module retains the first half of the maximum memory length of items and the latest half of the maximum memory length of items through a selective retention strategy, and discards the intermediate items.

[0075] In a specific embodiment of the present invention, the above-mentioned preference stabilization step also includes: when the historical interaction reaches a preset length, the PS module is activated; the PS module generates a preference summary from the historical interaction to capture key behavior patterns, wherein the behavior patterns include the types of frequently selected targets and context dependencies.

[0076] Specifically, in a specific embodiment of the present invention, Figure 3 As shown in Figure 1, the prompt template used in LLM Ranker includes the outputs from the memory compression (MC) and preference stabilization (PS) modules, labeled as “Preference” and “Compressed Memory” respectively. “Rec List” represents the output from the target model.

[0077] In a specific embodiment of the present invention, from the item space of the target recommendation system Randomly select the initial item As the user's initial interaction item, Uni(·) represents a random distribution. At this point, the sorter's memory is empty, no user preferences have been established, and the interaction history is x (1) =[i1]. (1) Input into the target recommendation system to obtain a recommendation list. The sorter uses the memory compression (MC) module and the preference stabilization (PS) module to process historical interaction information to update its preferences and memory. The memory compression module compresses historical interactions into the most representative items to reduce computational overhead and information processing burden. When the number of historical interactions exceeds the predetermined memory size, the MC module retains only the earliest and latest interaction items through a selective retention strategy. The preference stabilization module generates a user preference summary by analyzing historical interaction information to ensure that the sorter can follow a consistent preference pattern when selecting items. Then Figure 3 As shown, the sorter selects the next interaction item i2 based on the information provided by the memory compression module and the preference stabilization module, as well as the recommendation list of the target recommendation system, to form x (2) =[i1,i2]. After each interaction, new items will be added to the interaction history to form x (n) =[i1,i2,…,i n ]. Continue the above process until the predetermined interaction history length is reached.

[0078] like Figure 4 As shown, the workflow of the Preference Stabilization (PS) module and the Prompt template used: LLM receives the platform description (“Platform Description”) and user history data (“History”) and generates a summary of user preferences.

[0079] Memory Compression (MC) Module: The historical interactions capture important user preferences and chronological information. To address the computational burden brought by overly long histories, the present invention introduces a Memory Compression (MC) module. Specifically, the MC module selects the most representative interaction items based on the input sequence x = [i1, i2, …, i T and a predefined maximum memory length size. When the sequence length is less than the memory length, i.e., T ≤ size, the MC module directly returns the complete interaction history; when the sequence length exceeds the memory length, i.e., T > size, the MC module retains the first items and the latest items through a selective retention strategy and discards the intermediate items. Formally, the processing of the MC module can be expressed as: This strategy retains long-term stable information while also maintaining recent behavioral patterns.

[0080] Preference Stability (PS) Module: To maintain consistent selection behavior, the LLM sorter requires stable preference guidance from historical interactions. The Preference Stability module addresses this need by extracting and maintaining a preference profile to guide the sorter's decisions. Considering the complexity of user preferences and the limited information in short sequences, the PS module is activated when the interaction history reaches a predefined length n (0 < n ≤ size). As Figure 4 shown, the LLM processes the historical interactions and generates a preference summary that captures key behavioral patterns, including frequently selected item types and context dependencies. By incorporating this stable preference information into the decision-making process, the sorter can simulate consistent behavioral patterns, ensuring that its actions are consistent with the long-term trends of user behavior rather than being overly influenced by short-term fluctuations.

[0081] In a specific embodiment of the present invention, the debiasing processing step 102 further includes:

[0082] Calculating the expected value using the sampling times of the LLM sorter to measure whether the coverage rate of the target object in the preset historical interactions is achieved; shuffling the order of the recommended list to be input to the LLM sorter so that the result selected by the sorter is not affected by the position of the target object in the recommended list.

[0083] Specifically, in a specific embodiment of the present invention, to reduce the bias introduced by the autoregressive generation framework and the LLM, debiasing techniques are adopted. The present invention combines a random sampler with the LLM sorter to ensure the diversity of the items selected each time, avoiding overexposure and underexposure of certain items due to exposure bias. Specifically, the present invention expects that 90% of the items are covered in the generated data (i.e., the number of items in the generated data represents the item space If the size, i.e., the total number of items, then the expected number of sampling times K using the random sampler can be calculated according to the formula In addition, to eliminate the influence of positional bias in the LLM, the present invention shuffles the order of the recommendation list to be input into the LLM to ensure that the result selected by the sorter is not affected by the position of the item in the recommendation list, thus better simulating the behavior of real users.

[0084] Acceleration process: To accelerate the autoregressive data generation process, in each round of generation, the sorter not only selects one item but multiple items by modifying the internal prompt of the LLM sorter. Multiple items can be selected in each iteration. This method not only reduces the computational overhead but also better simulates the selection behavior of multiple items that the user is interested in simultaneously in the recommendation list.

[0085] In a specific embodiment of the present invention, the above model stealing step 103 includes: minimizing the difference between the target model and the surrogate model through the knowledge distillation method; the distillation loss is optimized by comparing the output of the target model with the output of the surrogate model to make the ranking of the surrogate model as close as possible to the ranking of the target model, thereby achieving model stealing.

[0086] Through the above synthetic data generation process, the present invention generates a series of synthetic data that conforms to the user behavior pattern. Next, in a specific embodiment of the present invention, these generated synthetic data are used to train a surrogate model. Specifically, the present invention uses the sequence set generated by the above process to query the target model to construct a surrogate data set where each x represents a user interaction sequence, B is the size of the data set, and the target model f t takes the interaction history x of a certain user as input, and the recommended result output is represented by f t (x).

[0087] After that, the difference between the target model and the surrogate model is minimized through the knowledge distillation method. The distillation loss is optimized by comparing the output of the target model with the output of the surrogate model. The goal is to make the ranking of the surrogate model as close as possible to the ranking of the target model, thereby effectively replicating the function of the target model. The specific calculation method of the distillation loss is as follows:

[0088]

[0089] where represents the score of the surrogate model for an item in the top k recommendations of the target recommendation system, arranged according to the item sorting given by the target recommendation system For example, if then For the scores of the substitute model on randomly sampled negative sample items, λ1 and λ2 are hyperparameters used to control the loss. This loss function makes the behavior of the substitute model as consistent as possible with the target model by minimizing the ranking difference between the output of the substitute model and the output of the target model.

[0090] As described above, the method of the present invention can be preferably implemented.

[0091] The present invention proposes a novel model stealing framework. Combining a memory compression module, a preference stability module, and a debiasing processing technique, the present invention can effectively generate synthetic data similar to the training data of the target recommendation system, overcoming the problems of poor data generation quality and excessive bias in existing methods. Then, knowledge distillation technology is used for model stealing, significantly improving the effect of the model stealing attack and the fidelity of the model.

[0092] Compared with the prior art, the method of the present invention has the following advantages: The present invention generates higher-quality synthetic data by using large language models (LLMs) to simulate real user behaviors, enhancing the effect of the model stealing attack (MEA). First, the sorter receives a recommendation list through interaction with the target model, and then processes the historical interaction information through a memory compression (MC) module and a preference stability (PS) module. The memory compression module helps the sorter effectively process overly long historical records by selectively retaining the most representative historical interactions, thus avoiding excessive computational overhead. The preference stability module extracts a summary of the user's preferences from the historical interactions, ensuring that the sorter can maintain consistent preferences when selecting items and avoiding preference drift. The present invention introduces a debiasing processing mechanism to further improve the diversity and quality of data generation.

[0093] In this way, the present invention can generate more representative and comprehensive synthetic data, enabling the attacker to steal the target model more efficiently. At the same time, the generated data can better reflect the preferences and behavior patterns of real users, thus enhancing the effect and accuracy of the attack.

[0094] Embodiment 2

[0095] As Figure 5 shown, the embodiment of the present application provides a model stealing attack system for a recommendation system, adopting the model stealing attack method for a recommendation system as described above. The system includes:

[0096] A data generation module 201 based on a large language model: Simulating real user behaviors through the LLM sorter of the large language model to generate synthetic data conforming to the user behavior pattern; wherein, the LLM sorter includes: A memory compression MC module: Used to selectively retain the valid historical interactions stored in the LLM; A preference stability PS module: Used to extract a summary of the user's preferences from the historical interactions stored in the LLM;

[0097] Bias removal processing module 202: After shuffling the recommendation list to be input into the LLM sorter through a random sampler, it is used in combination with the LLM sorter so that the result selected by the sorter is not affected by the position of the item in the recommendation list;

[0098] Model stealing module 203: Based on synthetic data, train a surrogate model and use the knowledge distillation method to achieve model stealing.

[0099] Embodiment III

[0100] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the model stealing attack method for the recommendation system described above are implemented.

[0101] Embodiment IV

[0102] The embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the model stealing attack method for the recommendation system as described above are implemented.

[0103] In addition, combined with Figure 1 The model stealing attack method for the recommendation system described in the embodiment of the present application can be implemented by an electronic device, such as a computer device. Figure 6 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.

[0104] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 6 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.

[0105] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured with one or more integrated circuits for implementing the embodiments of the present application.

[0106] The memory 82 can be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.

[0107] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the model stealing attack methods for the recommendation system in the above embodiments.

[0108] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0109] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A model stealing attack method for a recommendation system, characterized in that, The method includes: Data generation step based on a large language model: Simulating real user behavior through a large language model (LLM) sorter to generate synthetic data that conforms to the user behavior pattern; wherein, the LLM sorter includes: a Memory Compression (MC) module: for selectively retaining valid historical interactions stored in the LLM; a Preference Stability (PS) module: for extracting a summary of user preferences from the historical interactions stored in the LLM. Model stealing step: Based on the synthetic data, training a surrogate model and implementing model stealing using the knowledge distillation method.

2. The model stealing attack method for a recommendation system according to claim 1, characterized in that, The method further includes: Debiasing processing step: After shuffling the recommendation list to be input into the LLM sorter by a random sampler, it is used in combination with the LLM sorter, so that the result selected by the sorter is not affected by the position of the item in the recommendation list.

3. The model stealing attack method for a recommendation system according to claim 1 or 2, characterized in that The data generation step based on the large language model further includes: Memory compression step: When the number of historical interactions exceeds the predetermined memory size, the MC module retains the earliest and latest interaction items in the historical interactions through a selective retention strategy. Preference stability step: By analyzing the historical interaction information, generating a preference summary of the user, so that the LLM sorter follows a consistent preference pattern when selecting items.

4. The model stealing attack method for a recommendation system according to claim 2, characterized in that, The debiasing processing step further includes: Calculating the expected value using the sampling times of the LLM sorter to measure whether the coverage rate of the target item in the preset historical interactions is achieved. Shuffling the recommendation list to be input into the LLM sorter, so that the result selected by the sorter is not affected by the position of the target item in the recommendation list.

5. The method for model stealing attack against a recommendation system according to claim 3, wherein The memory compression step further includes: The MC module selects the most representative interaction items according to the sequence of input historical interactions and the maximum memory length after compression. When the number of historical interaction sequences is less than or equal to the maximum memory length, the MC module returns the complete interaction history. When the number of historical interaction sequences is greater than the maximum memory length, the MC module retains the first half of the maximum memory length of items and the last half of the maximum memory length of items through a selective retention strategy, and discards the intermediate items.

6. The method for model stealing attack against a recommendation system according to claim 3, wherein The preference stability step further includes: When the historical interaction reaches the preset length, the PS module is activated. The PS module generates a preference summary from the historical interactions to capture key behavior patterns, wherein the behavior patterns include the types of frequently selected target items and context dependencies.

7. The method for model stealing attack against a recommendation system according to claim 1, characterized in that The model stealing step includes: Minimizing the difference between the target model and the surrogate model through the knowledge distillation method. The distillation loss is optimized by comparing the output of the target model with the output of the surrogate model, so that the ranking of the surrogate model is as close as possible to the ranking of the target model to achieve model stealing.

8. A model stealing attack system for a recommendation system, which adopts the model stealing attack method for a recommendation system as described in any one of claims 1-7, characterized in that, The system includes: Large Language Model-based Data Generation Module: Simulate real user behavior through the Large Language Model (LLM) sorter to generate synthetic data that conforms to the user behavior pattern; wherein, the LLM sorter includes: Memory Compression (MC) Module: used to selectively retain the valid historical interactions stored in the LLM; Preference Stability (PS) Module: used to extract the summary of user preferences from the historical interactions stored in the LLM. Model Stealing Module: Based on the synthetic data, train a surrogate model and implement model stealing using the knowledge distillation method. Debiasing Processing Module: After shuffling the recommendation list to be input into the LLM sorter through a random sampler, it is used in combination with the LLM sorter so that the result selected by the sorter is not affected by the position of the item in the recommendation list.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the model stealing attack method for a recommendation system described in any one of claims 1-7.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps of the model stealing attack method for a recommendation system described in any one of claims 1 to 7.