Unsupervised black box member reasoning attack method and system based on generative collaborative learning
Through the unsupervised black box member inference attack method with generative collaborative learning, heuristic tasks are used to generate pseudo-labels and offset vector training discriminators, solving the applicability and stability of the black box model in the prior art, and achieving high-performance member inference attacks.
Patent Information
- Application Number
- CN202510399011.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing large language model member reasoning attack method cannot be applied to completely black box models. The heuristic method is limited by cue sensitivity and performance fluctuations, making it difficult to stably adapt to different models and data sets.
The unsupervised black box member inference attack method with generative collaborative learning is adopted to generate pseudo-labels through multiple heuristic tasks, and the pre-trained encoder calculates the offset vector, and the training discriminator for member probability prediction, realizing automated attacks without relying on label information.
Implementing high-performance member reasoning under completely black box conditions improves the robustness and stability of the method and adapts to diversified actual needs.
Smart Images

Figure CN120278201A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross - technical field of large language model privacy security and unsupervised machine learning, and particularly to an unsupervised black - box membership inference attack method and system based on generative collaborative learning. Background Art
[0002] Large Language Models (LLMs) have demonstrated excellent performance in multiple fields. This is mainly due to their powerful generalization ability obtained from pre - training on massive datasets, and their parameter scale has exceeded hundreds of billions (e.g., GPT - 3 has exceeded 175 billion parameters). This data - driven training paradigm brings two key problems: First, the training data may contain un - desensitized personal privacy information (such as medical records, social media conversations), copyrighted content, or benchmark test datasets; Second, the model may over - fit to some samples during training and may generate text with a high confidence level when interacting with users, thus indirectly leaking sensitive information. Research shows that when a specific trigger prefix is input, models such as GPT - 2 can reproduce email addresses and phone numbers in the training data completely.
[0003] Membership Inference Attack (MIA) aims to determine whether a given data sample belongs to the training set of a target model. Its essence is to distinguish the behavioral differences of the model between "seen data" (training members) and "unseen data" (non - training members). However, now, more and more model developers are reluctant to disclose information about the source of their training corpora to avoid legal and ethical disputes. Therefore, membership inference attack techniques are widely used to evaluate the privacy risks of language models.
[0004] Existing large language model membership inference attack methods mainly include gray - box access methods and prompt - based heuristic methods.
[0005] Gray - box methods assume that the model shows a stronger memory pattern for the trained data, which is reflected in the probability distribution of tokens output inside the model. Therefore, by applying a certain calculation method to the token probabilities of the input sequence, a membership score is obtained, and then it is classified as a member / non - member through a threshold. Therefore, they must have access to the token - level probability distribution of the model. This makes them inapplicable to fully black - box models, such as commercial large language models like the Claude model and the ChatGPT model, which follow the interaction pattern of "input prompt → output response".
[0006] Prompt - based heuristic methods reveal potential memories through the responses generated by the model. They design specific task prompts for the model to answer, and the model's memory of the data can be reflected in its answers to induce the model to expose its memory behavior of the training data for membership inference. However, it is limited by the high sensitivity of large - language models to prompts. For example, adding different guiding words (such as "Please continue:" vs. "Complete the following text:") to the same prefix may lead to significantly different generation results. This sensitivity makes heuristic methods need to repeatedly adjust the prompt templates for different models and datasets, greatly increasing the deployment cost and posing challenges in terms of stability. In addition, in cross - model or cross - domain scenarios, the performance of heuristic methods fluctuates significantly. For example, the DE - COP algorithm performs excellently on book texts (BookTection dataset), but on academic papers (arXivTection dataset), the accuracy drops by nearly 30% due to the increased complexity of the text structure. This instability makes it difficult for existing methods to adapt to diverse practical needs. Summary of the Invention
[0007] In view of this, embodiments of the present invention provide an unsupervised black - box membership inference attack method and system based on generative collaborative learning to eliminate or improve one or more defects existing in the prior art.
[0008] On the one hand, the present invention provides an unsupervised black - box membership inference attack method based on generative collaborative learning, and the method includes the following steps:
[0009] Input the input samples into a preset plurality of heuristic tasks respectively, and calculate the task scores of each heuristic task; input the task scores into a generator, and aggregate the task scores through the generator to obtain the final membership score of the input sample; screen the input sample based on the final membership score and a preset confidence threshold; generate corresponding pseudo - labels according to the final membership score of the screened input sample; the pseudo - labels include training members and non - training members;
[0010] Input the screened input samples into a large - language model to generate output responses; calculate the offset vector between the input sample and the output response based on a pre - trained encoder;
[0011] Use the offset vector and the pseudo - labels generated by the generator to train the discriminator so that the discriminator can predict the membership probability of the input sample according to the offset vector; use the membership probability generated by the discriminator as the new pseudo - label of the input sample to optimize the generator; iteratively execute the cross - supervised training of the generator and the discriminator until the discriminator converges; use the trained discriminator as the final membership inference attack model;
[0012] Input the sample to be detected into the large language model to generate an output response for the sample to be detected; generate an offset vector between the sample to be detected and its output response based on the pre-trained encoder, and input the offset vector into the membership inference attack model to generate the membership probability of the sample to be detected, so as to determine whether the sample to be detected is a training member.
[0013] In some embodiments of the present invention, after calculating the task scores of each heuristic task, it further includes:
[0014] Normalize the task scores, and the normalization formula is:
[0015]
[0016] where s i represents the task score of the i-th heuristic task; s i ′ represents the task score of the i-th heuristic task after normalization.
[0017] In some embodiments of the present invention, aggregating the task scores through the generator to obtain the final membership score of the input sample includes:
[0018] Perform a weighted sum on the normalized task scores to obtain the final membership score of the input sample, and the calculation formula is:
[0019]
[0020] where s member represents the final membership score of the input sample; w i represents the weight of the i-th heuristic task; s i ′ represents the task score of the i-th heuristic task after normalization, and N represents the total number of the heuristic tasks.
[0021] In some embodiments of the present invention, screening the input sample based on the final membership score and a preset confidence threshold includes:
[0022] Calculate the distribution of the final membership score; according to the distribution of the final membership score and the preset confidence threshold, calculate the dynamic membership threshold and non-membership threshold through the inverse cumulative distribution function, and the calculation formula is:
[0023] θ member =F -1 (1-α);
[0024] θ non-member =F -1 (α);
[0025] where θ member represents the member threshold; θ non-member represents the non - member threshold; F -1 (·) represents the inverse cumulative distribution function; α represents the preset confidence threshold;
[0026] Select the input samples with the highest final member score greater than or equal to the member threshold and the lowest final member score less than or equal to the non - member threshold, accounting for the preset proportion of the confidence threshold.
[0027] In some embodiments of the present invention, generating corresponding pseudo - labels according to the final member scores of the filtered input samples includes:
[0028]
[0029] where represents the pseudo - label; s member represents the final member score of the input sample; θ member represents the member threshold; θ non-member represents the non - member threshold.
[0030] In some embodiments of the present invention, calculating the offset vector between the input sample and the output response based on a pre - trained encoder includes:
[0031] Generating embedding vectors of the input sample and the output response based on the pre - trained encoder, and calculating the offset vector between the input sample embedding vector and the output response embedding vector. The definition of the offset vector satisfies the following formula:
[0032] Δe = e(y)-e(x);
[0033] where Δe represents the offset vector; e(x) represents the embedding representation of the input sample x; e(y) represents the embedding representation of the output response y;
[0034] Constructing the training set of the discriminator based on the offset vector and the corresponding pseudo - label. The definition of the training set satisfies the following formula:
[0035]
[0036] where represents the training set of the discriminator ; Δe represents the offset vector; represents the pseudo - label.
[0037] In some embodiments of the present invention, training the discriminator using the offset vector and the pseudo-labels generated by the generator further includes:
[0038] Constructing a binary cross-entropy loss between the membership probability generated by the discriminator and the pseudo-labels generated by the generator, and training the discriminator with the goal of minimizing the binary cross-entropy loss;
[0039] The formula for the binary cross-entropy loss is:
[0040]
[0041] where, represents the binary cross-entropy loss of the discriminator ; I(·) represents the indicator function; represents the pseudo-label of the i-th input sample; N represents the total number of input samples; represents the membership probability predicted by the discriminator for the i-th input sample.
[0042] In some embodiments of the present invention, before inputting the input sample into the generator, it further includes:
[0043] Initializing the weights of each heuristic task in the generator;
[0044] The heuristic tasks at least include multiple-choice, name filling, self-checking, output consistency, and recall continuation.
[0045] In some embodiments of the present invention, taking the membership probability generated by the discriminator as the new pseudo-label of the input sample includes:
[0046] Calculating the distribution of the membership probability; according to the distribution of the membership probability and a preset confidence threshold, calculating a dynamic threshold through the inverse cumulative distribution function to screen out high-confidence input samples;
[0047] Using the screened input samples for training the generator again, and taking the membership probability generated by the discriminator as the new pseudo-label.
[0048] On the other hand, the present invention also provides an unsupervised black-box membership inference attack system based on generative collaborative learning, which is characterized in that when the system is executed, it implements the steps of any one of the methods mentioned above.
[0049] The present invention provides an unsupervised black-box membership inference attack method and system based on generative collaborative learning, including: inputting samples into multiple heuristic tasks, calculating the task scores of each heuristic task, and then using a generator to aggregate all task scores to generate sample pseudo-labels, capturing the model memory features from different dimensions without relying on any annotation information. Filter the generated pseudo-labels, only retain the samples with high confidence at both ends, and discard the intermediate noisy samples to improve the sample quality and ensure the robustness of the model in complex scenarios. Use a pre-trained encoder to quantify the offset vector between the sample input prompt and the output response to obtain the internal state of the large language model, replacing the token probability distribution information in the traditional grey-box method to break through the black-box limitation. The discriminator takes the offset vector as input and is trained based on the pseudo-labels generated by the generator, enabling the discriminator to predict the membership probability of the sample according to the offset vector. During training, the generator provides pseudo-labels for the training of the discriminator, and the membership probability generated by the discriminator serves as the new pseudo-labels of the samples to help optimize the generator. The two are iteratively optimized through a cross-supervision mechanism until the discriminator converges. Based on the trained discriminator, high-performance membership inference can be achieved under completely black-box and unsupervised conditions, promoting the practical development of large language model privacy and security technologies.
[0050] Additional advantages, objects, and features of the present invention will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following, or can be learned from the practice of the present invention. The objects and other advantages of the present invention can be realized and obtained by the structure specifically pointed out in the specification and the drawings.
[0051] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. Description of the Drawings
[0052] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:
[0053] Figure 1 It is a schematic diagram of the steps of an unsupervised black-box membership inference attack method based on generative collaborative learning in an embodiment of the present invention.
[0054] Figure 2 It is an overall framework diagram of an unsupervised black-box membership inference attack method based on generative collaborative learning in an embodiment of the present invention. Detailed Embodiments
[0055] To make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the embodiments and the accompanying drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.
[0056] Herein, it should also be noted that in order to avoid obscuring the present invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less relevant to the present invention are omitted.
[0057] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0058] Herein, it should also be noted that if not otherwise specified, the term "connection" in this document can refer not only to direct connection, but also to indirect connection with an intermediate.
[0059] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0060] It should be emphasized here that the step labels mentioned below are not intended to limit the order of the steps. Instead, it should be understood that the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0061] To solve the problems in the existing large language model membership inference attack methods, where the grey-box method relies on the internal token probability distribution of the model and cannot be applied to completely black-box models, and the prompt-based heuristic method is limited by prompt sensitivity, poor generalization ability, large performance fluctuations, and difficulty in stably adapting to different models and datasets, the present invention proposes an unsupervised black-box membership inference attack method based on generative cooperative learning (Generative Cooperative Learning framework for Unsupervised Membership Inference Attacks, GCL-UMIA). Through the interaction and cooperation between the heuristic task and the discriminative task, an automated membership inference attack without relying on any labeled data is achieved, as Figure 1 shown, the method includes the following steps S101 to S104:
[0062] Step S101: Input the input samples into a preset plurality of heuristic tasks respectively, calculate the task scores of each heuristic task; input the task scores into a generator, and aggregate the task scores through the generator to obtain the final membership score of the input samples; screen the input samples based on the final membership score and a preset confidence threshold; generate corresponding pseudo-labels according to the final membership scores of the screened input samples. Among them, the pseudo-labels include training members and non-training members.
[0063] Step S102: Input the screened input samples into a large language model to generate an output response; calculate the offset vector between the input samples and the output response based on a pre-trained encoder.
[0064] Step S103: Use the offset vector and the pseudo-labels generated by the generator to train a discriminator, so that the discriminator can predict the membership probability of the input samples according to the offset vector; use the membership probability generated by the discriminator as the new pseudo-labels of the input samples, and use the input samples and their new pseudo-labels to train the generator again; iteratively execute the cross-supervised training of the generator and the discriminator until the discriminator converges; use the trained discriminator as the final membership inference attack model.
[0065] Step S104: Input the sample to be detected into a large language model to generate an output response for the sample to be detected; generate an offset vector between the sample to be detected and its output response based on a pre-trained encoder, input the offset vector into the membership inference attack model, and generate the membership probability of the sample to be detected to determine whether the sample to be detected is a training member.
[0066] As Figure 2 shown, it is the overall framework diagram of an unsupervised black-box membership inference attack method based on generative collaborative learning. The framework includes a generator, a discriminator and a collaborative framework. The generator generates pseudo-labels through a preset plurality of heuristic tasks, and these pseudo-labels are filtered before training the discriminator to remove noise. The discriminator is trained based on the difference between the input and output embedding vectors and the generated pseudo-labels. The new pseudo-labels generated by the trained discriminator are used to improve the generator. This process is iterated alternately, and finally a trained discriminator that can automatically perform membership inference attacks is obtained.
[0067] Specifically, in step S101, the generator is described.
[0068] The goal of the generator is to evaluate whether a given text (i.e., the input sample) belongs to the training members of a specific large language model through a plurality of heuristic tasks, so as to generate pseudo-labels for the training of the discriminator.
[0069] In some embodiments, to generate pseudo-labels, five heuristic tasks related to membership inference attacks are considered, namely: multiple-choice, name filling, self-checking, output consistency, and recall continuation. The scores of each task reflect different aspects of membership characteristics.
[0070] Multiple-choice: The large language model is required to select the original text from a given passage and its rewritten version. The model is more likely to correctly answer samples that appear in its training data.
[0071] Name filling: Mask the proper nouns in the sample and require the large language model to predict the masked word. This task can reveal whether the model has memorized specific training content.
[0072] Self-checking: Require the large language model to determine whether the given text belongs to its training dataset, enabling the model to evaluate its familiarity with the given sample.
[0073] Output consistency: The model generates multiple continuations for the same input and calculates the consistency between them. For trained samples, the model tends to generate consistent responses.
[0074] Recall continuation: Divide the sample into a prefix and a suffix, and require the large language model to continue writing from the prefix. The semantic similarity between the continuation and the suffix reflects the model's memory level.
[0075] In some embodiments, during the system cold start phase, the weights of each heuristic task are first initialized. Taking the above embodiments as an example, there are five heuristic tasks in the generator: multiple-choice, name filling, self-checking, output consistency, and recall continuation. Based on the reliability verified by existing research, the weight of the multiple-choice task is set to 0.6, the self-checking task is set to 0.3, and the remaining tasks (name filling, output consistency, recall continuation) are each set to 0.033.
[0076] Input the input sample into a preset multiple heuristic tasks, calculate the task scores of the input sample in each heuristic task respectively, and then the generator aggregates the task scores to obtain the final membership score of the input sample.
[0077] In some embodiments, to facilitate the subsequent aggregation of each task score, the task scores are first normalized to a consistent range. The normalization process is shown in formula (1):
[0078]
[0079] where s i represents the task score of the i-th heuristic task; s i ′Represents the task score of the $i$-th heuristic task after normalization, ensuring comparability between different tasks.
[0080] Considering that the weight of each task should vary according to the input samples and the large language model, and to enhance the generalization ability and stability of the method, in this invention, the weights of each heuristic task are regarded as variables affected by the input samples, and a neural network is designed to calculate the weights of each heuristic task. The calculation formula is shown in Formula (2):
[0081] w i = Softmax(W·x + b); (2)
[0082] where, $w$ i represents the weight of the $i$-th heuristic task, which is part of the parameters of the generator, with an initial value randomly set and can be optimized through subsequent training; $W$ represents the weight matrix; $x$ represents the semantic representation of the input sample; $b$ represents the bias term. After applying the Softmax function, the weight of each heuristic task is mapped to the interval [0, 1].
[0083] Based on the normalized task scores and weights, weighted aggregation is performed to obtain the final membership score of the input sample. The calculation formula is shown in Formula (3):
[0084]
[0085] where, $s$ member represents the final membership score of the input sample; $w$ i represents the weight of the $i$-th heuristic task; $s$ i ′ represents the normalized task score of the $i$-th heuristic task, and $N$ represents the total number of heuristic tasks. Exemplarily, $N = 5$.
[0086] To improve the reliability of the pseudo-labels and reduce noise, the input samples are screened based on the final membership scores of the input samples and a preset confidence threshold. The reliable input samples are selected for the training of the discriminator. That is, a confidence threshold $\alpha$ is preset. In each iteration, only the top $\alpha$ input samples with the highest and lowest final membership scores are selected as reliable training members and non-training members for the training of the discriminator.
[0087] Specifically, first calculate the distribution of the final membership scores of all input samples, and calculate the dynamic membership threshold and non-membership threshold through the inverse cumulative distribution function. The calculation formulas are shown in Formulas (4) and (5) respectively:
[0088] $\theta$ member = $F$ -1 (1 - $\alpha$); (4)
[0089] $\theta$non-member = F -1 (α); (5)
[0090] Among them, θ member represents the member threshold; θ non-member represents the non - member threshold; F -1 (·) represents the inverse cumulative distribution function; α represents the preset confidence threshold.
[0091] Select the top α - proportion of input samples with the highest and greater than or equal to the member threshold (s member ≥ θ member ), and the lowest and less than or equal to the non - member threshold (s member ≤ θ non-member ). Preferably, α is set to 0.15 or 0.2.
[0092] In some embodiments, a mapping function is introduced to generate corresponding pseudo - labels according to the final member scores of the filtered input samples, as shown in formula (6):
[0093]
[0094] Among them, represents the pseudo - label; s member represents the final member score of the input sample; θ member represents the member threshold; θ non-member represents the non - member threshold.
[0095] After screening, represents reliable training members, represents reliable non - training members, are low - confidence samples and will be directly filtered out. Only the input samples corresponding to reliable training members and non - training members are used to train the discriminator.
[0096] In step S102, the input of the discriminator is described.
[0097] One of the reasons why existing grey - box access member inference attack methods perform excellently is that they use the output token distribution of the large - language model as the input of the discriminator, and these distributions encode fine - grained features reflecting the internal state of the model. However, in the black - box setting, only the "input prompt" and "output response" of the large - language model can be accessed, and the internal state cannot be obtained. To solve this problem, the present invention proposes to obtain input data from the black - box large - language model, that is, to use the changes of the "input prompt" and "output response" in the embedding space as the input of the discriminator. This idea is based on the following observation: the "output response" is affected by both the "input prompt" and the internal parameters of the large - language model, that is, the offset between the "input prompt" and the "output response" indirectly reflects the influence of the internal state of the large - language model.
[0098] Based on the above analysis, in the present invention, the filtered input samples are input into the large language model to generate an output response. Then, based on the pre-trained encoder (such as the BERT model, etc.), the embedding vectors of the input samples and the output response are generated, and the offset vector between the embedding vector of the input samples and the embedding vector of the output response is calculated. The definition of the offset vector satisfies the following formula (7):
[0099] Δe = e(y) - e(x); (7)
[0100] Wherein, Δe represents the offset vector; e(x) represents the embedding representation of the input sample x; e(y) represents the embedding representation of the output response y.
[0101] By calculating the offset vector between the input samples and the output response, more accurate and comprehensive input features can be obtained for the discriminator, so that it is not necessary to rely on the internal state information of the large language model.
[0102] In step S103, cross-supervision training is performed on the generator and the discriminator.
[0103] Based on step S101, the generator provides the pseudo-labels required for the discriminator's training.
[0104] Based on the offset vector calculated by formula (7) and the pseudo-labels generated by the generator, a training set for the discriminator is constructed. The definition of the training set satisfies the following formula (8):
[0105]
[0106] Wherein, represents the training set of the discriminator ; Δe represents the offset vector; represents the pseudo-label.
[0107] The discriminator is a binary classification network. By inputting the training set into the discriminator for prediction, the membership probability of each input sample can be obtained.
[0108] The discriminator is trained using the training set, that is, the discriminator takes the offset vector and the pseudo-labels generated by the generator as inputs and outputs the membership probability of each input sample.
[0109] The membership probability generated by the discriminator is used as the new pseudo-label of the corresponding input sample to optimize the generator.
[0110] In some embodiments, during the generator training optimization process, a loss function is constructed between the final member scores and the new pseudo-labels generated by the discriminator. With the goal of minimizing this loss function, the generator is optimized, that is, the weight parameters of each heuristic task are optimized. The weight update process enables the generator to adapt to the feature distributions of different samples and the response characteristics of the target model, thereby enhancing the aggregation effect of the heuristic tasks.
[0111] In some embodiments, consistent with the input sample screening mechanism mentioned above, the distribution of the membership probabilities of each input sample is calculated. Based on the distribution of the membership probabilities and a preset confidence threshold, a dynamic threshold is calculated through the inverse cumulative distribution function to screen out new pseudo-labels with high confidence. The screened new pseudo-labels are then used in the optimization of the generator to filter out low-confidence pseudo-labels.
[0112] In some embodiments, during the discriminator training process, a binary cross-entropy loss is constructed between the membership probabilities generated by the discriminator and the pseudo-labels generated by the generator. With the goal of minimizing the binary cross-entropy loss, the discriminator is trained. Among them, the binary cross-entropy loss is shown in formula (9):
[0113]
[0114] Among them, represents the binary cross-entropy loss of the discriminator ; I(·) represents the indicator function; represents the pseudo-label of the i-th input sample; N represents the total number of input samples; represents the membership probability predicted by the discriminator for the i-th input sample.
[0115] Based on the above training steps, the generator and the discriminator alternately use the pseudo-labels generated by each other for training, and improve the membership detection performance by iteratively optimizing the pseudo-labels. The generator optimizes the weights of the heuristic tasks to generate more reliable pseudo-labels, and the discriminator uses high-quality pseudo-labels to optimize the classification boundary. The iteration continues until the discriminator converges, such as the AUC score of the discriminator fluctuates less than 1%, and the training stops. The trained discriminator is used as the final membership inference attack model.
[0116] In step S104, the trained discriminator, that is, the membership inference attack model, can be directly deployed in the black-box large language model scenario. The automated inference includes the following steps:
[0117] Obtain the sample to be detected, input the sample to be detected into the large language model, and generate the corresponding output response.
[0118] Based on the pre-trained encoder (such as the BERT model), an offset vector between the input representation and the output response of the sample to be detected is generated.
[0119] Input the offset vector into the membership inference attack model to generate the membership probability of the sample to be detected. According to a preset probability threshold, for example, the probability threshold is set to 0.5. If the generated membership probability is greater than 0.5, it is determined that the sample to be detected belongs to the training members; otherwise, it belongs to non-training members.
[0120] Correspondingly to the above method, the present invention further provides an unsupervised black-box membership inference attack system based on generative collaborative learning. When the system is executed, it can implement the steps of the unsupervised black-box membership inference attack method based on generative collaborative learning.
[0121] Correspondingly to the above method, the present invention further provides an electronic device, which includes a computer device. The computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the electronic device implements the steps of the method described above.
[0122] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the foregoing edge computing server deployment method. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0123] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement it in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link.
[0124] It should be clear that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present invention is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0125] In the present invention, features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0126] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An unsupervised black-box membership inference attack method based on generative collaborative learning, characterized in that The method includes the following steps: Input the input samples into a plurality of preset heuristic tasks respectively, and calculate the task scores of each heuristic task; input the task scores into a generator, and aggregate the task scores through the generator to obtain the final membership score of the input samples; screen the input samples based on the final membership score and a preset confidence threshold; generate corresponding pseudo-labels according to the final membership score of the screened input samples; the pseudo-labels include training members and non-training members; Input the screened input samples into a large language model to generate an output response; calculate the offset vector between the input samples and the output response based on a pre-trained encoder; Train the discriminator using the offset vector and the pseudo-labels generated by the generator, so that the discriminator can predict the membership probability of the input samples according to the offset vector; use the membership probability generated by the discriminator as the new pseudo-labels of the input samples to optimize the generator; iteratively execute the cross-supervised training of the generator and the discriminator until the discriminator converges; use the trained discriminator as the final membership inference attack model; Input the sample to be detected into a large language model to generate the output response of the sample to be detected; generate the offset vector between the sample to be detected and its output response based on the pre-trained encoder, and input the offset vector into the membership inference attack model to generate the membership probability of the sample to be detected, so as to determine whether the sample to be detected is a training member.
2. The unsupervised black-box membership inference attack method based on generative collaborative learning according to claim 1, wherein After calculating the task scores of each heuristic task, it further includes: Normalize the task scores, and the normalization formula is: where, s i represents the task score of the i-th heuristic task; s i ′ represents the normalized task score of the i-th heuristic task.
3. The unsupervised black-box membership inference attack method based on generative collaborative learning according to claim 2, wherein Aggregating the task scores through the generator to obtain the final membership score of the input samples, including: Performing weighted summation on the normalized task scores to obtain the final membership score of the input samples, and the calculation formula is: where s member represents the final member score of the input sample; w i represents the weight of the i-th heuristic task; s i ′ represents the task score of the i-th heuristic task after normalization, and N represents the total number of the heuristic tasks.
4. The unsupervised black-box membership inference attack method based on generative collaborative learning according to claim 1, characterized in that Screening the input samples based on the final membership score and a preset confidence threshold, including: Calculating the distribution of the final membership score; according to the distribution of the final membership score and the preset confidence threshold, calculating the dynamic membership threshold and non-membership threshold through the inverse cumulative distribution function, and the calculation formula is: θ member = F -1 (1 - α); θ non-member = F -1 (α); where θ member represents the member threshold; θ non-member represents the non-member threshold; F -1 (·) represents the inverse cumulative distribution function; α represents the preset confidence threshold; Screen out the input samples with the highest final membership score greater than or equal to the membership threshold and the lowest final membership score less than or equal to the non-membership threshold, and the proportion is the preset confidence threshold.
5. The unsupervised black-box membership inference attack method based on generative collaborative learning according to claim 4, wherein Generating corresponding pseudo-labels according to the final membership score of the screened input samples, including: Among them, represents the pseudo-label; s member represents the final membership score of the input sample; θ member represents the membership threshold; θ non-member represents the non-membership threshold.
6. The unsupervised black-box membership inference attack method based on generative collaborative learning according to claim 1, characterized in that Calculating the offset vector between the input samples and the output response based on a pre-trained encoder, including: Generating the embedding vectors of the input samples and the output response based on the pre-trained encoder, and calculating the offset vector between the input sample embedding vector and the output response embedding vector, and the definition of the offset vector satisfies the following formula: Δe = e(y) - e(x); where, Δe represents the offset vector; e(x) represents the embedding representation of the input sample x; e(y) represents the embedding representation of the output response y; Construct a training set for the discriminator based on the offset vector and the corresponding pseudo-label, and the definition of the training set satisfies the following formula: Among them, represents the training set of the discriminator ; Δe represents the offset vector; represents the pseudo-label.
7. The unsupervised black-box membership inference attack method based on generative collaborative learning according to claim 1, wherein Training the discriminator using the offset vector and the pseudo-label generated by the generator further includes: Constructing a binary cross-entropy loss between the membership probability generated by the discriminator and the pseudo-label generated by the generator, and training the discriminator with the goal of minimizing the binary cross-entropy loss; The formula for the binary cross-entropy loss is: Among them, represents the binary cross-entropy loss of the discriminator ; I(·) represents the indicator function; represents the pseudo-label of the i-th input sample; N represents the total number of input samples; represents the membership probability predicted by the discriminator for the i-th input sample.
8. The unsupervised black-box membership inference attack method based on generative collaborative learning according to claim 1, wherein Before inputting the input sample into the generator, it further includes: Initializing the weights of each heuristic task in the generator; The heuristic tasks at least include multiple-choice, name filling, self-checking, output consistency, and recall continuation.
9. The unsupervised black-box membership inference attack method based on generative collaborative learning according to claim 1, characterized in that Regarding the membership probability generated by the discriminator as the new pseudo-label of the input sample includes: Calculating the distribution of the membership probability; according to the distribution of the membership probability and a preset confidence threshold, calculating a dynamic threshold through the inverse cumulative distribution function to screen out high-confidence input samples; Using the screened input samples for the training of the generator again, and regarding the membership probability generated by the discriminator as the new pseudo-label.
10. An unsupervised black-box membership inference attack system based on generative collaborative learning, characterized in that, When the system is executed, it implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Data information label processing method of large language model
CN117453921A
Method and apparatus for determining and using controllable direction of GAN space
CN119631082A
KR20240085182A
Cited By
Retrieval enhancement generation system member leakage risk assessment method based on few queries
CN121478951A