Large model adversarial sample pair generation method and device, electronic equipment and storage medium
By looping into image text pairs and updating image text pairs with predicted jailbreak probability, an adversarial sample with stronger jailbreak capability is generated, solving the problem of inefficiency in the existing technology and achieving high efficiency in adversarial sample generation.
Patent Information
- Application Number
- CN202510622926.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing multimodal large model has low efficiency in adversarial sample generation methods, and the jailbreak attack method based on adversarial attacks creates direct conflict with security alignment, which is high in time.
By looping the image text pair into the target big model, obtaining the hidden state, and inputting it into the probability prediction model to obtain the predicted jailbreak probability, update the image text pair based on the predicted jailbreak probability until the end condition is reached, and an adversarial sample pair is generated. The predicted jailbreak probability is used to guide the update of the image text pair, avoiding offsetting the impact of secure alignment, reducing the number of iterations, and improving the generation efficiency.
Adversarial samples with stronger jailbreak capabilities are generated, which reduces the number of iterations, shortens the generation time, and improves the generation efficiency of adversarial sample pairs.
Smart Images

Figure CN120544002A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of multimodal large-model adversarial samples. Specifically, the present application relates to a method, device, electronic device, computer-readable storage medium, and computer program product for generating large-model adversarial samples. Background Art
[0002] With the rapid development of artificial intelligence (AI), large multimodal models have demonstrated outstanding performance across a wide range of multimodal tasks. However, these models are also vulnerable to adversarial examples. By adding specific perturbations to an original, clean image, deep neural networks can misjudge the predictions. These perturbed images are called adversarial examples.
[0003] In the existing technology, multimodal jailbreak attacks based on adversarial attacks usually perform jailbreaks by aligning model outputs with harmful text. This method directly conflicts with security alignment, has high time costs, and has low efficiency in generating adversarial samples. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium and computer program product for generating large-model adversarial sample pairs, aiming to solve the technical problem of low efficiency of existing adversarial sample generation methods.
[0005] In a first aspect, a method for generating large-model adversarial sample pairs is provided, the method comprising:
[0006] The image-text pair is looped into the target model to obtain the corresponding hidden state;
[0007] Input the hidden state into the probability prediction model to obtain the predicted jailbreak probability;
[0008] Based on the predicted jailbreak probability, the image-text pair is updated until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair; the adversarial sample pair is the updated image-text pair.
[0009] Optionally, the target large model includes multiple conversion modules, each conversion module outputs a corresponding hidden state, and there are multiple probability prediction models, and the probability prediction models correspond to the conversion modules one by one;
[0010] Input the hidden state into the probability prediction model to obtain the predicted jailbreak probability, including:
[0011] Each hidden state is input into the corresponding probability prediction model to obtain the corresponding predicted jailbreak probability.
[0012] Optionally, based on the predicted jailbreak probability, updating the image-text pair until a first predetermined end condition is reached, terminating the update, and obtaining the adversarial sample pair, including:
[0013] Based on each predicted jailbreak probability, the image-text pair is updated until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair.
[0014] Optionally, based on each predicted jailbreak probability, updating the image-text pair until a first predetermined end condition is reached, terminating the update, and obtaining an adversarial sample pair, including:
[0015] Determine the mean square error loss function based on the predicted jailbreak probability and the target probability;
[0016] Based on the mean square error loss function, the parameters of the image-text pair are updated until the first predetermined end condition is reached, and the update is terminated to obtain the adversarial sample pair.
[0017] Optionally, based on the predicted jailbreak probability, updating the image-text pair until a first predetermined end condition is reached, terminating the update, and obtaining the adversarial sample pair, including:
[0018] Determine the jailbreak loss function based on the predicted jailbreak probability;
[0019] The parameters of the image-text pair are updated based on the jailbreak loss function until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair.
[0020] Optionally, updating parameters of the image-text pair based on the jailbreak loss function until a first predetermined end condition is reached, and then terminating the update to obtain an adversarial sample pair, including:
[0021] Update the parameters of the image-text pair based on the jailbreak loss function and determine the number of iterations of the image-text pair;
[0022] If the number of iterations reaches the first predetermined end condition, the update is terminated and the adversarial sample pair is obtained.
[0023] Optionally, obtaining a training sample set; the training sample set includes multiple training samples;
[0024] Input the training samples into the target large model one by one to obtain corresponding multiple response results;
[0025] Based on multiple response results, determine the corresponding reference jailbreak probability;
[0026] Obtain the hidden state of the target large model for the training sample, input the hidden state into the initial prediction model, and obtain the predicted jailbreak probability;
[0027] Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until a second predetermined end condition is reached, and the training is terminated to obtain a trained probability prediction model.
[0028] Optionally, based on multiple response results, a corresponding reference jailbreak probability is determined, including:
[0029] Determine the content evaluation result of each response result based on the preset content evaluation component;
[0030] Based on the content evaluation results, the corresponding reference jailbreak probability is determined.
[0031] Optionally, content assessment results include harmful;
[0032] Based on the content evaluation results, the corresponding reference jailbreak probability is determined, including:
[0033] The number of response results whose content is evaluated as harmful is taken as the first number;
[0034] Based on the number of reply results and the first number, a reference jailbreak probability is determined.
[0035] Optionally, the second predetermined termination condition includes that the probability difference is less than or equal to a preset difference threshold;
[0036] Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until a second predetermined end condition is reached, and the training is terminated to obtain a trained probability prediction model, including:
[0037] The absolute value of the difference between the predicted jailbreak probability and the reference jailbreak probability is taken as the probability difference;
[0038] The parameters of the initial probability model are updated based on the probability difference until the probability difference is less than or equal to the preset difference threshold, and the training is terminated to obtain a trained probability prediction model.
[0039] In a second aspect, a large-model adversarial sample pair generation device is provided, the device comprising:
[0040] The state acquisition module is used to cyclically input the image-text pair into the target model to obtain the corresponding hidden state;
[0041] The probability prediction module is used to input each hidden state into the probability prediction model to obtain the predicted jailbreak probability;
[0042] The sample generation module is used to update the image-text pair based on the predicted jailbreak probability until a first predetermined end condition is reached, terminate the update, and obtain an adversarial sample pair; the adversarial sample pair is the updated image-text pair.
[0043] According to a third aspect, an electronic device is provided, comprising:
[0044] A memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any method in the first aspect of the present application.
[0045] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the large model adversarial sample pair generation method shown in any one of the first aspects of this application.
[0046] In a fifth aspect, a computer program product is provided, comprising a computer program, which implements the steps of any one of the methods in the first aspect of the present application when the computer program is executed by a processor.
[0047] The beneficial effects of the technical solution provided by the embodiments of the present application are:
[0048] The large-model adversarial sample pair generation method provided in the present application obtains the corresponding hidden state by cyclically inputting the image-text pair into the target large model, and inputting the hidden state into the probability prediction model to obtain the predicted jailbreak probability. Based on the predicted jailbreak probability, the image-text pair is updated until the end condition is reached to obtain the adversarial sample pair, and the hidden state is associated with the jailbreak probability. The predicted jailbreak probability is used to guide the update of the image-text pair, thereby obtaining an adversarial sample with stronger jailbreak capability. There is no need to offset the influence of security alignment, which can reduce the number of iterations for generating adversarial sample pairs, shorten the time, and improve the generation efficiency of adversarial sample pairs.
[0049] Furthermore, the reference jailbreak probability of the training sample is determined based on the output of the large model, the corresponding hidden state is obtained, and the hidden state is input into the initial prediction model to obtain the predicted jailbreak probability. Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until the predetermined end condition is reached, and a trained probability prediction model is obtained. The relationship between the hidden state and the jailbreak probability is constructed through the probability prediction model, which can quickly realize the update of image-text pairs based on the predicted jailbreak probability, effectively improving the efficiency of generating adversarial sample pairs.
[0050] In addition, the target large model can include multiple conversion modules, each conversion module corresponds to the output of a hidden state, each conversion module corresponds to a probability prediction model, the hidden state is input into the corresponding probability prediction model, and the predicted jailbreak probability corresponding to the hidden state is obtained. The mean square error loss function is determined by multiple predicted jailbreak probabilities, and the image-text pair is updated. The hidden state of each conversion module is fully considered, so that the updated adversarial sample pair has a stronger jailbreak ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0052] Figure 1 A schematic diagram of an application scenario of a method for generating large-model adversarial sample pairs provided in an embodiment of the present application;
[0053] Figure 2 A flowchart of a method for generating large-model adversarial sample pairs provided in an embodiment of the present application;
[0054] Figure 3 A schematic diagram of an image in a method for generating large-model adversarial sample pairs provided in an embodiment of the present application;
[0055] Figure 4 A schematic diagram of a process for generating a probability prediction model in a large-model adversarial sample pair generation method provided in an embodiment of the present application;
[0056] Figure 5 A flowchart illustrating an example of a method for generating large-model adversarial sample pairs provided in an embodiment of the present application;
[0057] Figure 6 A schematic diagram of the structure of a large-model adversarial sample pair generation device provided in an embodiment of the present application;
[0058] Figure 7 A schematic structural diagram of an electronic device applicable to a method for generating large-model adversarial samples provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] The following describes the embodiments of the present application in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0060] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the present technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the element can be directly connected or coupled to the other element, or it can refer to the element and the other element establishing a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The terms "or", "and / or", "including at least one of the following", etc. used in this application may be interpreted as inclusive, or mean any one or any combination. For example, "including at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C", and for another example, "A, B or C" or "A, B and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C".
[0061] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0062] In the specific implementation of this application, any data related to an object, such as data involved in the object's use of an application, is required. When the embodiments of this application are applied to specific products or technologies, the permission or consent of the object must be obtained, and the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. In other words, if any of the above-mentioned data related to an object is involved in the embodiments of this application, such data must be obtained with the authorization and consent of the object and in compliance with the relevant laws, regulations, and standards of the relevant countries and regions.
[0063] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0064] First, the technical terms involved in this application are introduced and explained:
[0065] Transformation module: It is the basic building block in the Transformer architecture of the neural network model based on the self-attention mechanism. It is a deep learning component for processing sequence data. The transformation module mainly consists of two parts: the self-attention mechanism and the feedforward neural network.
[0066] Hidden state: In large language models, hidden state usually refers to the changes in the model's internal state when processing input data. These states include but are not limited to contextual information or hidden layer outputs. The hidden state is maintained through a self-attention mechanism. Each conversion module processes the input and generates a hidden state, which is then used as the input for the next block.
[0067] Jailbreak attack: It is an attack method against machine learning models, whose purpose is to make the model make the wrong predictions expected by the attacker. This attack usually involves adversarial samples.
[0068] Adversarial samples: In adversarial machine learning, the attacker generates adversarial samples that are carefully designed to cause the model to misbehave in its input space. If the model makes incorrect predictions for these adversarial samples, it can be considered that the model has been "jailbroken" to some extent.
[0069] In the prior art, jailbreak attacks against large multimodal models mainly include two implementation methods: visual cue word injection and adversarial attack. Visual cue word injection is to jailbreak by embedding jailbreak instructions into the visual modality, such as the relatively poor security alignment of the visual modality; adversarial attack is to jailbreak by setting a specific loss function related to jailbreaking and maximizing this loss. Existing multimodal jailbreak attacks based on adversarial attacks usually jailbreak by aligning the model output with harmful text. However, this method directly conflicts with security alignment because security alignment aligns harmful input and harmless response, which is exactly the opposite of the goal of the above-mentioned jailbreak attack. Therefore, the above-mentioned jailbreak attack requires more iterations (usually thousands of rounds) and a larger perturbation range to offset the impact of security alignment. The time to generate adversarial samples is long and the efficiency is low.
[0070] The large-model adversarial sample generation method, device, electronic device, computer-readable storage medium and computer program product provided in this application are intended to solve at least one of the above technical problems in the prior art.
[0071] In response to at least one of the above-mentioned technical problems or areas that need improvement in the relevant technologies, the present application proposes a large-model adversarial sample pair generation method, device, electronic device, computer-readable storage medium and computer program product. The large-model adversarial sample pair generation method provided by the scheme obtains the corresponding hidden state by cyclically inputting the image-text pair into the target large model, and inputting the hidden state into the probability prediction model to obtain the predicted jailbreak probability. Based on the predicted jailbreak probability, the image-text pair is updated until the end condition is reached to obtain the adversarial sample pair, and the hidden state is associated with the jailbreak probability. The predicted jailbreak probability is used to guide the update of the image-text pair, thereby obtaining an adversarial sample with stronger jailbreak capability. There is no need to offset the influence of security alignment, which can reduce the number of iterations of adversarial sample generation, shorten the time, and improve the generation efficiency of adversarial sample pairs.
[0072] Furthermore, the reference jailbreak probability of the training sample is determined based on the output of the large model, the corresponding hidden state is obtained, and the hidden state is input into the initial prediction model to obtain the predicted jailbreak probability. Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until the predetermined end condition is reached, and a trained probability prediction model is obtained. The relationship between the hidden state and the jailbreak probability is constructed through the probability prediction model, which can quickly realize the update of image-text pairs based on the predicted jailbreak probability, effectively improving the efficiency of generating adversarial sample pairs.
[0073] In addition, the target large model can include multiple conversion modules, each conversion module corresponds to the output of a hidden state, each conversion module corresponds to a probability prediction model, the hidden state is input into the corresponding probability prediction model, and the predicted jailbreak probability corresponding to the hidden state is obtained. The mean square error loss function is determined by multiple predicted jailbreak probabilities, and the image-text pair is updated. The hidden state of each conversion module is fully considered, so that the updated adversarial sample pair has a stronger jailbreak ability.
[0074] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0075] Figure 1 A schematic diagram of an application scenario of the large-model adversarial sample pair generation method provided in an embodiment of the present application, wherein the application environment may include a terminal 101 or a server, and the terminal 101 is configured with a large-model adversarial sample pair generation system.
[0076] Specifically, terminal 101 cyclically inputs the image-text pair into the target large model to obtain the corresponding hidden state, inputs the hidden state into the probability prediction model to obtain the predicted jailbreak probability, and updates the image-text pair based on the predicted jailbreak probability until the first predetermined end condition is reached, ends the update, and obtains the adversarial sample pair, which is the updated image-text pair.
[0077] The above application scenario is only an example and does not limit the application scenario of the large-model adversarial sample generation method of this application.
[0078] Technicians in this technical field can understand that the terminal can be a smart phone (such as an Android phone, iOS phone, etc.), a tablet computer, a laptop computer, a digital broadcast receiver, a MID (Mobile Internet Devices), a PDA (Personal Digital Assistant), a desktop computer, a smart home appliance, a vehicle-mounted terminal (such as a vehicle-mounted navigation terminal, a vehicle-mounted computer, etc.), a smart speaker, a smart watch, etc. The terminal and the server can be directly or indirectly connected through wired or wireless communication, but are not limited to this.
[0079] The server may include a server that is equipped with a computer capable of processing database operations. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server or server cluster that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. The specific requirements can also be determined based on the actual application scenario requirements, which are not limited here.
[0080] In some possible implementations, taking the execution subject as a large model adversarial sample pair generation system as an example, the embodiment of the present application provides a large model adversarial sample pair generation method, such as Figure 2 As shown, the following steps may be included:
[0081] S210, the image-text pair is cyclically input into the target large model to obtain the corresponding hidden state.
[0082] The image-text pair includes image content and text content. There can be multiple image-text pairs. The image content of the multiple image-text pairs is the same, but the text content is different. The image content and text content in the image-text pair may be related or unrelated.
[0083] Among them, the target large model can be a multimodal large model that can process images and texts, including language-driven visual analysis models and large-scale visual language models.
[0084] Specifically, an image-text pair is obtained and input into a target large model. At this time, the target large model may be a white-box model, and hidden state information generated when the target large model processes the image-text pair is obtained. The hidden state refers to the change in the internal state of the model when processing input data. These states include but are not limited to context information or hidden layer output.
[0085] During the specific implementation process, according to the preset default pictures and test text sets, the default pictures are respectively combined with multiple text data in the test text set to obtain multiple image-text pairs to obtain a test sample set. The test text set is obtained by dividing the preset adversarial data set, wherein the adversarial data set can be a combination of the adversarial sample benchmark test set AdvBench, the strong rejection attack set StrongREJECT and the jailbreak benchmark test set JailbreakBench. The test text set can also include a mini jailbreak data set miniJailbreakV 28K for evaluating the robustness of multimodal large language models to various jailbreak attacks.
[0086] In the specific implementation process, the image content is preset, such as Figure 3 As shown, the content of the image is not limited to pandas. This is just an example and does not limit the content and style of the image.
[0087] S220: Input the hidden state into the probability prediction model to obtain the predicted jailbreak probability.
[0088] Specifically, the probability prediction model includes the correlation between the hidden state and the jailbreak probability. The hidden state is input into the probability prediction model to obtain the predicted jailbreak probability output by the model. Based on the predicted jailbreak probability, the parameters of the image-text pair corresponding to the hidden state are updated to obtain the adversarial sample pair.
[0089] S230, based on the predicted jailbreak probability, updating the image-text pair until a first predetermined end condition is reached, ending the update, and obtaining an adversarial sample pair.
[0090] Among them, the adversarial sample pairs are updated image-text pairs.
[0091] Specifically, the predicted jailbreak probability is compared with 1 (i.e., 100% probability of jailbreak), and the gap between the two is determined, so as to update the corresponding image-text pair so that the gap between the predicted jailbreak probability and 1 gradually becomes smaller, and the first predetermined end condition is reached, and the update is ended to obtain an adversarial sample pair. There can be multiple image-text pairs. After updating multiple image-text pairs, multiple adversarial sample pairs are obtained to generate adversarial samples.
[0092] In the specific implementation process, the image-text pair can be updated by updating the pixel parameters of the image or the text. Two constraints are imposed when updating the image parameters. The first is to constrain the pixel value of each pixel to be between 0 and 1 to ensure that the pixel is within the legal range. The second is the constraint on the disturbance range. It is necessary to ensure that the disturbance range is small enough so that it is difficult for the human eye to detect it, while misleading the model to ensure that the adversarial noise is highly invisible.
[0093] During implementation, the features of an image-text pair, such as each pixel in the image within the pair, are adjusted in the direction of the gradient. If the gradient is positive, the feature's value is increased; if the gradient is negative, the feature's value is decreased. The magnitude of the adjustment is typically proportional to the gradient and multiplied by a small coefficient, called the learning rate or step size, to ensure that the update perturbation is not too large. Specific image parameter adjustments can include increasing the pixelization of the image, changing the pixel's color, contrast, or brightness, or adding noise to the image. The image parameter adjustment methods are not limited to the examples given and can also be set based on actual needs.
[0094] During the specific implementation process, the first predetermined end condition may include: the accuracy rate, the loss function value reaches a threshold value, and the number of iterations reaches a threshold value, etc. Specifically, the first predetermined end condition may include the accuracy rate reaching a preset threshold value, comparing the predicted jailbreak probability with 1, determining the predicted accuracy rate, and updating the parameters of the image-text pair according to the predicted accuracy rate until the corresponding prediction accuracy rate reaches the accuracy rate threshold, ending the update, and obtaining the adversarial sample pair. For example, the accuracy rate threshold is set to 80%, then when the prediction accuracy rate reaches 0.8, it is considered that the update can be ended. It can be flexibly set according to actual needs.
[0095] In some possible implementations, the above steps of inputting the hidden state into the probability prediction model to obtain the predicted jailbreak probability include:
[0096] Each hidden state is input into the corresponding probability prediction model to obtain the corresponding predicted jailbreak probability.
[0097] Among them, the target large model includes multiple conversion modules, each conversion module corresponds to outputting a hidden state, there are multiple probability prediction models, and the probability prediction models correspond to the conversion modules one by one.
[0098] There may be a sequence relationship between the conversion modules, that is, the input of one conversion module is the output of the next conversion module.
[0099] Specifically, the trained probability prediction model is connected to the corresponding conversion module, and then the test data is input into the target multimodal large model to obtain the hidden state output by each conversion module of its text processing part. After obtaining the hidden state, these hidden states are input into the probability prediction model at the corresponding position to obtain multiple predicted jailbreak probabilities.
[0100] During the specific implementation process, the intermediate hidden states corresponding to each conversion module are obtained. Each intermediate hidden state can be input into the corresponding probability prediction model, or a target hidden state can be generated based on multiple intermediate hidden states, and the target hidden state can be input into the probability prediction model to obtain the predicted jailbreak probability. The probability prediction model used here is also obtained by training based on the hidden state composed of multiple intermediate states. The specific method of obtaining the predicted jailbreak probability can be selected according to actual needs.
[0101] In some possible implementations, the above steps include updating the image-text pair based on the predicted jailbreak probability until a first predetermined end condition is reached, and then terminating the update to obtain the adversarial sample pair, including:
[0102] Based on each predicted jailbreak probability, the image-text pair is updated until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair.
[0103] Specifically, the predicted jailbreak probability corresponding to each hidden state is obtained, so that each predicted jailbreak probability becomes higher and higher, and the predicted jailbreak probability becomes closer and closer to 1, and the image-text pair is updated until the first predetermined end condition is reached, and the update is terminated to obtain the adversarial sample pair.
[0104] In the specific implementation process, in order to generate adversarial samples with jailbreak capabilities, the input image is updated in the direction of maximizing the predicted jailbreak probability, that is, minimizing the difference between the jailbreak probability predicted by each probability prediction model and 1 (1 represents a 100% probability of jailbreak) to obtain the gradient. After obtaining the gradient, the gradient is used to iteratively update the input image to gradually increase its jailbreak probability, wherein the loss function can be a mean square error loss function.
[0105] In some possible implementations, the above steps update the image-text pair based on each predicted jailbreak probability until a first predetermined end condition is reached, and then terminate the update to obtain an adversarial sample pair, including:
[0106] Determine the mean square error loss function based on the predicted jailbreak probability and the target probability;
[0107] Based on the mean square error loss function, the parameters of the image-text pair are updated until the first predetermined end condition is reached, and the update is terminated to obtain the adversarial sample pair.
[0108] Among them, the target probability can be set based on actual needs.
[0109] Specifically, based on the predicted jailbreak probability corresponding to each hidden state and the target probability, a mean square error loss function is determined. Based on the mean square error loss function, the parameters of the image-text pair are updated until the first predetermined end condition is reached, the update is terminated, and an adversarial sample pair is obtained.
[0110] In the specific implementation process, the jailbreak probability optimization JPO is performed using the probability prediction model. The jailbreak ability of the input is enhanced by maximizing the predicted jailbreak probability of the input hidden state. The jailbreak probability of the input is maximized by minimizing the loss function. The loss function can be:
[0111]
[0112] Among them, L JPO is the loss function, is the predicted jailbreak probability, 1 represents 100% jailbreak probability, L MSE is the formula for calculating the mean square error loss function.
[0113] Specifically, the gradient of the above loss function for the image-text pair is calculated. The gradient direction is the direction that can minimize the loss function. The gradient is used to update the image in the image-text pair:
[0114]
[0115] Where x′ img As image parameters, x′ can be img Initialized to x imh , α is the single step length, sign is the sign function, x′ img L JPO is the gradient of the loss function with respect to the input. After multiple rounds of iterations, x′ img The hidden state of the input will get closer and closer to the hidden state with a high jailbreak probability, so its jailbreak ability is enhanced. Compared with existing methods, the jailbreak probability optimization method improves the jailbreak ability of the input by maximizing the jailbreak probability of the hidden state, avoiding direct conflict with security alignment, so it does not require too many iterations and a large perturbation range. In addition, the jailbreak probability optimization method only uses a probabilistic prediction model to perform a single-digit regression task, which is faster and easier to converge than existing methods.
[0116] In the specific implementation process, the formula of the mean square error loss function is as follows:
[0117]
[0118] Among them, y is the true value, corresponding to the target probability, is the predicted value, corresponding to the predicted jailbreak probability, n is the number of samples, corresponding to the number of hidden states, and i is the index of the sample.
[0119] In some possible implementations, the above steps include updating the image-text pair based on the predicted jailbreak probability until a first predetermined end condition is reached, and then terminating the update to obtain the adversarial sample pair, including:
[0120] Determine the jailbreak loss function based on the predicted jailbreak probability;
[0121] The parameters of the image-text pair are updated based on the jailbreak loss function until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair.
[0122] Specifically, based on the predicted jailbreak probability, a jailbreak loss function is determined, and parameters of the image-text pair are updated based on the jailbreak loss function until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair.
[0123] In the specific implementation process, the first predetermined end condition may include the predicted probability reaching a preset threshold, gradient updating the predicted jailbreak probability to update the parameters of the image-text pair until the predicted probability reaches the preset threshold, ending the training, and obtaining a probability prediction model; the first predetermined end condition may also include the convergence of the jailbreak loss function, determining the jailbreak loss function based on the predicted jailbreak probability, updating the parameters of the image-text pair based on the jailbreak loss function until the jailbreak loss function converges, ending the training, and obtaining a trained probability prediction model; the first predetermined end condition may also include the probability difference being less than or equal to a preset difference threshold, taking the absolute value of the difference between the predicted jailbreak probability and the reference jailbreak probability as the probability difference, updating the parameters of the initial probability model based on the probability difference until the probability difference is less than or equal to the preset difference threshold, ending the training, and obtaining a trained probability prediction model.
[0124] In some possible implementations, in the above steps, updating the parameters of the image-text pair based on the jailbreak loss function until a first predetermined end condition is reached, and then ending the update to obtain an adversarial sample pair includes:
[0125] Update the parameters of the image-text pair based on the jailbreak loss function and determine the number of iterations of the image-text pair;
[0126] If the number of iterations reaches the first predetermined end condition, the update is terminated and the adversarial sample pair is obtained.
[0127] Specifically, the parameters of the image-text pair are updated based on the jailbreak loss function. Since the response result is uncertain, the jailbreak loss function may not converge, so the number of iterations of the image-text pair can be determined. If the number of iterations reaches the first predetermined end condition, the update is terminated to obtain the adversarial sample pair. In actual use, the number of iterations can be set to 150 or 200 rounds, which is much lower than the existing technology.
[0128] During the specific implementation process, the number of iterations required to generate adversarial samples can be pre-set, the number of iterations can be recorded when the image-text pair is updated, and the parameters of the image-text pair are updated based on the jailbreak loss function. When the number of iterations reaches the preset number of iterations, the update is stopped to obtain the adversarial sample pair. The existing technology may require 5,000 rounds to generate a universal adversarial sample, while this scheme only needs 200 rounds to generate a universal adversarial sample, reducing the time cost. The existing method does not limit the perturbation range, while the perturbation range of this method is very small and has higher invisibility, which is due to avoiding direct conflict with security alignment. Therefore, this scheme is superior to the existing technology in terms of attack success rate, number of attack iterations and perturbation range.
[0129] In some possible implementations, such as Figure 4 As shown, the above method also includes:
[0130] S410: Obtain a training sample set.
[0131] The training sample set includes multiple training samples.
[0132] Specifically, a training text set can be obtained from a preset adversarial data set, and each text in the training text set can be combined with an image to obtain multiple training samples to obtain a training sample set, wherein the training sample can be a training image-text pair consisting of a default image and a text.
[0133] In the specific implementation process, the training sample set can be obtained by combining a training text set and a default image. The training text set can be obtained by dividing the adversarial dataset. The adversarial dataset can include the adversarial sample evaluation dataset AdvBench, the attack rejection dataset StrongREJECT and the jailbreak adversarial dataset JailbreakBench. The adversarial dataset can be divided into a training text set and a test text set according to the proportion. For example, the adversarial dataset can be divided into a training text set and a test text set according to the ratio of 0.8 and 0.2. The training text set is used to train the probability prediction model, and the test text set can be used to generate adversarial samples.
[0134] S420: Input the training samples into the target large model one by one to obtain corresponding multiple response results.
[0135] One training sample corresponds to multiple response results, and the number of response results can be preset.
[0136] Specifically, the training samples are input into the target large model one by one to obtain multiple response results corresponding to each training sample. The images and texts in the training samples are processed and analyzed respectively to obtain corresponding response results. For one training sample, there can be multiple different reasonable responses, and the number of responses can be set based on actual needs. For example, considering the accuracy and time cost, 20 responses can be generated for each training sample.
[0137] During the specific implementation process, the training samples may include images and texts. The training samples are input into the target large model, the images in the training samples are input into the image processing module in the target large model, and the texts in the training samples are input into the text processing module in the target large model. After extracting the features of the corresponding images and the corresponding texts respectively, the image features are mapped to the text space using a connector and spliced with the text features. The text units are processed one by one using an autoregressive method to generate a response result.
[0138] S430: Determine a corresponding reference jailbreak probability based on the multiple reply results.
[0139] Specifically, for each reply result, the preset judgment rules are used to evaluate the harmfulness of the reply content, determine the evaluation results corresponding to each reply result, and determine the corresponding reference jailbreak probability based on the various reply results of the same training sample, where the harmfulness can be used to indicate whether the corresponding reply result can cause the large model to jailbreak.
[0140] During the specific implementation process, the preset judgment rules may include comparing the reply results with the harmful data set, and determining the harmfulness assessment results based on the similarity obtained from the comparison. It may also include inputting the reply results into a harmful text assessment model to obtain the assessment results, or obtaining the assessment results based on manual selection. All specific methods that can determine whether the reply results can cause the large model to jailbreak can be used, and the specific selection is based on actual needs without limitation.
[0141] S440, obtaining the hidden state of the target large model for the training sample, inputting the hidden state into the initial prediction model, and obtaining the predicted jailbreak probability.
[0142] The initial prediction model may be a multi-layer perceptron (MLP) model, which includes at least three layers: an input layer, one or more hidden layers, and an output layer.
[0143] Specifically, the hidden state of the target large model when processing the training sample is obtained, the hidden state is used as the input of the initial prediction model, the reference jailbreak probability is used as the label, and the predicted jailbreak probability output by the initial prediction model is obtained. Based on the reference jailbreak probability, the initial prediction model is trained and the model parameters are modified to make the predicted jailbreak probability closer and closer to the reference jailbreak probability. The relationship between the hidden state and the jailbreak probability is constructed by the model, which helps to use the jailbreak probability to generate adversarial samples.
[0144] During the specific implementation process, the target large model may include multiple conversion modules, each conversion module can output the corresponding hidden state. Obtaining the hidden state of the target large model is to obtain the hidden state of each module. Each conversion module is provided with a corresponding initial prediction model. The hidden state of each conversion module is input into the corresponding initial prediction model to obtain the respective predicted jailbreak probability, and the initial model parameters are updated based on the predicted jailbreak probability and the reference jailbreak probability.
[0145] S450 , based on the reference jailbreak probability and the predicted jailbreak probability, update the parameters of the initial probability model until a second predetermined end condition is reached, and the training is terminated to obtain a trained probability prediction model.
[0146] Among them, the second predetermined end condition may include: the accuracy, recall rate or loss function value reaches a threshold, the training time reaches a maximum value, the gradient change of the model is very small, and the performance begins to decline, etc.
[0147] Specifically, the reference jailbreak probability can be compared with the predicted jailbreak probability to update the parameters of the initial probability model so that the model meets the second predetermined end condition, the training is terminated, and a trained probability prediction model is obtained.
[0148] In a specific implementation, the second predetermined termination condition includes the convergence of the probability loss function, determining the probability loss function based on the reference jailbreak probability and the predicted jailbreak probability, and updating the parameters of the initial probability model based on the probability loss function until the probability loss function converges, terminating the training, and obtaining a trained probability prediction model. The second predetermined termination condition may also include the probability difference being less than or equal to a preset difference threshold, using the absolute value of the difference between the predicted jailbreak probability and the reference jailbreak probability as the probability difference, and updating the parameters of the initial probability model based on the probability difference until the probability difference is less than or equal to the preset difference threshold, terminating the training, and obtaining a trained probability prediction model.
[0149] During the specific implementation process, the second predetermined end condition may also include the accuracy reaching a preset threshold, comparing the reference jailbreak probability with the predicted jailbreak probability, determining the accuracy of the predicted jailbreak probability, and updating the parameters of the large model to be trained according to the accuracy until the accuracy reaches the preset threshold, ending the training, and obtaining a probability prediction model.
[0150] In the specific implementation process, the probability prediction model can also be in the form of a network. In order to improve the jailbreak ability of the input, a JPPN (Jailbreak Probability Prediction Network) can be constructed and trained. The JPPN is used to model the relationship between the hidden state of the input and the jailbreak probability. Specifically, a three-layer MLP model can be used to construct the JPPN. The input of the JPPN is the hidden state of the large model. The output of the JPPN is the prediction of the jailbreak probability of the hidden state of the input, which is a real number greater than or equal to 0 and less than or equal to 1. Since each conversion module in the target multimodal large model will output an intermediate state, the present invention will initialize a separate JPPN for each conversion module and connect it to the output position of the conversion module. The training process of the JPPN can use the mean square error loss function, and the optimization goal is to minimize the difference between the approximate jailbreak probability of an image-text pair and the jailbreak probability of the hidden state of this image-text pair predicted by the JPPN. In actual application, the number of training rounds of the probability prediction model can be set to 150 rounds. When the number of training rounds is reached, the parameter update is stopped to obtain the probability prediction network. The initial learning rate can also be set to 0.001, and the learning rate is reduced to 20% of the original every 50 rounds.
[0151] In some possible implementations, determining the corresponding reference jailbreak probability based on the multiple response results in the above steps includes:
[0152] Determine the content evaluation result of each response result based on the preset content evaluation component;
[0153] Based on the content evaluation results, the corresponding reference jailbreak probability is determined.
[0154] Among them, the content evaluation results include harmful and harmless.
[0155] Specifically, based on a preset content evaluation component, the content of the reply result is evaluated to determine whether the reply result is harmful content, and a content evaluation result of each reply result is obtained. According to whether the content evaluation result is harmful or harmless, an approximate jailbreak probability is calculated, and the approximate jailbreak probability is used as a reference jailbreak probability.
[0156] During the specific implementation process, the content evaluation component can score each reply, scoring replies including harmful content as 0, rejection or irrelevant content as 1, and positive guidance content as 2. 0 points can be regarded as a successful attack, and 1 and 2 points can be regarded as failed attacks. By obtaining the content evaluation results of multiple replies for an input, the approximate jailbreak probability of the input can be calculated, that is, the above-mentioned reference jailbreak probability. The approximate jailbreak probability is defined as the ratio of the number of successful jailbreaks to the number of replies in the multiple reply results corresponding to this input.
[0157] In some possible implementations, determining the corresponding reference jailbreak probability based on the content evaluation result in the above steps includes:
[0158] The number of response results whose content is evaluated as harmful is taken as the first number;
[0159] Based on the number of reply results and the first number, a reference jailbreak probability is determined.
[0160] Among them, the content evaluation results include harmful and harmless.
[0161] Specifically, for a training sample, the total number of corresponding reply results is determined, a first number of reply results whose content evaluation results are harmful is determined, and based on the relationship between the first number and the total number, a corresponding reference jailbreak probability is determined.
[0162] In the specific implementation process, the same input may successfully jailbreak or fail to jailbreak. The jailbreak probability can be defined by dividing the number of successful jailbreaks by the total number of queries. Specifically, the jailbreak probability of X on M is is defined as:
[0163]
[0164] Where N is the number of response results, M represents the multimodal target large model, and X represents the training sample of the input target large model, specifically X=(x img ,x txt ), x img is the image in the training sample, x txt is the text in the training sample, C is the standard for judging whether the model output is harmful, which can be achieved through the preset harmful judgment interface.
[0165] Furthermore, a limited number of replies can be generated, and the proportion of replies that successfully jailbreak can be calculated to obtain the jailbreak probability. The approximate value of jailbreak probability is used as the reference jailbreak probability, which can be defined as the approximate jailbreak probability
[0166]
[0167] in, is the approximate jailbreak probability, n is the number of finite response results, M represents the multimodal target large model, X = (x img ,x txt ) represents the training sample of the input target model, x img is the image in the training sample, x txtis the text in the training sample, C is the criterion for judging whether the model output is harmful. According to the law of large numbers, when the number of replies used to calculate the approximate probability of jailbreaking increases, will gradually converge to
[0168] In some possible implementations, in the above steps, updating the parameters of the initial probability model based on the reference jailbreak probability and the predicted jailbreak probability until a second predetermined end condition is reached, thereby terminating the training and obtaining a trained probability prediction model, including:
[0169] The absolute value of the difference between the predicted jailbreak probability and the reference jailbreak probability is taken as the probability difference;
[0170] The parameters of the initial probability model are updated based on the probability difference until the probability difference is less than or equal to the preset difference threshold, and the training is terminated to obtain a trained probability prediction model.
[0171] The second predetermined termination condition includes that the probability difference is less than or equal to a preset difference threshold.
[0172] Specifically, the absolute value of the difference between the predicted jailbreak probability and the reference jailbreak probability is used as the probability difference, the probability difference is reduced, and the parameters of the initial probability model are updated according to the probability difference until the probability difference is less than or equal to the preset difference threshold. The training is terminated to obtain a probability prediction model that can determine the possibility of jailbreaking a hidden state.
[0173] During the specific implementation process, in order to ensure that the probability prediction model can make relatively accurate predictions on the jailbreak probability of different inputs after training, it is necessary to test the accuracy of the probability prediction model. Since predicting the jailbreak probability is a regression problem, and the true value of the regression problem has a certain degree of randomness (this is because the value of the reference jailbreak probability depends on the response content generated by the multimodal target large model, and the response content is random), this application proposes a τ threshold accuracy indicator: when the absolute value of the difference between the predicted jailbreak probability and the true value is less than or equal to τ, the prediction is considered accurate, otherwise it is considered inaccurate. The calculation method of the τ threshold accuracy is the number of accurate predictions divided by the total number of predictions. For example, τ can be 0.2 or 0.3, which can be calculated according to actual conditions.
[0174] In the above embodiment, the image-text pair is cyclically input into the target large model to obtain the corresponding hidden state, the hidden state is input into the probability prediction model to obtain the predicted jailbreak probability, and based on the predicted jailbreak probability, the image-text pair is updated until the end condition is reached to obtain the adversarial sample pair, the hidden state is associated with the jailbreak probability, and the predicted jailbreak probability is used to guide the update of the image-text pair, thereby obtaining an adversarial sample with stronger jailbreak capability. There is no need to offset the impact of security alignment, which can reduce the number of iterations of adversarial sample generation, shorten the time, and improve the generation efficiency of adversarial sample pairs.
[0175] Furthermore, the reference jailbreak probability of the training sample is determined based on the output of the large model, the corresponding hidden state is obtained, and the hidden state is input into the initial prediction model to obtain the predicted jailbreak probability. Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until the predetermined end condition is reached, and a trained probability prediction model is obtained. The relationship between the hidden state and the jailbreak probability is constructed through the probability prediction model, which can quickly realize the update of image-text pairs based on the predicted jailbreak probability, effectively improving the efficiency of generating adversarial sample pairs.
[0176] In addition, the target large model can include multiple conversion modules, each conversion module corresponds to the output of a hidden state, each conversion module corresponds to a probability prediction model, the hidden state is input into the corresponding probability prediction model, and the predicted jailbreak probability corresponding to the hidden state is obtained. The mean square error loss function is determined by multiple predicted jailbreak probabilities, and the image-text pair is updated. The hidden state of each conversion module is fully considered, so that the updated adversarial sample pair has a stronger jailbreak ability.
[0177] In one example, the large model adversarial sample generation method of the present application is as follows: Figure 5 As shown, this may include:
[0178] Input the image-text pair into the target large model to obtain the corresponding hidden state; wherein the target large model includes multiple conversion modules, each conversion module outputs a corresponding hidden state, and there are multiple probability prediction models, and the probability prediction models correspond to the conversion modules one by one;
[0179] Input each hidden state into the corresponding probability prediction model to obtain the corresponding predicted jailbreak probability;
[0180] Determine the mean square error loss function based on the predicted jailbreak probability and the target probability;
[0181] Based on the mean square error loss function, updating parameters of the image-text pair to determine whether a first predetermined end condition is met;
[0182] If the first predetermined end condition is met, the update is terminated and an adversarial sample pair is obtained;
[0183] If the first predetermined end condition is not met, the image-text pair is updated continuously until the first predetermined end condition is met, and the update is terminated to obtain an adversarial sample pair.
[0184] The above-mentioned large-model adversarial sample pair generation method obtains the corresponding hidden state by cyclically inputting the image-text pair into the target large model, and inputs the hidden state into the probability prediction model to obtain the predicted jailbreak probability. Based on the predicted jailbreak probability, the image-text pair is updated until the end condition is reached to obtain the adversarial sample pair, and the hidden state is associated with the jailbreak probability. The predicted jailbreak probability is used to guide the update of the image-text pair, thereby obtaining an adversarial sample with stronger jailbreak capability. There is no need to offset the influence of security alignment, which can reduce the number of iterations of adversarial sample generation, shorten the time, and improve the generation efficiency of adversarial sample pairs.
[0185] Furthermore, the reference jailbreak probability of the training sample is determined based on the output of the large model, the corresponding hidden state is obtained, and the hidden state is input into the initial prediction model to obtain the predicted jailbreak probability. Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until the predetermined end condition is reached, and a trained probability prediction model is obtained. The relationship between the hidden state and the jailbreak probability is constructed through the probability prediction model, which can quickly realize the update of image-text pairs based on the predicted jailbreak probability, effectively improving the efficiency of generating adversarial sample pairs.
[0186] In addition, the target large model can include multiple conversion modules, each conversion module corresponds to the output of a hidden state, each conversion module corresponds to a probability prediction model, the hidden state is input into the corresponding probability prediction model, and the predicted jailbreak probability corresponding to the hidden state is obtained. The mean square error loss function is determined by multiple predicted jailbreak probabilities, and the image-text pair is updated. The hidden state of each conversion module is fully considered, so that the updated adversarial sample pair has a stronger jailbreak ability.
[0187] The embodiment of the present application provides a large model adversarial sample generation device, such as Figure 6 As shown, the large model adversarial sample pair generation device 60 may include: a state acquisition module 610, a probability prediction module 620 and a sample generation module 630, wherein,
[0188] The state acquisition module 610 is used to cyclically input the image-text pair into the target large model to obtain the corresponding hidden state;
[0189] Probability prediction module 620, used to input each hidden state into the probability prediction model to obtain the predicted jailbreak probability;
[0190] The sample generation module 630 is used to update the image-text pair based on the predicted jailbreak probability until a first predetermined end condition is reached, terminate the update, and obtain an adversarial sample pair; the adversarial sample pair is the updated image-text pair.
[0191] Specifically, the sample generation module 630 may include a state acquisition module 610 and a probability prediction module 620 . The state acquisition module 610 and the probability prediction module 620 may be combined to directly update the image-text pair.
[0192] As an optional embodiment, in the device, the probability prediction module 620 is specifically configured to:
[0193] Each hidden state is input into the corresponding probability prediction model to obtain the corresponding predicted jailbreak probability.
[0194] As an optional embodiment, in the device, the sample generation module 630 is specifically configured to:
[0195] Based on each predicted jailbreak probability, the image-text pair is updated until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair.
[0196] As an optional embodiment, in the device, the sample generation module 630 is specifically configured to:
[0197] Determine the mean square error loss function based on the predicted jailbreak probability and the target probability;
[0198] Based on the mean square error loss function, the parameters of the image-text pair are updated until the first predetermined end condition is reached, and the update is terminated to obtain the adversarial sample pair.
[0199] As an optional embodiment, in the device, the sample generation module 630 is specifically configured to:
[0200] Determine the jailbreak loss function based on the predicted jailbreak probability;
[0201] The parameters of the image-text pair are updated based on the jailbreak loss function until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair.
[0202] As an optional embodiment, in the device, the sample generation module 630 is specifically configured to:
[0203] Update the parameters of the image-text pair based on the jailbreak loss function and determine the number of iterations of the image-text pair;
[0204] If the number of iterations reaches the first predetermined end condition, the update is terminated and the adversarial sample pair is obtained.
[0205] As an optional embodiment, the device further includes a probability model training module, which is specifically used to:
[0206] Obtain a training sample set; the training sample set includes multiple training samples;
[0207] Input the training samples into the target large model one by one to obtain corresponding multiple response results;
[0208] Based on multiple response results, determine the corresponding reference jailbreak probability;
[0209] Obtain the hidden state of the target large model for the training sample, input the hidden state into the initial prediction model, and obtain the predicted jailbreak probability;
[0210] Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until a second predetermined end condition is reached, and the training is terminated to obtain a trained probability prediction model.
[0211] As an optional embodiment, in the device, the probability model training module is specifically used to:
[0212] Determine the content evaluation result of each response result based on the preset content evaluation component;
[0213] Based on the content evaluation results, the corresponding reference jailbreak probability is determined.
[0214] As an optional embodiment, in the device, the probability model training module is specifically used to:
[0215] The number of response results whose content is evaluated as harmful is taken as the first number;
[0216] Based on the number of reply results and the first number, a reference jailbreak probability is determined.
[0217] As an optional embodiment, in the device, the probability model training module is specifically used to:
[0218] The absolute value of the difference between the predicted jailbreak probability and the reference jailbreak probability is taken as the probability difference;
[0219] The parameters of the initial probability model are updated based on the probability difference until the probability difference is less than or equal to the preset difference threshold, and the training is terminated to obtain a trained probability prediction model.
[0220] The large-model adversarial sample pair generation device provided in the present application obtains the corresponding hidden state by cyclically inputting the image-text pair into the target large model, and inputting the hidden state into the probability prediction model to obtain the predicted jailbreak probability. Based on the predicted jailbreak probability, the image-text pair is updated until the end condition is reached to obtain the adversarial sample pair, and the hidden state is associated with the jailbreak probability. The predicted jailbreak probability is used to guide the update of the image-text pair, thereby obtaining an adversarial sample with stronger jailbreak capability. There is no need to offset the influence of security alignment, which can reduce the number of iterations of adversarial sample generation, shorten the time, and improve the generation efficiency of adversarial sample pairs.
[0221] Furthermore, the reference jailbreak probability of the training sample is determined based on the output of the large model, the corresponding hidden state is obtained, and the hidden state is input into the initial prediction model to obtain the predicted jailbreak probability. Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until the predetermined end condition is reached, and a trained probability prediction model is obtained. The relationship between the hidden state and the jailbreak probability is constructed through the probability prediction model, which can quickly realize the update of image-text pairs based on the predicted jailbreak probability, effectively improving the efficiency of generating adversarial sample pairs.
[0222] In addition, the target large model can include multiple conversion modules, each conversion module corresponds to the output of a hidden state, each conversion module corresponds to a probability prediction model, the hidden state is input into the corresponding probability prediction model, and the predicted jailbreak probability corresponding to the hidden state is obtained. The mean square error loss function is determined by multiple predicted jailbreak probabilities, and the image-text pair is updated. The hidden state of each conversion module is fully considered, so that the updated adversarial sample pair has a stronger jailbreak ability.
[0223] The devices of the embodiments of the present application can execute the methods provided in the embodiments of the present application, and their implementation principles are similar and have corresponding technical effects. The actions performed by each module in the devices of the embodiments of the present application correspond to the steps in the methods of the embodiments of the present application. For detailed functional descriptions of each module of the device, please refer to the descriptions of the corresponding methods shown above, and will not be repeated here.
[0224] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, and the processor executes the above-mentioned computer program to implement the steps of the method provided in any optional embodiment of the present application. Compared with the prior art, it can be achieved: by cyclically inputting the image-text pair into the target large model to obtain the corresponding hidden state, inputting the hidden state into the probability prediction model to obtain the predicted jailbreak probability, updating the image-text pair based on the predicted jailbreak probability until the end condition is reached, obtaining the adversarial sample pair, associating the hidden state with the jailbreak probability, and using the predicted jailbreak probability to guide the update of the image-text pair, thereby obtaining an adversarial sample with stronger jailbreak capability, without the need to offset the influence of security alignment, and can reduce the number of iterations of adversarial sample generation, shorten the time, and improve the generation efficiency of adversarial sample pairs.
[0225] In an alternative embodiment, an electronic device is provided, such as Figure 7 As shown, Figure 7 The electronic device 7000 shown includes: a processor 7001 and a memory 7003. The processor 7001 and the memory 7003 are connected, for example, via a bus 7002. Optionally, the electronic device 7000 may further include a transceiver 7004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 7004 is not limited to one, and the structure of the electronic device 7000 does not constitute a limitation on the embodiments of the present application.
[0226] Processor 7001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 7001 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0227] Bus 7002 may include a path for transmitting information between the above components. Bus 7002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 7002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0228] The memory 7003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation here.
[0229] The memory 7003 is used to store the computer program for executing the embodiments of the present application, and the execution is controlled by the processor 7001. The processor 7001 is used to execute the computer program stored in the memory 7003 to implement the steps shown in the above method embodiments.
[0230] Among them, electronic devices include but are not limited to: servers, terminals or components that can implement the above-mentioned adversarial sample generation method.
[0231] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0232] It should be noted that the computer-readable storage medium mentioned above in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0233] An embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0234] The terms "first," "second," "third," "fourth," "1," "2," and the like (if any) in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than that shown or described in the drawings.
[0235] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0236] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart can include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times respectively. Under different scenarios at the execution time, the execution order of these sub-steps or stages can be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0237] The above description is only an optional implementation method for some implementation scenarios of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of this application, the use of other similar implementation methods based on the technical ideas of this application also falls within the protection scope of the embodiments of this application.
Claims
1. A method for generating large-model adversarial sample pairs, characterized by: include: The image-text pair is looped into the target model to obtain the corresponding hidden state; Inputting the hidden state into a probability prediction model to obtain a predicted jailbreak probability; Based on the predicted jailbreak probability, the image-text pair is updated until a first predetermined end condition is reached, and the update is terminated to obtain an adversarial sample pair; the adversarial sample pair is the updated image-text pair.
2. The method for generating large model adversarial sample pairs according to claim 1, characterized in that: The target large model includes multiple conversion modules, each of which outputs a hidden state. There are multiple probability prediction models, and the probability prediction models correspond to the conversion modules one by one. Inputting the hidden state into a probability prediction model to obtain a predicted jailbreak probability includes: Each of the hidden states is input into a corresponding probability prediction model to obtain a corresponding predicted jailbreak probability.
3. The method for generating large model adversarial sample pairs according to claim 2, characterized in that: The updating of the image-text pair based on the predicted jailbreak probability until a first predetermined end condition is reached, and the updating is terminated to obtain an adversarial sample pair, including: Based on each of the predicted jailbreak probabilities, the image-text pair is updated until the first predetermined end condition is reached, and the updating is terminated to obtain the adversarial sample pair.
4. The method for generating large model adversarial sample pairs according to claim 3, characterized in that: The updating of the image-text pair based on each of the predicted jailbreak probabilities until a first predetermined end condition is reached, and the updating is terminated to obtain the adversarial sample pair, comprising: Determining a mean square error loss function based on the predicted jailbreak probabilities and the target probabilities; Based on the mean square error loss function, parameters of the image-text pair are updated until the first predetermined end condition is reached, and the update is terminated to obtain the adversarial sample pair.
5. The method for generating large model adversarial sample pairs according to claim 1, characterized in that: The updating of the image-text pair based on the predicted jailbreak probability until a first predetermined end condition is reached, and the updating is terminated to obtain an adversarial sample pair, including: Determining a jailbreak loss function based on the predicted jailbreak probability; The parameters of the image-text pair are updated based on the jailbreak loss function until the first predetermined end condition is reached, and the update is terminated to obtain the adversarial sample pair.
6. The method for generating large model adversarial sample pairs according to claim 5, characterized in that: The step of updating the parameters of the image-text pair based on the jailbreak loss function until the first predetermined end condition is reached, and then ending the updating to obtain the adversarial sample pair includes: performing parameter updates on the image-text pair based on the jailbreak loss function, and determining the number of iterations of the image-text pair; If the number of iterations reaches the first predetermined end condition, the update is terminated to obtain an adversarial sample pair.
7. The method for generating large model adversarial sample pairs according to claim 1, characterized in that: The method further comprises: Acquire a training sample set; the training sample set includes multiple training samples; Input the training samples into the target large model one by one to obtain corresponding multiple response results; Determining a corresponding reference jailbreak probability based on the multiple reply results; Obtaining a hidden state of the target large model for the training sample, inputting the hidden state into an initial prediction model to obtain a predicted jailbreak probability; Based on the reference jailbreak probability and the predicted jailbreak probability, the parameters of the initial probability model are updated until a second predetermined end condition is reached, and the training is terminated to obtain the trained probability prediction model.
8. The method for generating large model adversarial sample pairs according to claim 7, characterized in that: The determining, based on the multiple reply results, a corresponding reference jailbreak probability includes: Determining a content evaluation result of each of the response results based on a preset content evaluation component; Based on the content evaluation result, a corresponding reference jailbreak probability is determined.
9. The method for generating large model adversarial sample pairs according to claim 8, characterized in that: The content assessment results include harmful; Determining a corresponding reference jailbreak probability based on the content evaluation result includes: The number of response results for which the content evaluation result is harmful is used as a first number; The reference jailbreak probability is determined based on the number of the reply results and the first number.
10. A large-model adversarial sample pair generation device, characterized in that: include: The state acquisition module is used to cyclically input the image-text pair into the target model to obtain the corresponding hidden state; A probability prediction module, configured to input each of the hidden states into a probability prediction model to obtain a predicted jailbreak probability; A sample generation module is used to update the image-text pair based on the predicted jailbreak probability until a first predetermined end condition is reached, end the update, and obtain an adversarial sample pair; the adversarial sample pair is an updated image-text pair.
Citation Information
Patent Citations
Large model jailbreak attack detection method
CN119377802A
Defense method and device, computer equipment and computer readable storage medium
CN119848833A
Video frame-based video multi-mode large model jailbreak attack method, system and equipment and medium
CN119862573A
Adversarially robust visual fingerprinting and image provenance models
US20230222762A1
Composite adversarial attack model training for neural networks
US20240412074A1
Cited By
Multi-modal large model availability evaluation method based on repeated output
CN121030417A