Multi-modal large model confrontation sample generation method and related equipment

By iteratively training the multimodal large model to generate adversarial samples in different data forms, and introducing consistency loss constraints, it solves the problem of high cost of adversarial samples generation and lack of semantic consistency in the prior art, and achieves efficient and semantic consistent multimodal adversarial sample generation.

CN120105086APending Publication Date: 2025-06-06BEIJING HONGTENG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411803719.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When generating adversarial samples of multimodal large models, the prior art needs to construct attack algorithms in different data forms separately. The cost is high and the generated adversarial samples do not have semantic consistency, resulting in poor joint attack effect in multimodal large models.

Method used

By inputting multimodal data sets into the target multimodal model multiple times for iterative training, adversarial samples in different data forms are generated, and semantic consistency between adversarial samples is ensured through consistency loss constraints.

Benefits of technology

It realizes the generation of multimodal adversarial samples with semantic consistency in a one-time adversarial attack, which improves the efficiency and effectiveness of adversarial sample generation and reduces the deployment complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105086A_ABST
    Figure CN120105086A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an adversarial sample generation method of a multi-modal large model and related equipment. The adversarial sample generation method for the multi-modal large model comprises the steps that a multi-modal data set is input into a target multi-modal large model for multiple times for iteration training, iteration loss and an adversarial sample set are obtained, and the multi-modal data set comprises multi-modal data in different data forms; the data form of each adversarial sample in the adversarial sample group corresponds to the multi-modal data; determining the consistency among the adversarial samples in the adversarial sample group to obtain a consistency loss; and according to the iteration loss and the consistency loss, performing parameter updating on the target multi-modal large model until a predetermined ending condition is reached, and ending iteration training to obtain a target confrontation sample group. According to the technical scheme provided by the embodiment of the invention, the multi-modal adversarial samples in different data forms can be generated by performing one-time attack on the target multi-modal large model, and the efficiency of generating the multi-modal adversarial samples is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer and communication technology, and more specifically, to a method for generating adversarial samples of a multimodal large model and related equipment. Background Art

[0002] With the rapid development of large model technology, multimodal large models have also made significant progress, but the vulnerability risks of multimodal large models have also received greater attention. This is because multimodal large models are also developed based on deep learning, and the inherent vulnerability of deep learning technology itself is inevitably inherited by large models. At the same time, the risks faced by multimodal large models when facing diverse data inputs are more variable and uncontrollable. Therefore, it is particularly important to improve the robustness of multimodal large models.

[0003] Therefore, there are many methods that directly migrate the previous technical solutions such as approximate word replacement for text, TextFooler and PGD for images to the corresponding data level of multimodal large models for adversarial attacks and adversarial sample generation. The above method is also applicable to multimodal large models based on the fragility of deep learning technology itself, and through continuous iterative optimization during the training process, adversarial samples with good attack effects are obtained.

[0004] Although the above-mentioned technical solution of migrating the original algorithm technology for a single data form to a multimodal large model has been proven to be effective, it has two major problems. First, it is too costly to construct different attack algorithms for the multiple data forms acceptable to the multimodal large model, which is equivalent to conducting multiple adversarial attack tests with different data forms on the same target model. Secondly, splitting the adversarial attack on the multimodal large model according to the data form artificially breaks the consistency within the multimodal large model. Therefore, the adversarial samples of different data forms are also unrelated to each other. Although they all have a certain degree of aggressiveness, if they are combined and input into the multimodal large model together, a less than ideal situation may occur. Summary of the invention

[0005] The embodiments of the present application provide a method and related equipment for generating adversarial samples of a multimodal large model, thereby achieving, at least to a certain extent, a one-time adversarial attack on the multimodal large model and simultaneously generating multimodal adversarial samples of different data forms and with semantic consistency.

[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0007] According to one aspect of an embodiment of the present application, a method for generating adversarial samples for a multimodal large model is provided, and the method for generating adversarial samples for the multimodal large model comprises: inputting a multimodal data group into a target multimodal large model for iterative training multiple times to obtain an iterative loss and an adversarial sample group, wherein the multimodal data group comprises multimodal data in different data forms, and the data form of each adversarial sample in the adversarial sample group corresponds to the multimodal data; judging the consistency between each adversarial sample in the adversarial sample group to obtain a consistency loss; and updating parameters of the target multimodal large model according to the iterative loss and the consistency loss until a predetermined end condition is reached, and then terminating the iterative training to obtain a target adversarial sample group.

[0008] In a possible implementation of the present application, the parameters of the target multimodal large model are updated according to the iterative loss and the consistency loss until a predetermined end condition is reached, the iterative training is terminated, and a target adversarial sample group is obtained, specifically including: obtaining a total loss according to the iterative loss and the consistency loss; according to the total loss, the parameters of the target multimodal large model are updated until a predetermined end condition is reached, the iterative training is terminated, and a target adversarial sample group is obtained.

[0009] In a possible implementation of the present application, the target multimodal large model is updated with parameters according to the total loss until a predetermined end condition is reached, the iterative training is terminated, and a target adversarial sample group is obtained, specifically including: inputting the adversarial sample group into the target multimodal large model to obtain an output result; if the output result is a correct result, the target multimodal large model is updated with parameters according to the total loss, and the next round of iterative training is performed; if the output result is an incorrect result, the predetermined end condition is reached, the iterative training is terminated, and the adversarial sample group is output as the target adversarial sample group.

[0010] In a possible implementation of the present application, the target multimodal large model is updated with parameters according to the total loss until a predetermined end condition is reached, the iterative training is terminated, and a target adversarial sample group is obtained, specifically including: according to the total loss, the target multimodal large model is updated with parameters, and the number of iterative training is updated; if the number of iterative training exceeds a predetermined iteration number threshold, the predetermined end condition is reached, the iterative training is terminated, and the adversarial sample group of this iteration is output as the target adversarial sample group.

[0011] In a possible implementation of the present application, the multimodal data group is input into the target multimodal large model for iterative training multiple times to obtain iterative loss and adversarial sample group, specifically including: inputting the multimodal data group into the target multimodal large model to obtain iterative loss and iterative gradient; and obtaining the adversarial sample group according to the iterative gradient.

[0012] In a possible implementation of the present application, obtaining the adversarial sample according to the iterative gradient specifically includes: calculating the adversarial noise according to the iterative gradient; adding the adversarial noise to each multimodal data of the multimodal data group for decoding processing to obtain a corresponding adversarial sample group.

[0013] In a possible implementation of the present application, the anti-noise can be obtained by the following formula:

[0014] nosie=ε·sign(grad)

[0015] Among them, nosie is the adversarial noise, ε is the perturbation size, grad is the iterative gradient, sign(grad) is the sign function, when grad>0, sign(grad) takes 1, when grad<0, sign(grad) takes -1.

[0016] In a possible implementation of the present application, determining the consistency between adversarial samples in the adversarial sample group to obtain a consistency loss specifically includes: inputting the adversarial sample group into a consistency discrimination model to obtain a consistency loss.

[0017] According to one aspect of an embodiment of the present application, an adversarial sample generation device for a multimodal large model is provided, and the adversarial sample generation device for the multimodal large model includes: a data group input module, used to input a multimodal data group into a target multimodal large model for iterative training multiple times to obtain an iterative loss and an adversarial sample group, wherein the multimodal data group includes multimodal data in different data forms, and the data form of each adversarial sample in the adversarial sample group corresponds to the multimodal data; a consistency judgment module, used to judge the consistency between each adversarial sample in the adversarial sample group to obtain a consistency loss; a target sample generation module, used to update parameters of the target multimodal large model according to the iterative loss and the consistency loss until a predetermined end condition is reached, thereby terminating the iterative training and obtaining a target adversarial sample group.

[0018] In a possible implementation of the present application, the target sample generation module specifically includes: a total loss determination submodule, used to obtain the total loss based on the iterative loss and the consistency loss; a target sample generation submodule, used to update the parameters of the target multimodal large model according to the total loss until a predetermined end condition is reached, the iterative training is ended, and a target adversarial sample group is obtained.

[0019] In a possible implementation of the present application, the target sample generation submodule specifically includes: a result output unit, which is used to input the adversarial sample group into the target multimodal large model to obtain an output result; a first judgment unit, which is used to update the parameters of the target multimodal large model according to the total loss if the output result is a correct result, and perform the next round of iterative training; a second judgment unit, which is used to meet a predetermined end condition if the output result is an incorrect result, end the iterative training, and output the adversarial sample group as the target adversarial sample group.

[0020] In a possible implementation of the present application, the target sample generation submodule specifically includes: an iterative update unit, used to update the parameters of the target multimodal large model according to the total loss, and update the number of iterative training; an iterative stop unit, used to meet a predetermined end condition if the number of iterative training exceeds a predetermined iteration number threshold, end the iterative training, and output the adversarial sample group of this iteration as the target adversarial sample group.

[0021] In a possible implementation of the present application, the data group input module specifically includes: a data group input submodule, which is used to input the multimodal data group into the target multimodal large model to obtain iterative loss and iterative gradient; an adversarial sample generation submodule, which is used to obtain an adversarial sample group according to the iterative gradient.

[0022] In a possible implementation of the present application, the adversarial sample generation submodule specifically includes: an adversarial noise determination unit, used to calculate the adversarial noise according to the iterative gradient; a data perturbation adding unit, used to add the adversarial noise to each multimodal data of the multimodal data group for decoding processing to obtain a corresponding adversarial sample group.

[0023] In a possible implementation of the present application, the anti-noise can be obtained by the following formula:

[0024] nosie=ε·sign(grad)

[0025] Among them, nosie is the adversarial noise, ε is the perturbation size, grad is the iterative gradient, sign(grad) is the sign function, when grad>0, sign(grad) takes 1, when grad<0, sign(grad) takes -1.

[0026] In a possible implementation of the present application, the consistency judgment module specifically includes: a consistency judgment submodule, which is used to input the adversarial sample group into a consistency discrimination model to obtain a consistency loss.

[0027] According to one aspect of an embodiment of the present application, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the adversarial sample generation method for a multimodal large model as described in the above embodiment is implemented.

[0028] According to one aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the adversarial sample generation method for a multimodal large model as described in the above embodiments.

[0029] A computer program product comprises one or more computer programs, characterized in that when the one or more computer programs are executed by one or more processors, the steps of the method for generating adversarial samples of a multimodal large model as described in the above-mentioned embodiments are implemented.

[0030] In the technical solutions provided by some embodiments of the present application, after inputting a multimodal data group into a target multimodal large model, multimodal adversarial samples in different data forms can be generated by attacking the target multimodal large model once, without the need to perform multiple iterative calculations according to the data form, thereby improving the efficiency of generating multimodal adversarial samples for the multimodal large model and reducing the deployment complexity, thereby satisfying the purpose of obtaining adversarial samples in multiple data forms through one-time deployment. At the same time, because multimodal adversarial samples are obtained through a one-time adversarial attack and still have certain high-level semantic connections with each other, it is more effective to use the adversarial samples generated by the embodiments of the present application to conduct adversarial attack tests on multimodal large models.

[0031] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification, are used to explain the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. In the drawings:

[0033] Figure 1 An exemplary implementation environment diagram is shown to which the technical solution of the embodiments of the present application can be applied.

[0034] Figure 2 A schematic flow chart of a method for generating adversarial samples for a multimodal large model provided in an embodiment of the present application is shown.

[0035] Figure 3 Shown according to Figure 2 A specific implementation flowchart of step S100 in the adversarial sample generation method for a multimodal large model shown in the corresponding embodiment.

[0036] Figure 4 Shown according to Figure 2 A specific implementation flowchart of step S300 in the adversarial sample generation method for a multimodal large model shown in the corresponding embodiment.

[0037] Figure 5 A schematic diagram of the structure of an adversarial sample generation device for a multimodal large model provided in an embodiment of the present application is shown.

[0038] Figure 6 A schematic diagram of the structure of a computer system of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0039] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more comprehensive and complete and fully convey the concept of the example embodiments to those skilled in the art.

[0040] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present application. However, those skilled in the art will appreciate that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, known methods, devices, realizations or operations are not shown or described in detail to avoid blurring the various aspects of the application.

[0041] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0042] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0043] Figure 1A schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of the present application can be applied is shown.

[0044] like Figure 1 As shown, the system architecture may include terminal devices (such as Figure 1 The embodiment of the present invention is a schematic diagram of a mobile phone 101, a tablet computer 102, a portable computer 103, a desktop computer, etc., a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as a wired communication link, a wireless communication link, etc.

[0045] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.

[0046] The user can use the terminal device to interact with the server 105 through the network 104 to receive or send messages, etc. The server 105 can be a server that provides various services. For example, the user uses the terminal device 103 (or the terminal device 101 or 102) to upload a multimodal data group to the server 105, and the server 105 can input the multimodal data group into the target multimodal large model to obtain iterative loss and adversarial noise, wherein the multimodal data group contains multimodal data in different data forms; the adversarial noise is added to each multimodal data of the multimodal data group for decoding processing to obtain a corresponding adversarial sample group, and the data form of each adversarial sample in the adversarial sample group corresponds to the multimodal data.

[0047] It should be noted that the adversarial sample generation method for the multimodal large model provided in the embodiment of the present application is generally executed by the server 105, and accordingly, the adversarial sample generation device for the multimodal large model is generally arranged in the server 105. However, in other embodiments of the present application, the terminal device may also have similar functions as the server, so as to execute the adversarial sample generation solution for the multimodal large model provided in the embodiment of the present application.

[0048] The implementation details of the technical solution of the embodiment of the present application are described in detail below:

[0049] Figure 2 A flowchart of a method for generating adversarial samples of a multimodal large model according to an embodiment of the present application is shown. The method for generating adversarial samples of a multimodal large model can be executed by a server, which can be Figure 1 Refer to the server shown in Figure 2As shown, the adversarial sample generation method of the multimodal large model at least includes:

[0050] S100, inputting the multimodal data group into the target multimodal large model for iterative training multiple times to obtain iterative loss and adversarial sample groups, wherein the multimodal data group includes multimodal data in different data forms, and the data form of each adversarial sample in the adversarial sample group corresponds to the multimodal data.

[0051] S200, determining the consistency between the adversarial samples in the adversarial sample group to obtain a consistency loss.

[0052] S300, updating the parameters of the target multimodal large model according to the iteration loss and the consistency loss until a predetermined end condition is reached, terminating the iterative training, and obtaining a target adversarial sample group.

[0053] In an embodiment of the present application, a multimodal data group is input into a target multimodal large model for iterative training multiple times, and an iteration is performed for each input. Each iteration generates an adversarial sample and an iterative loss. According to the iterative loss combined with the subsequent consistency loss, the parameters of the target multimodal large model are updated to complete an iteration. In the process of iterative training, when the predetermined end condition is reached, the iterative training is terminated, and the adversarial sample group obtained in this iteration can be output as the target adversarial sample group. That is, the embodiment of the present application can generate multimodal adversarial samples in different data forms when the target multimodal large model is attacked once, and there is no need to perform multiple iterative calculations according to the data form, which improves the efficiency of generating multimodal adversarial samples for the multimodal large model. At the same time, the adversarial attack process is integrated into one iteration, which also simplifies the test process, making the generation and testing of multimodal adversarial samples more consistent and convenient, reducing the deployment complexity, and meeting the purpose of obtaining adversarial samples in multiple data forms in one deployment. Moreover, because multimodal adversarial samples still have certain high-level semantic connections with each other through a one-time adversarial attack, it is more effective to use the adversarial samples generated by the embodiments of the present application to conduct multimodal large model adversarial attack tests. Since the multimodal adversarial samples in the target adversarial sample group of the present application have good semantic consistency, the adversarial tests conducted using these samples can more realistically simulate the attack scenarios in actual applications. The separate testing methods of the prior art may result in poor testing results, because adversarial samples in different data forms may not be able to effectively jointly attack the target model. By introducing consistency constraints when generating adversarial samples, the embodiments of the present application ensure that the generated adversarial samples can more effectively test the robustness of the model in a multimodal scenario, thereby improving the success rate of adversarial attacks.

[0054] In general, the embodiments of the present application effectively overcome the shortcomings of the prior art by improving generation efficiency, reducing deployment complexity, ensuring consistency, and improving test results. By integrating strategies for adversarial noise calculation and consistency constraints, the embodiments of the present application achieve optimization of multimodal adversarial sample generation and meet the needs of practical applications.

[0055] In S100, the multimodal data group includes multimodal data in different data forms. The multimodal data can be divided into text data, image data, audio data, video data, etc. according to different data forms.

[0056] It should be noted that in the following embodiments, a multimodal data group including text data T and image data I will be taken as an example for description.

[0057] Specifically, in some embodiments, the specific implementation of step S100 can refer to the following embodiments. Figure 2 The detailed description of step S100 in the method for generating adversarial samples of a multimodal large model shown in the corresponding embodiment, in the method for generating adversarial samples of a multimodal large model, step S100 may include the following steps:

[0058] S110, input the multimodal data set into the target multimodal large model to obtain iterative loss and iterative gradient.

[0059] S120: Obtain an adversarial sample group according to the iterative gradient.

[0060] In this embodiment, after the multimodal data set is input into the target multimodal large model, in addition to outputting the corresponding results, the loss function (i.e., iterative loss) Loss is also iterated. ori And gradient information (i.e. iterative gradient) grad.

[0061] After obtaining the iterative gradient, various gradient methods can be used to form adversarial perturbations, and then form adversarial samples to obtain adversarial sample groups.

[0062] In S110, Loss ori The calculation of Loss varies depending on the task of the target multimodal model. For example, if the target multimodal model is a classification model, then Loss ori It is calculated by the cross entropy loss function; if the target multimodal large model is a target detection large model, then Loss ori It is calculated by the intersection-over-union loss function; if the target multimodal large model is another large model, it can also be calculated by other corresponding methods.

[0063] In some embodiments of the present application, a multimodal data group is input into a target multimodal large model, i.e., the multimodal data in the multimodal data group is converted into a feature vector of the multimodal data (e.g., a text vector, an image vector), and then the multimodal data vector is input into the target multimodal large model.

[0064] The input multimodal data set can be cleaned and standardized to provide high-quality basic data for subsequent adversarial sample generation.

[0065] In S120, various gradient methods are used to form adversarial perturbations, and an adversarial sample group is obtained by combining the adversarial perturbations with various multimodal data.

[0066] Specifically, in some embodiments, the specific implementation of step S120 can refer to the following embodiments. Figure 3 The detailed description of step S120 in the adversarial sample generation method for a multimodal large model shown in the corresponding embodiment, in the adversarial sample generation method for a multimodal large model, step S120 may include the following steps:

[0067] According to the iterative gradient, adversarial noise is calculated.

[0068] The adversarial noise is added to each multimodal data of the multimodal data group for decoding processing to obtain a corresponding adversarial sample group.

[0069] In this embodiment, the adversarial noise can be obtained by fast gradient sign method (FGSM), projected gradient descent (PGD), C&W attack (Carlini&Wagner) or DeepFool, etc. Adversarial noise generation can discover the loopholes and weaknesses of the model under adversarial attacks, thereby helping to design more effective defenses.

[0070] After obtaining the adversarial noise, add the adversarial noise to the text data T and image data I for decoding processing to obtain the corresponding text adversarial samples. and image adversarial examples Where i represents the i-th iteration. By processing multimodal data in an integrated manner, the inconsistency problem caused by separate processing of data forms is solved. Multimodal adversarial samples with consistency can effectively perform comprehensive testing.

[0071] In some embodiments, when adversarial noise is added to text data T and image data I for decoding processing, pixel perturbation can be used on the image data I and semantic perturbation can be used on the text data T. The text data T and image data I can be text vectors and image vectors converted into feature vector forms.

[0072] In some embodiments, the generated adversarial noise may be in various forms so that the most suitable adversarial sample can be selected according to requirements.

[0073] In some embodiments, the above anti-noise can be derived from the following formula:

[0074] nosie=ε·sign(grad)

[0075] Among them, nosie is the adversarial noise, ε is the perturbation size, grad is the iterative gradient, sign(grad) is the sign function, when grad>0, sign(grad) takes 1, when grad<0, sign(grad) takes -1.

[0076] In S200, consistency loss is introduced to ensure the consistency of high-level semantics between text and image adversarial samples. Specifically, the consistency between adversarial samples is judged by auxiliary discriminant model and included in the calculation of the final loss function, which effectively maintains the connection between multimodal data. The adversarial samples generated in this way can not only attack the target large model, but also maintain semantic consistency, which enhances the effectiveness of adversarial attack.

[0077] Specifically, in some embodiments, the specific implementation of step S200 can refer to the following embodiments. Figure 2 The detailed description of step S200 in the method for generating adversarial samples of a multimodal large model shown in the corresponding embodiment, in the method for generating adversarial samples of a multimodal large model, step S200 may include the following steps:

[0078] The adversarial sample group is input into a consistency discrimination model to obtain a consistency loss.

[0079] In this embodiment, the auxiliary discriminant model (any effective open source text image consistency discriminant model can be selected) is used to compare the adversarial samples in the adversarial sample group (for example, text adversarial samples and image adversarial examples The consistency judgment between consist .

[0080] Specifically, the feature vector of each adversarial sample in the adversarial sample group is input into the input layer of the consistency discrimination model, and the input layer is preprocessed and then passes through the consistency judgment network to output the consistency score. Finally, the output layer (such as the classification layer, regression layer, etc.) obtains the consistency result and consistency loss according to the consistency score.

[0081] Since the consistency discrimination model is obtained through training with a large number of samples, maintenance only needs to be adjusted through samples. Compared with other estimation methods, the maintenance cost is lower. Basically, it only needs to maintain the adversarial sample group, which reduces maintenance costs and improves code stability.

[0082] Specifically, the training method of the above consistency discrimination model specifically includes:

[0083] An adversarial sample group sample set is obtained, where the adversarial sample group sample set includes a plurality of adversarial sample group samples, and each adversarial sample group sample is marked with a corresponding consistency label.

[0084] The samples of the adversarial sample group are input into the consistency discrimination model one by one to obtain consistency results.

[0085] According to the output consistency result and the consistency label, the parameters of the consistency discrimination model are updated until a predetermined end condition is reached, and the training is terminated to obtain a trained consistency discrimination model.

[0086] In an embodiment of the present application, during training, an adversarial sample group sample set including multiple adversarial sample group samples can be first obtained, and each adversarial sample group sample is marked with a corresponding consistency label; then the multiple adversarial sample group samples are divided into a training set, a validation set, and a test set according to a predetermined ratio, and then the parameters of the encoder and the decoder in the consistency discrimination model are adjusted and determined according to the adversarial sample group samples included in the training set, the validation set, and the test set to obtain a trained consistency discrimination model.

[0087] When training the model, the adversarial sample group sample set can be divided into a training set, a validation set, and a test set, and then training is performed based on the training set, validation is performed based on the validation set, and testing is performed based on the test set to obtain a trained consistency discrimination model.

[0088] Before training based on the training set, the adversarial sample group samples in the training set can be preprocessed. The preprocessing includes vector adjustment, normalization, data enhancement, and category encoding.

[0089] After obtaining the enhanced training set, the consistency discrimination model can be trained based on the enhanced training set to update the parameters and weights in the network.

[0090] Specifically, the adversarial sample group samples in the training set are input into the consistency discrimination model to obtain the consistency result output by the consistency discrimination model. The consistency result is compared with the consistency label and the loss function is calculated. Then the stochastic gradient descent method is used to minimize the loss function, and the parameters and weights in the consistency discrimination model are updated by back propagation until the loss function meets the predetermined conditions, such as convergence or less than a predetermined threshold.

[0091] In some embodiments, the consistency result includes a target contour and a target contour parameter, and the consistency label may include at least one of a contour label and a contour parameter label. When calculating the loss function, the loss function is obtained by comparing the target contour with the contour label and / or comparing the target contour parameter with the contour parameter label.

[0092] After training, the image consistency discrimination network with updated parameters of the training set can be verified based on the validation set. Specifically, the image consistency discrimination network is debugged based on the validation set data, and when the loss function meets the predetermined conditions, the model parameters of this stage are output. If the loss function does not meet the predetermined conditions, the hyperparameters such as the learning rate are automatically adjusted to train the next round of network model.

[0093] When the loss function calculated on the validation set meets the predetermined conditions, the parameters and weights can be retained, and then the retained parameters can be tested based on the test set. Specifically, the test set data is input so that the image consistency discrimination network with retained parameters and weights can output consistency results and model weights. The model loss and corresponding weights of multiple rounds are compared, and the model weight with the minimum loss is output to determine the trained consistency discrimination model.

[0094] After obtaining the trained consistency discrimination model, the consistency discrimination of the adversarial sample group can be completed based on the consistency discrimination model.

[0095] In addition, no data augmentation is required for the data input in the validation set and the test set.

[0096] In other embodiments, each adversarial sample in the adversarial sample group may be vectorized into a corresponding adversarial sample feature vector, and then the distances between the feature vectors (such as Euclidean distance, cosine vector) may be compared, and the consistency loss may be determined based on the distances between the feature vectors.

[0097] In S300, back propagation and iterative optimization are performed according to the obtained iterative loss and consistency loss until a predetermined end condition is reached, and the iterative training is terminated to obtain a target adversarial sample group.

[0098] Specifically, in some embodiments, the specific implementation of step S300 can be found in Figure 4 . Figure 4 is based on Figure 2 The detailed description of step S300 in the method for generating adversarial samples of a multimodal large model shown in the corresponding embodiment, in the method for generating adversarial samples of a multimodal large model, step S300 may include the following steps:

[0099] S310, obtaining a total loss according to the iteration loss and the consistency loss.

[0100] S320, updating the parameters of the target multimodal large model according to the total loss until a predetermined end condition is reached, terminating the iterative training, and obtaining a target adversarial sample group.

[0101] In an embodiment of the present application, a total loss is first calculated based on the iteration loss and the consistency loss, and then the parameters of the target multimodal large model are updated based on the total loss to complete one iteration.

[0102] During the iterative training process, when the predetermined end condition is reached, the iterative training ends, and the adversarial sample group obtained in this iteration can be output as the target adversarial sample group.

[0103] In S310 , the iteration loss and the consistency loss may be weightedly summed to obtain a final loss, ie, a total loss.

[0104] Specifically, in some embodiments, the total loss is calculated as follows:

[0105] Loss total =Loss ori +βLoss consist

[0106] Among them, Loss total is the total loss, and β is the weight parameter, which is defined at initialization.

[0107] In S320, there may be multiple predetermined end conditions, such as a multimodal adversarial sample generated in a certain iteration. and The adversarial attack is successful, or the maximum number of iterations has been reached.

[0108] Specifically, in some embodiments, the specific implementation of step S320 can refer to the following embodiments. Figure 4 The detailed description of step S320 in the method for generating adversarial samples of a multimodal large model shown in the corresponding embodiment, in the method for generating adversarial samples of a multimodal large model, step S320 may include the following steps:

[0109] Inputting the adversarial sample group into the target multimodal large model to obtain an output result;

[0110] If the output result is a correct result, the parameters of the target multimodal large model are updated according to the total loss, and the next round of iterative training is performed;

[0111] If the output result is an erroneous result, the predetermined end condition is reached, the iterative training is terminated, and the adversarial sample group is output as the target adversarial sample group.

[0112] In this embodiment, the predetermined end condition is that the generated adversarial sample group (for example, including text adversarial samples) and image adversarial examples ) to achieve the purpose of successful adversarial attack, that is, after the adversarial sample group is input into the target multimodal large model, the target multimodal large model can output wrong results.

[0113] When the adversarial sample group is input into the target multimodal large model and an incorrect result is obtained, it is proved that the adversarial sample is a qualified adversarial sample, which can make the target multimodal large model output a result different from the original data (i.e., the multimodal data that has not been perturbed).

[0114] It should be noted that for target multimodal large models with different tasks, the methods for determining whether their output results are erroneous results are different.

[0115] If the target multimodal large model is a classification large model, it is only necessary to compare the classification results obtained by inputting the adversarial sample group into the target multimodal large model with the classification results obtained by inputting the original data (i.e., the multimodal data that has not been disturbed) into the target multimodal large model to determine whether the output result is an incorrect result. When they are consistent, the output result is a correct result rather than an incorrect result; when they are inconsistent, the output result is an incorrect result.

[0116] If the target multimodal model is a target detection model, it is necessary to compare the detection results obtained by inputting the adversarial sample group into the target multimodal model with the detection results obtained by inputting the original data (i.e., the multimodal data that has not been disturbed) into the target multimodal model to see if they are consistent, or to compare the detection results obtained by inputting the adversarial sample group into the target multimodal model with the detection labels corresponding to the multimodal data to see if they are consistent, so as to determine whether the output result is an incorrect result. When they are consistent, the output result is a correct result rather than an incorrect result; when they are inconsistent, the output result is an incorrect result.

[0117] If the target multimodal large model is a generative large model, an auxiliary model is required to determine whether the output result is an erroneous result. Specifically, the generative result obtained by inputting the adversarial sample group into the target multimodal large model and the generative result obtained by inputting the original data (i.e., the multimodal data that has not been disturbed) into the target multimodal large model are input into the auxiliary model together to determine whether the output result is an erroneous result.

[0118] Specifically, in some other embodiments, the specific implementation of step S320 can refer to the following embodiments. Figure 4The detailed description of step S320 in the method for generating adversarial samples of a multimodal large model shown in the corresponding embodiment, in the method for generating adversarial samples of a multimodal large model, step S320 may include the following steps:

[0119] According to the total loss, parameters of the target multimodal large model are updated, and the number of iterative training is updated.

[0120] If the number of iterative training exceeds a predetermined threshold number of iterations, a predetermined termination condition is reached, the iterative training is terminated, and the adversarial sample is output as a target adversarial sample.

[0121] In this embodiment, the predetermined end condition is a predetermined iteration number threshold. When a multimodal data group has undergone multiple iterations, if it still cannot obtain a suitable adversarial sample after the predetermined number of iterations, the iterative training should be stopped at this time, and the adversarial sample of the current iteration should be output as the target adversarial sample to avoid unnecessary waste of resources.

[0122] The following describes an apparatus embodiment of the present application, which can be used to execute the method for generating adversarial samples of a multimodal large model in the above-mentioned embodiment of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the embodiment of the method for generating adversarial samples of a multimodal large model in the above-mentioned embodiment of the present application.

[0123] Figure 5 A block diagram of an adversarial sample generation device for a multimodal large model according to an embodiment of the present application is shown.

[0124] Reference Figure 5 As shown, according to an embodiment of the present application, an adversarial sample generating device 500 for a multimodal large model includes: a data group input module 510, a consistency judgment module 520 and a target sample generating module 530.

[0125] Among them, the data group input module 510 is used to input the multimodal data group into the target multimodal large model for iterative training multiple times to obtain an iterative loss and an adversarial sample group, wherein the multimodal data group contains multimodal data in different data forms, and the data form of each adversarial sample in the adversarial sample group corresponds to the multimodal data; the consistency judgment module 520 is used to judge the consistency between each adversarial sample in the adversarial sample group to obtain a consistency loss; the target sample generation module 530 is used to update the parameters of the target multimodal large model according to the iterative loss and the consistency loss until a predetermined end condition is reached, the iterative training is ended, and the target adversarial sample group is obtained.

[0126] In a possible implementation of the present application, the target sample generation module specifically includes: a total loss determination submodule, used to obtain the total loss based on the iterative loss and the consistency loss; a target sample generation submodule, used to update the parameters of the target multimodal large model according to the total loss until a predetermined end condition is reached, the iterative training is ended, and a target adversarial sample group is obtained.

[0127] In a possible implementation of the present application, the target sample generation submodule specifically includes: a result output unit, which is used to input the adversarial sample group into the target multimodal large model to obtain an output result; a first judgment unit, which is used to update the parameters of the target multimodal large model according to the total loss if the output result is a correct result, and perform the next round of iterative training; a second judgment unit, which is used to meet a predetermined end condition if the output result is an incorrect result, end the iterative training, and output the adversarial sample group as the target adversarial sample group.

[0128] In a possible implementation of the present application, the target sample generation submodule specifically includes: an iterative update unit, used to update the parameters of the target multimodal large model according to the total loss, and update the number of iterative training; an iterative stop unit, used to meet a predetermined end condition if the number of iterative training exceeds a predetermined iteration number threshold, end the iterative training, and output the adversarial sample group of this iteration as the target adversarial sample group.

[0129] In a possible implementation of the present application, the data group input module specifically includes: a data group input submodule, which is used to input the multimodal data group into the target multimodal large model to obtain iterative loss and iterative gradient; an adversarial sample generation submodule, which is used to obtain an adversarial sample group according to the iterative gradient.

[0130] In a possible implementation of the present application, the adversarial sample generation submodule specifically includes: an adversarial noise determination unit, used to calculate the adversarial noise according to the iterative gradient; a data perturbation adding unit, used to add the adversarial noise to each multimodal data of the multimodal data group for decoding processing to obtain a corresponding adversarial sample group.

[0131] In a possible implementation of the present application, the anti-noise can be obtained by the following formula:

[0132] nosie=ε·sign(grad)

[0133] Among them, nosie is the adversarial noise, ε is the perturbation size, grad is the iterative gradient, sign(grad) is the sign function, when grad>0, sign(grad) takes 1, when grad<0, sign(grad) takes -1.

[0134] In a possible implementation of the present application, the consistency judgment module specifically includes: a consistency judgment submodule, which is used to input the adversarial sample group into a consistency discrimination model to obtain a consistency loss.

[0135] In an embodiment of the present application, a multimodal data group is input into a target multimodal large model for iterative training multiple times, and an iteration is performed for each input. Each iteration generates an adversarial sample and an iterative loss. According to the iterative loss combined with the subsequent consistency loss, the parameters of the target multimodal large model are updated to complete an iteration. In the process of iterative training, when the predetermined end condition is reached, the iterative training is terminated, and the adversarial sample group obtained in this iteration can be output as the target adversarial sample group. That is, the embodiment of the present application can generate multimodal adversarial samples in different data forms when the target multimodal large model is attacked once, and there is no need to perform multiple iterative calculations according to the data form, which improves the efficiency of generating multimodal adversarial samples for the multimodal large model. At the same time, the adversarial attack process is integrated into one iteration, which also simplifies the test process, making the generation and testing of multimodal adversarial samples more consistent and convenient, reducing the deployment complexity, and meeting the purpose of obtaining adversarial samples in multiple data forms in one deployment. Moreover, because multimodal adversarial samples still have certain high-level semantic connections with each other through a one-time adversarial attack, it is more effective to use the adversarial samples generated by the embodiments of the present application to conduct multimodal large model adversarial attack tests. Since the multimodal adversarial samples in the target adversarial sample group of the present application have good semantic consistency, the adversarial tests conducted using these samples can more realistically simulate the attack scenarios in actual applications. The separate testing methods of the prior art may result in poor testing results, because adversarial samples in different data forms may not be able to effectively jointly attack the target model. By introducing consistency constraints when generating adversarial samples, the embodiments of the present application ensure that the generated adversarial samples can more effectively test the robustness of the model in a multimodal scenario, thereby improving the success rate of adversarial attacks.

[0136] In general, the embodiments of the present application effectively overcome the shortcomings of the prior art by improving generation efficiency, reducing deployment complexity, ensuring consistency, and improving test results. By integrating strategies for adversarial noise calculation and consistency constraints, the embodiments of the present application achieve optimization of multimodal adversarial sample generation and meet the needs of practical applications.

[0137] Figure 6A schematic diagram of the structure of a computer system suitable for implementing an electronic device of an embodiment of the present application is shown.

[0138] It should be noted that Figure 6 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0139] like Figure 6 As shown, the computer system includes a central processing unit (CPU) 1801, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1802 or the program loaded from the storage part 1808 to the random access memory (RAM) 1803, such as executing the method described in the above embodiment. In RAM 1803, various programs and data required for system operation are also stored. CPU 1801, ROM 1802 and RAM 1803 are connected to each other through bus 1804. Input / output (I / O) interface 1805 is also connected to bus 1804.

[0140] The following components are connected to the I / O interface 1805: an input section 1806 including a keyboard, a mouse, etc.; an output section 1807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to the I / O interface 1805 as needed. A removable medium 1811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1810 as needed so that a computer program read therefrom is installed into the storage section 1808 as needed.

[0141] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 1809, and / or installed from a removable medium 1811. When the computer program is executed by a central processing unit (CPU) 1801, various functions defined in the system of the present application are executed.

[0142] It should be noted that the computer-readable medium shown in the embodiment of the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by an instruction execution system, device or device or used in combination with it. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable computer program is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0143] The flowchart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. Wherein, each box in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above-mentioned module, program segment, or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0144] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. The names of these units do not, in some cases, constitute limitations on the units themselves.

[0145] As another aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiment.

[0146] The present specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figure 1 to Figure 4 The method of the embodiment shown in the figure can be specifically executed by referring to Figure 1 to Figure 4 The specific description of the illustrated embodiment will not be repeated here.

[0147] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0148] Through the description of the above implementation methods, it is easy for those skilled in the art to understand that the example implementation methods described here can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the implementation methods of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the implementation methods of the present application.

[0149] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common knowledge or customary technical means in the art that are not disclosed in the present application.

[0150] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for generating adversarial samples for a multimodal large model, characterized in that: The adversarial sample generation method of the multimodal large model includes: Inputting the multimodal data group into the target multimodal large model for iterative training multiple times to obtain iterative loss and adversarial sample groups, wherein the multimodal data group includes multimodal data in different data forms, and the data form of each adversarial sample in the adversarial sample group corresponds to the multimodal data; Determining the consistency between the adversarial samples in the adversarial sample group to obtain a consistency loss; According to the iterative loss and the consistency loss, the parameters of the target multimodal large model are updated until a predetermined end condition is reached, and the iterative training is terminated to obtain a target adversarial sample group.

2. The method for generating adversarial samples of a multimodal large model according to claim 1, characterized in that: The step of updating the parameters of the target multimodal large model according to the iteration loss and the consistency loss until a predetermined end condition is reached, terminating the iterative training, and obtaining the target adversarial sample group specifically includes: Obtaining a total loss according to the iteration loss and the consistency loss; According to the total loss, the parameters of the target multimodal large model are updated until a predetermined end condition is reached, and the iterative training is terminated to obtain a target adversarial sample group.

3. The method for generating adversarial samples of a multimodal large model as claimed in claim 2, characterized in that: The method of updating the parameters of the target multimodal large model according to the total loss until a predetermined end condition is reached, terminating the iterative training, and obtaining the target adversarial sample group specifically includes: Inputting the adversarial sample group into the target multimodal large model to obtain an output result; If the output result is a correct result, the parameters of the target multimodal large model are updated according to the total loss, and the next round of iterative training is performed; If the output result is an erroneous result, the predetermined end condition is reached, the iterative training is terminated, and the adversarial sample group is output as the target adversarial sample group.

4. The method for generating adversarial samples of a multimodal large model as claimed in claim 2, characterized in that: The method of updating the parameters of the target multimodal large model according to the total loss until a predetermined end condition is reached, terminating the iterative training, and obtaining the target adversarial sample group specifically includes: According to the total loss, updating the parameters of the target multimodal large model and updating the number of iterative training; If the number of iterative training exceeds a predetermined threshold number of iterations, a predetermined termination condition is reached, the iterative training is terminated, and the adversarial sample group of this iteration is output as a target adversarial sample group.

5. The method for generating adversarial samples of a multimodal large model according to claim 1, characterized in that: The multimodal data set is input into the target multimodal large model for multiple times for iterative training to obtain iterative loss and adversarial sample set, specifically including: Input the multimodal data set into the target multimodal large model to obtain iterative loss and iterative gradient; According to the iterative gradient, an adversarial sample group is obtained.

6. The method for generating adversarial samples of a multimodal large model as claimed in claim 5, characterized in that: The step of obtaining an adversarial sample group according to the iterative gradient specifically includes: Calculating adversarial noise according to the iterative gradient; The adversarial noise is added to each multimodal data of the multimodal data group for decoding processing to obtain a corresponding adversarial sample group.

7. A device for generating adversarial samples of a multimodal large model, characterized in that: The adversarial sample generating device of the multimodal large model comprises: A data group input module, used for inputting a multimodal data group into a target multimodal large model for iterative training multiple times to obtain an iterative loss and an adversarial sample group, wherein the multimodal data group includes multimodal data in different data forms, and the data form of each adversarial sample in the adversarial sample group corresponds to the multimodal data; A consistency judgment module, used to judge the consistency between the adversarial samples in the adversarial sample group and obtain a consistency loss; The target sample generation module is used to update the parameters of the target multimodal large model according to the iterative loss and the consistency loss until a predetermined end condition is reached, and the iterative training is ended to obtain a target adversarial sample group.

8. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for generating adversarial samples for a multimodal large model as described in any one of claims 1 to 6 is implemented.

9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the one or more processors to implement the adversarial sample generation method for a multimodal large model as described in any one of claims 1 to 6.

10. A computer program product comprising one or more computer programs, characterized in that When the one or more computer programs are executed by one or more processors, the steps of the method for generating adversarial samples of a multimodal large model described in any one of claims 1 to 6 are implemented.