Perturbed sample pair generation method

By adopting the enhanced and complementary spatial search methods in the vision-language pretrained model, images and text are synergistically perturbed, and perturbation strategies are optimized to improve generalization effect, the problem of poor anti-robustness evaluation in the existing technology is solved, and more efficient security evaluation is achieved.

CN119782770BActive Publication Date: 2025-06-17TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510280412.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-17
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

When the prior art evaluates the anti-robustness of vision-language pretrained models, the synergistic perturbation effect is poor and the perturbation process cannot be effectively optimized, resulting in poor generalization effect of the generated perturbation data.

Method used

Using a method based on reinforcement and complementary spatial search, multiple perturbers add perturbation data generated by perturbation strategies to the image and text respectively, optimize the perturbation strategy to improve the synergistic perturbation effect, and adjust the perturbation strategy through the target reward value until the preset conditions are met.

Benefits of technology

By optimizing the perturbation strategy, the generalization effect of the generated perturbation sample pair is improved and the security evaluation ability of multimodal models is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782770B_ABST
    Figure CN119782770B_ABST
Patent Text Reader

Abstract

The present invention provides a method for generating perturbed sample pairs and a method for evaluating the performance of a multimodal model based on reinforcement and complementary space search. The method for generating perturbed sample pairs includes: obtaining sample pairs and matching labels of the sample pairs; for each sample pair, repeatedly perform the following operations until the perturbation strategy meets a preset condition: for each of multiple perturbators, add perturbation data to the image and the text respectively to obtain the perturbed images and perturbed texts of the multiple perturbators respectively, where the perturbation data is generated based on the perturbation strategy; input the perturbed images and perturbed texts of the multiple perturbators into the multimodal model respectively to obtain multiple perturbed matching data; adjust the perturbation strategy based on the target reward value determined by the multiple perturbed matching data and the matching labels; determine the perturbed images and perturbed texts in the case where the perturbation strategy meets the preset condition as the target perturbed sample pairs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-modal model security, and more particularly to a method for generating perturbation sample pairs, a method, device, and equipment for evaluating the performance of a multi-modal model based on reinforcement and complementary space search. Background Art

[0002] Vision-language pre-trained models are usually pre-trained on a large number of image-text pairs and learn the cross-modal representations therein so that the models can understand and correlate the information in images and texts. Due to the emergence of various adversarial perturbations, the adversarial robustness of vision-language pre-trained models has attracted people's attention. The purpose of security evaluation is to identify and quantify the vulnerability of the model when facing adversarial attacks. Therefore, the security of the vision-language pre-trained model can be tested by designing perturbation samples and perturbing the vision-language pre-trained model.

[0003] In the related art, often, each modality in the vision-language pre-trained model is perturbed separately by various single-modal perturbation data. The effect of collaborative perturbation is not good. However, since the specific perturbation process inside the vision-language pre-trained model cannot be known, effective optimization cannot be performed, and due to excessive focus on significant features, the generalization effect of the generated perturbation data is poor. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method for generating perturbation sample pairs, a method, device, and equipment for evaluating the performance of a multi-modal model based on reinforcement and complementary space search.

[0005] According to a first aspect of the present invention, there is provided a method for generating perturbation sample pairs, including: obtaining a sample pair and a matching label of the sample pair, the sample pair including an image and a text, and the matching label characterizing the consistency of the image and the text; for the sample pair, repeatedly performing the following operations until the perturbation strategy meets a preset condition: for each of a plurality of perturbators, adding perturbation data to the image and the text respectively to obtain a perturbed image and a perturbed text of each of the plurality of perturbators, the perturbation data being generated based on the perturbation strategy; inputting the perturbed image and the perturbed text of each of the plurality of perturbators into a multi-modal model respectively to obtain a plurality of perturbed matching data; adjusting the perturbation strategy based on a target reward value determined by the plurality of perturbed matching data and the matching label; and determining the perturbed image and the perturbed text in the case where the perturbation strategy meets the preset condition as a target perturbation sample pair.

[0006] According to an embodiment of the present invention, the above-mentioned perturbation matching data includes a similarity and a perturbation matching result. The above-mentioned similarity characterizes the degree of consistency between the above-mentioned image and text. The above method further includes: for each of the above-mentioned perturbators, using an alignment reward function to determine a sub-alignment reward value based on the above-mentioned similarity; determining a sub-perturbation reward value based on the above-mentioned perturbation matching result and the above-mentioned matching label; determining a candidate reward value of the above-mentioned perturbator based on the above-mentioned sub-alignment reward value and the above-mentioned sub-perturbation reward value; determining the above-mentioned target reward value based on the candidate reward values of each of the above-mentioned multiple perturbators.

[0007] According to an embodiment of the present invention, the above-mentioned perturbation strategy includes an image perturbation direction, a text perturbation direction, and a strategy distribution parameter. The above-mentioned image perturbation direction characterizes the probability of adding perturbations to random sub-regions in the above-mentioned image. The above-mentioned text perturbation direction characterizes the probability of adding perturbations to random sub-texts in the above-mentioned text. The above-mentioned random sub-regions and the above-mentioned random sub-texts have a corresponding relationship. The above-mentioned strategy distribution parameter is used to control the consistency difference between the image perturbation caused by the above-mentioned image perturbation direction and the text perturbation caused by the above-mentioned text perturbation direction. The adjustment of the above-mentioned perturbation strategy based on the above-mentioned target reward value includes: determining the above-mentioned image perturbation direction and the above-mentioned text perturbation direction based on the sub-alignment reward value of the above-mentioned target reward value; adjusting the above-mentioned strategy distribution parameter based on the sub-perturbation reward value of the above-mentioned target reward value to obtain an adjusted strategy distribution function.

[0008] According to an embodiment of the present invention, each of the above-mentioned perturbators has a preset perturbation iteration value. The above-mentioned preset conditions include: the number of iterations reaches the above-mentioned preset perturbation iteration value or the above-mentioned perturbation strategy causes the above-mentioned target reward value to converge.

[0009] According to an embodiment of the present invention, adding the above-mentioned perturbation data to the above-mentioned image and the above-mentioned text respectively to obtain a perturbed image and a perturbed text includes: adding random noise to the above-mentioned image to obtain the above-mentioned perturbed image; performing synonym replacement or word embedding on sub-texts in the above-mentioned text to obtain the above-mentioned perturbed text.

[0010] According to an embodiment of the present invention, adding the above-mentioned perturbation data to the above-mentioned image and the above-mentioned text respectively to obtain a perturbed image and a perturbed text includes: adding random noise to the above-mentioned image to obtain the above-mentioned perturbed image; performing synonym replacement or word embedding on sub-texts in the above-mentioned text to obtain the above-mentioned perturbed text.

[0011] According to an embodiment of the present invention, adding random noise to the above-mentioned image to obtain the above-mentioned perturbed image includes: determining a target sub-region of the above-mentioned image based on the above-mentioned image perturbation direction; adding noise to the above-mentioned target sub-region to obtain the above-mentioned perturbed image.

[0012] According to a second aspect of the present invention, there is provided a method for evaluating the performance of a multimodal model based on enhanced and complementary space search, including: obtaining a target perturbation sample pair, where the target perturbation sample pair includes target perturbation image data and target perturbation text data, and the target perturbation sample pair is obtained by using the method of the first aspect; using the target perturbation sample pair to evaluate the performance of the multimodal model to obtain a performance evaluation result, and the performance evaluation result is used to improve the security of the multimodal model.

[0013] A third aspect of the present invention provides a device for generating perturbation sample pairs, including: an acquisition module, configured to acquire a sample pair and a matching label of the sample pair, where the sample pair includes an image and text, and the matching label represents the consistency of the image and the text; a perturbation data adding module, configured to, for each of a plurality of perturbators, add perturbation data to the image and the text respectively to obtain a perturbed image and a perturbed text of each of the plurality of perturbators, and the perturbation data is generated based on the perturbation strategy; an input module, configured to input the perturbed images and the perturbed texts of each of the plurality of perturbators into a multimodal model respectively to obtain a plurality of perturbed matching data; an adjustment module, configured to adjust the perturbation strategy based on a target reward value determined by the plurality of perturbed matching data and the matching label; a determination module, configured to determine the perturbed image and the perturbed text when the perturbation strategy meets the preset conditions as the target perturbation sample pair.

[0014] A fourth aspect of the present invention provides a device for evaluating the performance of a multimodal model based on enhanced and complementary space search, including: a target sample pair acquisition module, configured to acquire a target perturbation sample pair, where the target perturbation sample pair includes target perturbation image data and target perturbation text data, and the target perturbation sample pair is obtained by using the method of the first aspect; a performance evaluation module, configured to use the target perturbation sample pair to evaluate the performance of the multimodal model to obtain a performance evaluation result, and the performance evaluation result is used to improve the security of the multimodal model.

[0015] A fifth aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, where the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0016] A sixth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0017] The seventh aspect of the present invention further provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of the above method.

[0018] According to an embodiment of the present invention, multiple perturbators respectively add perturbations generated by a perturbation strategy to an image and a text, to obtain respective perturbed images and perturbed texts of the multiple perturbators, input the respective perturbed images and perturbed texts of the multiple perturbators into a multimodal model to obtain multiple perturbed matching data, and optimize the perturbation strategy in the perturbed multimodal model according to a target reward value determined by the perturbed data and the matching labels, thereby optimizing the collaborative perturbation in the perturbation process and enhancing the generalization effect of the generated perturbed sample pairs. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features and advantages of the present invention will become clearer.

[0020] Figure 1 The application scenario diagram of the perturbed sample pair generation method, the multimodal model performance evaluation method based on reinforcement and complementary space search, and the device according to the embodiment of the present invention is shown.

[0021] Figure 2 The flowchart of the perturbed sample pair generation method according to the embodiment of the present invention is shown.

[0022] Figure 3 The comparison schematic diagram of the perturbation results generated by multiple perturbators and the sample pairs according to the embodiment of the present invention is shown;

[0023] Figure 4 The architecture schematic diagram according to the embodiment of the present invention is shown;

[0024] Figure 5 The sample pair generation example diagram according to the embodiment of the present invention is shown;

[0025] Figure 6 The flowchart of the multimodal model performance evaluation method based on reinforcement and complementary space search according to the embodiment of the present invention is shown;

[0026] Figure 7 The structural block diagram of the perturbed sample pair generation device according to the embodiment of the present invention is shown;

[0027] Figure 8 The structural block diagram of the multimodal model performance evaluation device based on reinforcement and complementary space search according to the embodiment of the present invention is shown;

[0028] Figure 9 The block diagram of the electronic device suitable for implementing the perturbed sample pair generation method according to the embodiment of the present invention is shown. Detailed Implementation Modes

[0029] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, obviously, one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0030] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0031] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0032] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0033] Vision-language pre-training models are usually pre-trained on a large number of image-text pairs and learn the cross-modal representations therein so that the models can understand and associate the information in images and texts. Due to the emergence of various adversarial perturbations, the adversarial robustness of vision-language pre-training models has attracted people's attention. The purpose of security evaluation is to identify and quantify the vulnerability of the model when facing adversarial perturbations. Therefore, the security of vision-language pre-training models can be revealed by designing perturbation methods. Co-perturbation is a perturbation method designed specifically for multi-modal models. Co-perturbation achieves a higher perturbation success rate on multi-modal models by leveraging the intrinsic consistency and complementarity between modalities, thereby achieving a high perturbation success rate for multi-modal models. Research has shown that vision-language pre-training models are more sensitive to co-perturbation than single-modal models.

[0034] Current research on collaborative perturbations of vision - language pre - trained models mainly focuses on white - box perturbations and transfer - based perturbations. These perturbations require all or part of the prior knowledge of the perturbing vision - language pre - trained model, and it is challenging to obtain the structural knowledge of the black - box victim model. Moreover, it is difficult to effectively collaborate on perturbations between images and texts. Currently, only the downstream task outputs of the vision - language pre - trained model can be accessed. In addition, existing vision - language pre - trained collaborative perturbation methods iteratively optimize adversarial samples on a set of image - text pairs, which may ignore other potential perturbation targets and lead to local overfitting. The overfitting problem mentioned in the present invention refers to the fact that adversarial samples overly focus on significant features at the expense of details, which will affect the generalization ability of adversarial samples on tasks and models. This is because these methods use embedding guidance and self - enhancement to improve the diversity of adversarial instances on the optimization path to increase adversarial transferability. However, since they only search for optimization targets in fixed image - text pairs, the performance and generalization ability of the perturbations are relatively weak.

[0035] In view of this, an embodiment of the present invention provides a method for generating perturbation sample pairs. The method includes: obtaining a sample pair and a matching label of the sample pair, where the sample pair includes an image and a text, and the matching label represents the consistency between the image and the text; for the sample pair, repeatedly perform the following operations until the perturbation strategy meets a preset condition: for each of multiple perturbing agents, add perturbation data to the image and the text respectively to obtain the perturbed image and perturbed text of each perturbing agent, where the perturbation data is generated based on the perturbation strategy; input the perturbed images and perturbed texts of multiple perturbing agents into a multimodal model respectively to obtain multiple perturbed matching data; adjust the perturbation strategy based on the target reward value determined by the multiple perturbed matching data and the matching label; determine the perturbed image and perturbed text when the perturbation strategy meets the preset condition as the target perturbation sample pair.

[0036] Figure 1 The application scenario diagram of the method for generating perturbation sample pairs, the method for evaluating the performance of a multimodal model based on reinforcement and complementary space search, and the device according to an embodiment of the present invention is shown.

[0037] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0038] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0039] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on.

[0040] The server 105 can be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0041] It should be noted that the method for generating perturbation sample pairs and the method for evaluating the performance of a multimodal model based on reinforcement and complementary space search provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the device for generating perturbation sample pairs and the device for evaluating the performance of a multimodal model based on reinforcement and complementary space search provided by the embodiments of the present invention can generally be set in the server 105. The method for generating perturbation sample pairs and the method for evaluating the performance of a multimodal model based on reinforcement and complementary space search provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the device for generating perturbation sample pairs and the device for evaluating the performance of a multimodal model based on reinforcement and complementary space search provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0042] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0043] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the scenario described below Figures 2 to 5A detailed description is given of the method for generating perturbation sample pairs according to an embodiment of the present invention.

[0044] Figure 2 The flowchart of the method for generating perturbation sample pairs according to an embodiment of the present invention is shown.

[0045] As Figure 2 shown, the method for generating perturbation sample pairs in this embodiment includes operation S210 to operation S250.

[0046] In operation S210, a sample pair and the matching label of the sample pair are obtained.

[0047] Among them, the sample pair includes an image and text, and the matching label characterizes the consistency between the image and the text.

[0048] For the sample pair, the following S210 to S240 are repeatedly executed until the perturbation strategy meets the preset conditions:

[0049] In operation S220, for each of multiple perturbators, perturbation data is respectively added to the image and the text to obtain the perturbed image and the perturbed text of each perturbator.

[0050] Among them, the perturbation data is generated based on the perturbation strategy.

[0051] According to an embodiment of the present invention, the process of finding the optimal perturbation strategy can be modeled as a strategy search process in a complementary space according to the Markov decision process.

[0052] The process of finding the optimal perturbation strategy can be modeled by the following formula (1).

[0053]

[0054] Among them, S is the state set, A is the action set, is the perturbation strategy, R is the reward function, γ is the discount rate, I is the perturbation search space, and t represents a finite number of steps.

[0055] For a multimodal model, the state set S includes the feature information of the image, the feature information of the text, and the interaction information between the image and the text. The state set S can be described by the following formula (2).

[0056]

[0057] Among them, represents the feature information of the image, represents the feature information of the text, represents the interaction information between the two modalities.

[0058] According to an embodiment of the present invention, considering the differences in features between images and texts, the perturbations to the multimodal model are divided into two parts. By adding perturbation variables, the images and texts are perturbed simultaneously. The perturbation performed by each perturber can be represented by the following formula (3).

[0059]

[0060] Among them, represents the perturbation to the image, and the perturbed image can be expressed as , represents the perturbation to the text, and the perturbed text can be expressed as .

[0061] In operation S230, the perturbed images and perturbed texts of multiple perturbers are respectively input into the multimodal model to obtain multiple perturbed matching data.

[0062] Figure 3 FIG. shows a comparison schematic diagram of the perturbation results generated by multiple perturbers and the sample pairs according to an embodiment of the present invention.

[0063] As Figure 3 shown, the figure may include a sample pair 310, the perturbed data 320 generated by the first perturber, and the perturbed data 330 generated by the second perturber; the sample pair 310 includes an image 311 and the text "This is a cat lying on the lawn" 312; the perturbed data 320 generated by the first perturber includes a first perturbed image 321 and the first perturbed text "This is a cat sitting on the lawn" 322; the perturbed data 330 generated by the second perturber includes a second perturbed image 331 and the second perturbed text "This is a cat lying prone on the lawn" 332. By respectively inputting the perturbed data 320 generated by the first perturber and the perturbed data 330 generated by the second perturber into the multimodal model, first perturbed matching data and second perturbed matching data can be obtained.

[0064] In operation S240, based on the target reward value determined by multiple perturbed matching data and matching labels, the perturbation strategy is adjusted.

[0065] According to an embodiment of the present invention, the first perturbed matching data and the second perturbed matching data obtained above can be respectively calculated through a preset reward function to obtain a first perturbation reward value and a second perturbation reward value. The target reward value is selected according to the foregoing two reward values, thereby adjusting the perturbation strategy.

[0066] In operation S250, the perturbed images and perturbed texts when the perturbation strategy meets the preset conditions are determined as the target perturbation sample pairs.

[0067] According to an embodiment of the present invention, multiple perturbators respectively add perturbations generated by a perturbation strategy to an image and text, obtaining respective perturbed images and perturbed texts of the multiple perturbators, inputting the respective perturbed images and perturbed texts of the multiple perturbators into a multimodal model to obtain multiple perturbed matching data, and optimizing the perturbation strategy in the perturbed multimodal model according to a target reward value determined by the perturbed data and matching labels, thereby optimizing the collaborative perturbation in the perturbation process and enhancing the generalization effect of the generated perturbed sample pairs.

[0068] According to an embodiment of the present invention, the above-mentioned perturbed matching data includes a similarity and a perturbed matching result, the similarity characterizes the degree of consistency between the image and the text, and the method further includes: for each perturbator, using an alignment reward function to determine a sub-alignment reward value based on the similarity; determining a sub-perturbation reward value based on the perturbed matching result and the matching label; determining a candidate reward value of the perturbator based on the sub-alignment reward value and the sub-perturbation reward value; and determining a target reward value based on the respective candidate reward values of the multiple perturbators.

[0069] The candidate reward value of each perturbator can be described by the following formula (4).

[0070]

[0071] Wherein, is the sub-alignment reward value, is the sub-perturbation reward value, can be calculated by the following formula (5).

[0072]

[0073] Wherein, represents the perturbation to the image, represents the perturbation to the text, represents the probability of data alignment after respectively applying perturbations to the image and the text, represents minimizing the alignment loss, that is, maximizing the alignment probability.

[0074] According to an embodiment of the present invention, the sub-perturbation reward value can represent the error rate of the multimodal model under the condition of perturbation.

[0075] The overall environment in the image, text, and their interaction process can be represented by the following formula (6).

[0076]

[0077] Wherein, I represents an overall environment including an image, text, and their interaction process, I v represents the image environment, I t represents the text environment, I c represents the environment of the interaction process.

[0078] According to an embodiment of the present invention, in order to determine a target reward value from candidate reward values, the state and action set of each step of each disturber can be recorded. The disturbance strategy can be regarded as the probability of taking action A in state S, which can be expressed by the following formula (7).

[0079]

[0080] Among them, s0 represents the state at the 0th step, a0 represents the action taken in state s0, s1 represents the state at the 1st step, and a1 represents the action taken in state s1.

[0081] According to an embodiment of the present invention, by aligning the reward function to determine the sub-alignment reward value based on similarity, the semantic consistency between the image and the text is evaluated, and the sub-disturbance reward value is determined according to the disturbance matching result and the matching label, and then the disturbance result of the disturber is evaluated. The sub-alignment reward value and the sub-disturbance reward value are combined to determine the target reward value to optimize the generated sample pair, and then the multi-modal model can be fully performance-tested.

[0082] According to an embodiment of the present invention, each of the above-mentioned disturbers has a preset disturbance iteration value, and the preset conditions include: the number of iterations reaches the preset disturbance iteration value or the disturbance strategy makes the target reward value converge.

[0083] According to an embodiment of the present invention, each of the above-mentioned disturbers is assigned a preset disturbance iteration value, which represents the number of steps that can be used in the entire search process. At the beginning of the search, each disturber is assigned a preset iterative disturbance value, which can be freely used before the disturbance ends. That is, for each disturbance of the disturber, it is not necessary to reach the alignment state every time, as long as the maximum disturbance rate is achieved before reaching the alignment state.

[0084] According to an embodiment of the present invention, the above-mentioned disturbance strategy includes an image disturbance direction, a text disturbance direction, and a policy distribution parameter. The image disturbance direction represents the probability of adding disturbances to random sub-regions in the image, and the text disturbance direction represents the probability of adding disturbances to random sub-texts in the text. The random sub-regions and the random sub-texts have a corresponding relationship. The policy distribution parameter is used to control the consistency difference between the image disturbance caused by the image disturbance direction and the text disturbance caused by the text disturbance direction. Adjusting the disturbance strategy based on the target reward value includes: determining the image disturbance direction and the text disturbance direction based on the sub-alignment reward value of the target reward value; adjusting the policy distribution parameter based on the sub-disturbance reward value of the target reward value to obtain an adjusted policy distribution function.

[0085] According to an embodiment of the present invention, the disturbance strategy can be represented by the following formula (8).

[0086]

[0087] Among them, d v represents the image perturbation direction, and d l represents the text perturbation direction, p v represents the image perturbation intensity, and p l represents the text perturbation intensity, and θ represents the policy distribution parameter.

[0088] According to the embodiments of the present invention, the tendency of policy selection can be changed by adjusting the value of the policy distribution parameter θ, so as to optimize the perturbation effect and help to find the optimal perturbation strategy in different scenarios.

[0089] According to the embodiments of the present invention, adding perturbation data to the image and text respectively to obtain a perturbed image and a perturbed text includes: adding random noise to the image to obtain a perturbed image; performing synonym replacement or word embedding on sub-texts in the text to obtain a perturbed text.

[0090] According to the embodiments of the present invention, the above image perturbation can also be performed on the features extracted by the neural network of the image, so as to change the feature distribution. The image perturbation direction can also be to perturb in directions such as color and texture in the image, and the present invention does not limit this.

[0091] According to the embodiments of the present invention, the above text perturbation direction can be, for example, determining whether to perform synonym replacement or word embedding on the text, adjusting the sentence structure and other perturbation directions, and the present invention does not limit this.

[0092] According to the embodiments of the present invention, the image perturbation intensity and the text perturbation intensity can be adjusted adaptively, and can also be adjusted according to requirements.

[0093] According to the embodiments of the present invention, the sub-regions of the perturbation generated in the image are controlled through the image perturbation direction in the perturbation strategy, and random noise is added in the sub-regions to generate a perturbed image; the sub-texts that need to be perturbed are determined through the text perturbation direction in the perturbation strategy, and synonym replacement or word embedding is performed on the sub-texts for perturbation, so as to quickly generate diverse perturbation data. The consistency difference between the image perturbation and the text perturbation is controlled through the policy distribution in the perturbation strategy, which can improve the co-perturbation effect of the image and the text.

[0094] According to the embodiments of the present invention, adding random noise to the image to obtain a perturbed image includes: determining the target sub-region of the image based on the image perturbation direction; adding noise to the target sub-region to obtain a perturbed image.

[0095] Exemplarily, for the above Figure 3When performing perturbations, it is found that adding noise to the region where the cat is located in the image and replacing the verb in "This is a cat lying on the lawn" with another word will result in a relatively high sub-alignment reward value. Then, in the next iteration process, it will be more inclined to determine the target sub-region as the region where the cat is located in the image, and determine the verb "lying" in the text as the target sub-text for adjustment.

[0096] According to the embodiments of the present invention, the corresponding cumulative reward will be calculated. Our ultimate goal is to find a perturbation strategy that maximizes the reward, that is, the optimal perturbation effect, which can be expressed by the following formula (9).

[0097]

[0098] Wherein, represents the perturbation strategy, R(τ) represents the cumulative reward of the trajectory τ, pπ(τ) represents the probability of the trajectory τ occurring under the perturbation strategy , and dπ represents the state distribution under the strategy π.

[0099] To obtain the above perturbation strategy, for each perturber, initialize the search state S = 0 and deduce forward from the last time step . When calculating the perturbation strategy, in fact, it is calculating the optimal perturbation strategy at step t and each state in the strategy . The sub-perturbation reward value can be expressed as xrj, where , the sub-perturbation reward value characterizes the loss generated by the perturber j when perturbing the perturbed object r. For the sake of concise representation of the formula, we stack xrj into an mn-dimensional vector , that is, the policy distribution parameter. In the policy distribution parameter, the probability of transferring the state of the perturbed object to is expressed as , where represents the true state of all perturbed objects in step s t , and this probability is between 0 and 1. To optimize to maximize the total expected reward of the perturbed objects from to to under a specific true state vector

[0100]

[0101] Wherein, represents the probability of transferring to state given the policy distribution parameter θ, state and state transition . represents the optimal value function in the state case, T represents the total number of time steps, and the above formula (10) needs to satisfy the conditions shown in formulas (11) to (13).

[0102] Figure 4 shows a schematic architecture diagram according to an embodiment of the present invention.

[0103] As Figure 4 shown, the real picture and real sample pair are input into the multimodal model and processed by different modules therein. The step of finding the target perturbation strategy is regarded as a problem of finding the maximum value of the target reward value through a finite number of steps in the complementary space established by the image and text. The search space is gradually compressed, and the perturbation strategy is optimized through inter-step and intra-step constraints, so as to determine the target perturbation strategy and generate a perturbation sample pair for the multimodal model.

[0104] Figure 5 shows an example diagram of sample pair generation according to an embodiment of the present invention.

[0105] As Figure 5 shown, among them, the above picture and the real description "A businessman wearing a yellow tie shows a frustrated expression" are input into the method according to the embodiment of the present invention for perturbation iteration. It can be seen that in step 10, "businessman" in the text becomes "person", in step 24, "yellow" in the text becomes "black", in step 45, "shows" becomes "wears", and in step 91, "expression" becomes "suit". The following picture and the real description "A man stands on a ladder cleaning the window" are input into the method according to the embodiment of the present invention for iteration. It can be seen that in step 9, "man" in the text becomes "person"; in step 36, "cleaning" becomes "wiping", in step 60, it changes from "wiping" to "hitting", and in step 79, it changes from "hitting" to "striking".

[0106] Figure 6 shows a flowchart of a multimodal model performance evaluation method based on reinforcement and complementary space search according to an embodiment of the present invention.

[0107] As Figure 6 shown, the perturbation sample pair generation method of this embodiment includes operations S610 to S620.

[0108] In operation S610, obtain the target perturbation sample pair.

[0109] Among them, the target perturbation sample pair includes target perturbation image data and target perturbation text data, and the target perturbation sample pair is obtained by using the above perturbation sample pair generation method.

[0110] In operation S620, the performance of the multi-modal model is evaluated using the target perturbation sample pair, and a performance evaluation result is obtained.

[0111] Among them, the performance evaluation result is used to improve the security of the multi-modal model.

[0112] According to the embodiments of the present invention, by inputting the perturbation sample pair into the multi-modal model, comparing the recognition result of the multi-modal model with the correct result of the original sample, calculating the accuracy rate, and recording the stability of the model after inputting the perturbation sample pair to determine the robustness of the model, so as to find the security vulnerabilities of the multi-modal model for targeted optimization and improvement of the multi-modal model, and increase the security of the multi-modal model.

[0113] Based on the above perturbation sample pair generation method, the present invention also provides a perturbation sample pair generation device. The following will be combined with Figure 7 This device will be described in detail.

[0114] Figure 7 Fig. shows a structural block diagram of a perturbation sample pair generation device according to an embodiment of the present invention.

[0115] As Figure 7 shown, the perturbation sample pair generation device 700 of this embodiment includes an acquisition module 710, a perturbation data addition module 720, an input module 730, an adjustment module 740, and a determination module 750.

[0116] The acquisition module 710 is used to acquire a sample pair and a matching label of the sample pair. The sample pair includes an image and text, and the matching label characterizes the consistency between the image and the text. In one embodiment, the acquisition module 710 can be used to perform the operation S210 described above, which will not be elaborated here.

[0117] The perturbation data addition module 720 is used to add perturbation data to the image and text respectively for each of multiple perturbators, to obtain the perturbed images and perturbed texts of the multiple perturbators respectively. The perturbation data is generated based on a perturbation strategy. In one embodiment, the perturbation data addition module 720 can be used to perform the operation S220 described above, which will not be elaborated here.

[0118] The input module 730 is used to input the perturbed images and perturbed texts of the multiple perturbators into the multi-modal model respectively, to obtain multiple perturbation matching data. In one embodiment, the input module 730 can be used to perform the operation S230 described above, which will not be elaborated here.

[0119] The adjustment module 740 is used to adjust the perturbation strategy based on the target reward value determined by the multiple perturbation matching data and the matching label. In one embodiment, the adjustment module 740 can be used to perform the operation S240 described above, which will not be elaborated here.

[0120] The determination module 750 is configured to determine the perturbed image and the perturbed text when the perturbation strategy meets the preset conditions as a target perturbation sample pair. In one embodiment, the determination module 750 may be configured to perform the operation S250 described above, which will not be elaborated herein.

[0121] According to an embodiment of the present invention, any multiple modules among the acquisition module 710, the perturbation data addition module 720, the input module 730, the adjustment module 740, and the determination module 750 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the acquisition module 710, the perturbation data addition module 720, the input module 730, the adjustment module 740, and the determination module 750 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in a suitable combination of any several of them. Alternatively, at least one of the acquisition module 710, the perturbation data addition module 720, the input module 730, the adjustment module 740, and the determination module 750 may be at least partially implemented as a computer program module, and when the computer program module is run, it may perform the corresponding functions.

[0122] Based on the above multi-modal model performance evaluation method based on reinforcement and complementary space search, the present invention also provides a multi-modal model performance evaluation device based on reinforcement and complementary space search. The following will be combined with Figure 8 to describe this device in detail.

[0123] Figure 8 Fig. shows a structural block diagram of a multi-modal model performance evaluation device based on reinforcement and complementary space search according to an embodiment of the present invention.

[0124] As Figure 8 shown, the multi-modal model performance evaluation device 800 of this embodiment based on reinforcement and complementary space search includes a target sample pair acquisition module 810 and a performance evaluation module 820.

[0125] The target sample pair acquisition module 810 is configured to acquire target perturbation sample pairs, where a target perturbation sample pair includes target perturbation image data and target perturbation text data, and the target perturbation sample pair is obtained by using the above-described perturbation sample pair generation method. In one embodiment, the target sample pair acquisition module 810 may be configured to perform the operation S610 described above, which will not be elaborated herein.

[0126] The performance evaluation module 820 is configured to use the target perturbation sample pairs to evaluate the performance of the multimodal model and obtain a performance evaluation result, where the performance evaluation result is used to improve the security of the multimodal model. In one embodiment, the performance evaluation module 820 may be configured to perform the operation S620 described above, which will not be elaborated herein.

[0127] According to an embodiment of the present invention, any multiple of the target sample pair acquisition module 810 and the performance evaluation module 820 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the target sample pair acquisition module 810 and the performance evaluation module 820 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the target sample pair acquisition module 810 and the performance evaluation module 820 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0128] Figure 9 The block diagram of an electronic device suitable for implementing the perturbation sample pair generation method according to an embodiment of the present invention is shown.

[0129] As Figure 9As shown, the electronic device 900 according to an embodiment of the present invention includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. The processor 901 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 901 can also include on-board memory for caching purposes. The processor 901 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0130] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to an embodiment of the present invention by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the program can also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 can also perform various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.

[0131] According to an embodiment of the present invention, the electronic device 900 can further include an input / output (I / O) interface 905, and the input / output (I / O) interface 905 is also connected to the bus 904. The electronic device 900 can further include one or more of the following components connected to the input / output (I / O) interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage section 908 as needed.

[0132] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods according to the embodiments of the present invention are implemented.

[0133] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903.

[0134] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method for generating perturbation samples provided by the embodiments of the present invention.

[0135] When the computer program is executed by the processor 901, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0136] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 909, and / or be installed from the removable medium 911. The program code included in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0137] In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, the above functions defined in the system of the embodiments of the present invention are performed. According to the embodiments of the present invention, the systems, devices, apparatuses, modules, units, etc. described above can be implemented by computer program modules.

[0138] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0140] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0141] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A method for generating perturbation sample pairs, characterized in that: The method comprises: Acquire a sample pair and a matching label of the sample pair, the sample pair comprising an image and a text, the matching label representing the consistency of the image and the text; For the sample pair, the following operations are repeated until the perturbation strategy meets the preset conditions: For each of the multiple disturbers, adding disturbance data to the image and the text respectively to obtain a disturbed image and a disturbed text of each of the multiple disturbers, wherein the disturbance data is generated based on the disturbance strategy; Inputting the disturbance images and the disturbance texts of the multiple disturbers into a multimodal model respectively to obtain a plurality of disturbance matching data; Adjusting the perturbation strategy based on the plurality of perturbation matching data and the target reward value determined by the matching tags; The disturbed image and the disturbed text when the perturbation strategy satisfies the preset condition are determined as a target perturbation sample pair.

2. The method according to claim 1, characterized in that The disturbance matching data includes similarity and disturbance matching results, wherein the similarity represents the consistency between the image and the text, and the method further includes: For each of the perturbers, Determining a sub-alignment reward value based on the similarity using an alignment reward function; Determining a sub-perturbation reward value based on the perturbation matching result and the matching label; determining a candidate reward value for the perturbator based on the sub-alignment reward value and the sub-disturbance reward value; The target reward value is determined based on candidate reward values ​​for each of the plurality of disturbers.

3. The method according to claim 2, characterized in that The perturbation strategy includes an image perturbation direction, a text perturbation direction and a strategy distribution parameter, wherein the image perturbation direction represents the probability of adding perturbation to a random sub-region in the image, and the text perturbation direction represents the probability of adding perturbation to a random sub-text in the text, wherein the random sub-region and the random sub-text have a corresponding relationship, and the strategy distribution parameter is used to control the consistency difference between the image perturbation caused by the image perturbation direction and the text perturbation caused by the text perturbation direction, and the perturbation strategy is adjusted based on the target reward value, including: Determining the image perturbation direction and the text perturbation direction based on the sub-alignment reward value of the target reward value; The policy distribution parameter is adjusted based on the sub-disturbance reward value of the target reward value to obtain an adjusted policy distribution parameter.

4. The method according to claim 1, characterized in that Each of the perturbers has a preset perturbation iteration value, and the preset conditions include: The number of iterations reaches the preset disturbance iteration value or the disturbance strategy makes the target reward value converge.

5. The method according to claim 3, characterized in that: The adding disturbance data to the image and the text respectively to obtain the disturbance image and the disturbance text comprises: Adding random noise to the image to obtain the disturbed image; The perturbed text is obtained by performing synonym replacement or word embedding on the sub-text in the text.

6. The method according to claim 5, characterized in that The adding random noise to the image to obtain the disturbed image comprises: determining a target sub-region of the image based on the image disturbance direction; Noise is added to the target sub-region to obtain the disturbed image.

7. A multimodal model performance evaluation method based on reinforcement and complementary space search, characterized in that: The method comprises: Obtain a target perturbation sample pair, wherein the target perturbation sample pair includes target perturbation image data and target perturbation text data, and the target perturbation sample pair is obtained by using the method described in any one of claims 1 to 6; The target perturbation sample pairs are used to perform a performance evaluation on the multimodal model to obtain a performance evaluation result, and the performance evaluation result is used to adjust the multimodal model.

8. A disturbance sample pair generating device, characterized in that: The device comprises: An acquisition module, used to acquire a sample pair and a matching label of the sample pair, wherein the sample pair includes an image and a text, and the matching label represents the consistency of the image and the text; A disturbance data adding module, used for adding disturbance data to the image and the text respectively for each of the multiple disturbers, so as to obtain the disturbance image and the disturbance text of each of the multiple disturbers, wherein the disturbance data is generated based on the disturbance strategy; An input module, used for inputting the disturbance images and the disturbance texts of the multiple disturbers into the multimodal model to obtain multiple disturbance matching data; An adjustment module, configured to adjust the perturbation strategy based on the plurality of perturbation matching data and a target reward value determined by the matching tags; A determination module is used to determine the disturbed image and the disturbed text as a target disturbance sample pair when the disturbance strategy meets a preset condition.

9. A multimodal model performance evaluation device based on enhanced and complementary space search, characterized in that: The device comprises: A target sample pair acquisition module, used to acquire a target perturbation sample pair, wherein the target perturbation sample pair includes target perturbation image data and target perturbation text data, and the target perturbation sample pair is obtained by using the method described in any one of claims 1 to 6; The performance evaluation module is used to use the target perturbation sample pair to perform performance evaluation on the multimodal model to obtain a performance evaluation result, and the performance evaluation result is used to improve the security of the multimodal model.

10. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Semantic depolarization attack method for vision-language pre-training model

    CN118332328A

  • Figure graph security improvement method and device based on confrontation prompt mining

    CN119443201A