Adversarial sample generation method and device of multi-modal fusion model

By generating adversarial samples of the multimodal fusion model using the modal alternation selection algorithm and Bessel function, the problem of insufficient robustness and security of the multimodal fusion model in complex environments is solved, and stronger adaptability and attack effect are achieved.

CN119399593BActive Publication Date: 2025-10-10BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411229248.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2025-10-10
Estimated Expiration
2044-09-03

AI Technical Summary

Technical Problem

Existing multimodal fusion models fail to fully utilize the information interaction between modalities in adversarial attacks, resulting in insufficient robustness and security, especially poor performance in complex environments.

Method used

A modal alternating selection algorithm is used to generate adversarial patches that have obvious effects in each modality through Bessel functions, and these patches are optimized in an iterative process. Finally, these patches are injected into the multimodal fusion model to generate optimal adversarial samples.

Benefits of technology

The robustness and security of the multimodal fusion model in adversarial attacks are improved, enabling it to perform better in complex environments and better adapt to changing real-world scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399593B_ABST
    Figure CN119399593B_ABST
Patent Text Reader

Abstract

The application provides a kind of multi-modal fusion model's adversarial sample generation method and device, the method comprises: obtaining multi-modal image, multi-modal image is input to multi-modal fusion model and obtains initial fusion image, multi-modal image includes multiple single modal images;Determine multiple initial adversarial samples based on multi-modal image;Each initial adversarial sample is injected into each single modal image to obtain each single modal perturbation image, based on each single modal perturbation image and each single modal image, through multi-modal fusion model, the semi-perturbation fusion image corresponding to each single modal perturbation image is obtained, based on each semi-perturbation fusion image and initial fusion image, determine the optimal initial adversarial sample corresponding to each single modal image;Based on each optimal initial adversarial sample, generate multiple sub-adversarial samples, determine the optimal adversarial sample of the multi-modal fusion model from multiple sub-adversarial samples.The application can improve the robustness and security of multi-modal fusion model in adversarial attack.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and in particular, to an adversarial sample generation method and device of a multi-modal fusion model. BACKGROUND

[0002] Deep neural networks (DNNs) have achieved very significant performance improvements on a variety of visual tasks, and thus, their security has received increasing attention. As one of the most basic tasks in computer vision, object detection is widely used in a large number of scenarios, including video supervision, autonomous driving, and human-computer interaction, and investigating the potential risks of object detection models has become an important task. Previous work has revealed that existing deep neural network-based detectors are vulnerable to adversarial perturbations, which is undoubtedly a potential huge security risk for deploying this type of object detector in systems with high security requirements. Adversarial attack methods are an important means to test the robustness and security of models.

[0003] Currently, a common method for generating adversarial samples is to add slight perturbations invisible to the human eye to the entire visible light photo; although this method is very effective, these perturbations cannot be achieved in the real world because it is impossible to do so for every pixel in the attack photo. To enhance the feasibility of adversarial attacks, patch-based attack methods are proposed, which only attack a small area where pixels are concentrated; in patch-based attack methods, adversarial patches are generated and injected into images to mislead deep network models. Recent investigations have shown that adversarial patches are highly concealed and can be implemented in the real world; for example, Patch-fool can mislead ViTs (Vision Transformer) in various visual tasks by adding a small patch to the visible light image; Invisible-Cloak can make the detector ignore the wearer of the clothes by printing a prepared pattern on the clothes; Naturalistic Attack can make the detected person hold a printed photo and be ignored by the detector; the above methods can detect the robustness and security of the model to some extent, but the current mature adversarial attack means is single-modal attack, that is, single-modal attack usually uses visible light images as attack targets, generates adversarial samples by perturbing the pixels of the target image or injecting patches, which requires the attacked object to have obvious texture information and be sensitive to pixel perturbation or patch attack; thus, the single-modal attack method can only be applied in high-visibility conditions and has poor effect in low-visibility environments such as heavy fog or night. Therefore, single-modal attack cannot cope with complex and variable real-world environments.

[0004] To address the challenges of single-modal attacks, multimodal fusion detection has attracted attention, and corresponding adversarial attack methods have emerged. For example, Unified-Attack simultaneously attacks visible light and infrared images. Currently, multimodal attack methods are based on single-modal attacks. While they can attack multiple modalities simultaneously, existing multimodal attack methods treat multiple modalities independently, attacking each modality separately. This means that each modality uses a separate adversarial sample. This method does not consider the information interaction between modalities and ignores the interactivity between modalities, resulting in poor performance of the model after adversarial training. Therefore, improving the robustness and security of multimodal fusion models in adversarial attacks is a technical problem that needs to be solved urgently. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a method and apparatus for generating adversarial samples for a multimodal fusion model to eliminate or improve one or more defects in the prior art.

[0006] One aspect of the present invention provides a method for generating adversarial samples for a multimodal fusion model, the method comprising the following steps:

[0007] Acquiring a multimodal image, and inputting the multimodal image into a multimodal fusion model to obtain an initial fused image, wherein the multimodal image includes multiple single-modal images;

[0008] Determining a plurality of initial adversarial examples based on the multimodal image;

[0009] Injecting each of the initial adversarial samples into each of the unimodal images to obtain each unimodal perturbation image, obtaining a semi-perturbation fusion image corresponding to each of the unimodal perturbation images through a multimodal fusion model based on each of the unimodal perturbation images and each of the unimodal images, and determining an optimal initial adversarial sample corresponding to each of the unimodal images based on each of the semi-perturbation fusion images and the initial fusion image;

[0010] A plurality of sub-adversarial samples are generated based on each of the optimal initial adversarial samples, and an optimal adversarial sample of the multimodal fusion model is determined from the plurality of sub-adversarial samples.

[0011] In some embodiments of the present invention, obtaining a semi-perturbed fusion image corresponding to each unimodal perturbation image based on each unimodal perturbation image and each unimodal image through a multimodal fusion model includes:

[0012] The first unimodal perturbation image and other second unimodal images in the multimodal image except the first unimodal image corresponding to the first unimodal perturbation image are input into the multimodal fusion model to obtain a semi-perturbation fusion image corresponding to the first unimodal perturbation image.

[0013] In some embodiments of the present invention, determining the optimal initial adversarial sample corresponding to each of the single-modal images based on each of the semi-perturbed fused images and the initial fused image includes:

[0014] Calculating the confidence difference between each of the semi-perturbed fused images and the initial fused image;

[0015] Determining the initial adversarial sample size corresponding to each of the semi-perturbed fusion images and the degree of deviation of the perturbation region contour before and after the attack;

[0016] The optimal initial adversarial sample corresponding to each of the single-modal images is determined based on the confidence difference, the size of the initial adversarial sample, and the degree of deviation of the perturbation area contour.

[0017] In some embodiments of the present invention, the calculation formula for the optimal initial adversarial example is:

[0018]

[0019] Among them, M * represents the optimal initial adversarial sample, M represents the initial adversarial sample, x vis represents the first single modality image, x inf represents the second single modality image, L s Represents the confidence difference, L d Indicates the degree of deviation of the disturbance area contour, L a represents the initial adversarial sample size, and λ1 and λ2 represent coefficients.

[0020] In some embodiments of the present invention, determining a plurality of initial adversarial examples based on the multimodal image includes:

[0021] Determining, based on the size of the multimodal image, a plurality of sample center point coordinates and anchor point position coordinates corresponding to each of the sample center point coordinates;

[0022] Based on the anchor point position coordinates corresponding to the center point coordinates of each sample, a sample contour curve corresponding to the center point coordinates of each sample is obtained by using a Bessel function;

[0023] Each initial adversarial sample is determined based on each of the sample contour curves.

[0024] In some embodiments of the present invention, the slopes of two curves passing through the same anchor point in each of the sample contour curves are the same.

[0025] In some embodiments of the present invention, generating a plurality of sub-adversarial examples based on each of the optimal initial adversarial examples, and determining the optimal adversarial example of the multimodal fusion model from the plurality of sub-adversarial examples includes:

[0026] Select, mutate, and crossover based on each of the optimal initial adversarial samples to obtain multiple sub-adversarial samples;

[0027] Injecting each sub-adversarial sample into the multimodal image to obtain a multimodal perturbation image corresponding to each sub-adversarial sample;

[0028] Inputting the multimodal perturbation image into the multimodal fusion model to obtain a fully perturbation fusion image;

[0029] An optimal adversarial sample of the multimodal fusion model is determined based on each of the fully perturbed fusion images and the initial fusion image.

[0030] In some embodiments of the present invention, the Bezier curve is:

[0031] B i (t) = f B (P i ,D,E,P i+1 )=(1-t) 3 P i +3t(1-t) 2 E+3t 2 (1-t)D+t 3 P i+1 t∈[0,1];

[0032] Among them, P i represents the i-th anchor point, P i+1 represents the i+1th anchor point, and D and E represent two control points.

[0033] According to another aspect of the present invention, a system for generating adversarial samples for a multimodal fusion model is also disclosed, including a processor, a memory, and a computer program stored in the memory, wherein the processor is used to execute the computer program. When the computer program is executed, the system implements the steps of the method described in any of the above embodiments.

[0034] According to yet another aspect of the present invention, a computer-readable storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any of the above embodiments are implemented.

[0035] The adversarial sample generation method and device for the multimodal fusion model of the present invention adopt a modal alternation selection algorithm to fully utilize the image information of the single modality, select the initial adversarial samples with outstanding attack effects in each single modality, and generate new sub-adversarial samples based on the outstanding initial adversarial samples under each selected single modality. The new sub-adversarial samples are then used to attack the multimodal fusion image, and the optimal adversarial sample of the multimodal fusion model is determined based on the attack effect. This method makes good use of the interactive information between multiple modalities and has stronger adaptability to complex environments, thereby improving the robustness and security of the multimodal fusion model in adversarial attacks.

[0036] Additional advantages, objects, and features of the present invention will be set forth in part in the following description and will become apparent to those skilled in the art upon examination of the following or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the structures particularly pointed out in the description and drawings.

[0037] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings described herein are intended to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention. The components in the drawings are not drawn to scale, but are merely for the purpose of illustrating the principles of the present invention. To facilitate the illustration and description of certain portions of the present invention, corresponding portions in the drawings may be exaggerated, that is, may be larger than other components in an exemplary device actually manufactured according to the present invention. In the drawings:

[0039] Figure 1 Schematic diagram of the process of generating adversarial samples for a multimodal fusion model in one embodiment of the present invention.

[0040] Figure 2 Schematic diagram of the effect of adversarial sample generation of a multimodal fusion model according to an embodiment of the present invention.

[0041] Figure 3 Schematic diagram of a flow chart of a method for generating adversarial samples for a multimodal fusion model according to another embodiment of the present invention.

[0042] Figure 4 This is a schematic diagram of contour generation based on Bessel functions according to an embodiment of the present invention.

[0043] Figure 5 Schematic diagram of adversarial sample boundary constraints according to an embodiment of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0045] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.

[0046] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0047] The current single-mode adversarial attack method has a relatively simple attack mode, and most of them attack on visible light images, which cannot adapt to complex application scenarios. Multi-mode adversarial attacks do not utilize the fusion information of multimodal images, but attack on their own modes without considering the information interaction between different modes, resulting in poor performance of the multimodal fusion model after adversarial training. Therefore, in response to the problem that there are few application scenarios for current single-mode attacks and the information interaction ignored by multimodal attacks, this application proposes a method and device for generating adversarial samples for a multimodal fusion model. This method and device starts from the shape, which has a significant effect on each mode. With the help of Bessel functions, an adversarial patch with the same shape in all modes is designed. It is continuously optimized in the iterative process and finally injected into images of different modes at the same time, thereby obtaining the optimal adversarial sample of the multimodal fusion model. The adversarial sample takes into account the information of multimodal images and focuses on information interaction, and has the potential to work in complex environments, so that the multimodal fusion model trained based on the adversarial sample has good robustness and security.

[0048] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0049] Figure 1 FIG. 1 is a flow chart of a method for generating adversarial samples of a multimodal fusion model according to an embodiment of the present invention. Figure 1 As shown, the adversarial sample generation method includes at least steps S10 to S40.

[0050] Step S10: Acquire a multimodal image, and input the multimodal image into a multimodal fusion model to obtain an initial fused image, wherein the multimodal image includes multiple single-modal images.

[0051] In this step, the multimodal image is a pair of paired unimodal images. Specifically, these paired unimodal images are fed into the multimodal fusion model to generate an initial fused image. Furthermore, to ensure a high-quality initial fused image, the initial fused image generated by the multimodal fusion model is further inspected using an object detector to determine whether the confidence level of objects in the image is below a preset threshold. If so, the initial fused image generated by the multimodal fusion model is deemed too poor to be effective against attacks, and is therefore ignored.

[0052] Step S20: determining a plurality of initial adversarial examples based on the multimodal image.

[0053] This step is to initialize the adversarial samples. The number of initial adversarial samples is set based on needs. The initial adversarial samples are used to attack each unimodal image.

[0054] Exemplarily, determining multiple initial adversarial samples based on the multimodal image specifically includes the following steps: determining multiple sample center point coordinates and anchor point position coordinates corresponding to each sample center point coordinate based on the size of the multimodal image; obtaining a sample contour curve corresponding to each sample center point coordinate through a Bessel function based on the anchor point position coordinates corresponding to each sample center point coordinate; and determining each initial adversarial sample based on each sample contour curve.

[0055] The size of the multimodal image can also be understood as the size of the image received by the multimodal fusion model. That is, this step determines the positions of multiple groups of center point coordinates and anchor points based on the image size input by the multimodal fusion model. Then, all anchor points in each group are smoothly connected through the Bessel function to obtain a smooth sample contour, thereby determining an initial adversarial sample.

[0056] Step S30: Injecting each of the initial adversarial samples into each of the unimodal images to obtain each unimodal perturbation image, obtaining a semi-perturbation fusion image corresponding to each of the unimodal perturbation images through a multimodal fusion model based on each of the unimodal perturbation images and each of the unimodal images, and determining the optimal initial adversarial sample corresponding to each of the unimodal images based on each of the semi-perturbation fusion images and the initial fusion image.

[0057] In this step, each initial adversarial example is injected into an image of a single modality. The perturbed and unperturbed unimodal images are fused using a multimodal fusion model, and the attack effectiveness of the initial adversarial example in the semi-perturbed fused image is measured. This step achieves alternating optimization of each modality, ultimately determining the optimal initial adversarial example for each modality.

[0058] Exemplarily, obtaining a semi-perturbed fusion image corresponding to each unimodal perturbation image based on each of the unimodal perturbation images and each of the unimodal images through a multimodal fusion model includes: inputting a first unimodal perturbation image and other second unimodal images in the multimodal image except the first unimodal image corresponding to the first unimodal perturbation image into the multimodal fusion model to obtain a semi-perturbed fusion image corresponding to the first unimodal perturbation image. In this embodiment, the semi-perturbed fusion image is obtained by fusing the perturbed first unimodal image with the other unperturbed unimodal images; for example, when there are two unimodal images and two initial adversarial samples, four unimodal perturbation images are obtained based on the two initial adversarial samples, the two unimodal images are defined as unimodal image A and unimodal image B, and the four unimodal perturbation images are defined as the first unimodal perturbation image A, the second unimodal perturbation image A, the first unimodal perturbation image B, and the second unimodal perturbation image B; In the first iteration, the multimodal fusion model fuses the first unimodal perturbation image A and the unimodal image B; in the second iteration, the multimodal fusion model fuses the second unimodal perturbation image A and the second unimodal image B; in the third iteration, the multimodal fusion model fuses the unimodal image A and the first unimodal perturbation image B; in the fourth iteration, the multimodal fusion model fuses the unimodal image A and the second unimodal perturbation image B; thereby finding the optimal initial adversarial samples corresponding to the unimodal image A and the unimodal image B respectively.

[0059] When determining the optimal initial adversarial sample for each unimodal image, the optimal initial adversarial sample can be calculated based on an adaptation function, i.e., the attack results of each initial adversarial sample on each unimodal image are quantified using the adaptation function. In one embodiment, determining the optimal initial adversarial sample for each unimodal image based on each semi-perturbed fused image and the initial fused image includes: calculating the confidence difference between each semi-perturbed fused image and the initial fused image; determining the size of the initial adversarial sample for each semi-perturbed fused image and the degree of deviation of the perturbed region contour before and after the attack; and determining the optimal initial adversarial sample for each unimodal image based on each confidence difference, the size of the initial adversarial sample, and the degree of deviation of the perturbed region contour. In this embodiment, the adaptation function is related to the confidence difference, the size of the initial adversarial sample, and the degree of deviation of the perturbed region contour.

[0060] For example, the calculation formula for the optimal initial adversarial sample (adaptation function) is:

[0061]

[0062] Among them, M * represents the optimal initial adversarial sample, M represents the initial adversarial sample, x vis represents the first single modality image, x infrepresents the second single modality image, L s Represents the confidence difference, L d Indicates the degree of deviation of the disturbance area contour, L a represents the initial adversarial sample size, and λ1 and λ2 represent coefficients.

[0063] In the above steps, the modality alternation selection algorithm based on information fusion determines the optimal initial adversarial sample for each modality. The modality alternation selection algorithm takes the initial adversarial sample as input, transforms it to obtain an image mask, and injects this image mask into one of the unimodal images to obtain the perturbed image corresponding to that modality, while ensuring that the images of the other modalities remain unchanged. The injection method of the initial adversarial sample is shown in the following formula: in, represents a single-modal perturbation image, x vis represents the single modality image before perturbation, C vis represents the control matrix corresponding to the unimodal image, ⌒ represents the Hadamard product of the matrix, and M is the mask matrix that characterizes the adversarial sample.

[0064] After a single-modal image is subjected to a single-modal attack, all paired modal images (one of which is a single-modal perturbation image) are used as input. A semi-perturbation fused image is generated through a multimodal fusion model. The initial adversarial example's attack effectiveness on the corresponding modality is then tested using an adaptation function. The remaining modalities are attacked in a similar manner, and their adaptation function values ​​are calculated. By calculating the attack effectiveness of all initial adversarial examples on all single-modal images, the initial adversarial example with the best attack effectiveness is selected as the master, accelerating convergence.

[0065] Step S40: generating a plurality of sub-adversarial samples based on each of the optimal initial adversarial samples, and determining an optimal adversarial sample of the multimodal fusion model from the plurality of sub-adversarial samples.

[0066] In this step, based on the optimal initial adversarial sample that performs well on each single-modal image determined in step S30, multiple sub-adversarial samples are derived. The derived sub-adversarial samples attack images of all modalities at the same time, and the optimal adversarial sample is determined by measuring the attack effect of each sub-adversarial sample.

[0067] Specifically, this step derives multiple sub-adversarial samples through a genetic algorithm, that is, generates multiple sub-adversarial samples based on each of the optimal initial adversarial samples, and determines the optimal adversarial sample of the multimodal fusion model from the multiple sub-adversarial samples. This may specifically include: selecting, mutating, and crossover based on each of the optimal initial adversarial samples to obtain multiple sub-adversarial samples; injecting each sub-adversarial sample into the multimodal image to obtain a multimodal perturbation image corresponding to each sub-adversarial sample; inputting the multimodal perturbation image into the multimodal fusion model to obtain a fully perturbed fusion image; and determining the optimal adversarial sample of the multimodal fusion model based on each of the fully perturbed fusion images and the initial fusion image. In this embodiment, the sub-adversarial samples are simultaneously injected into multiple unimodal images, and the multiple unimodal images perturbed by the sub-adversarial samples serve as input images of the multimodal fusion model, so that the multimodal fusion model outputs a fully perturbed fusion image corresponding to the sub-adversarial sample. The attack effect of the sub-adversarial sample on the fused image can be further measured based on a preset adaptation function; thus, based on the attack effects of multiple sub-adversarial samples, the optimal sub-adversarial sample is found as the optimal adversarial sample for the multimodal fusion model. It can be understood that the calculation formula for the optimal adversarial sample in this step is similar to the calculation formula for the optimal initial adversarial sample (adaptation function) in step S30, except that the confidence difference in this step is the difference between the confidence levels of the fully perturbed fused image and the initial fused image.

[0068] In addition, in order to ensure the rationality and quality of the generated initial adversarial sample, the position and size of the initial adversarial sample need to be specified; at the same time, in order to ensure the smoothness of the connection between different curve segments, the anchor points of the generated patch need to appear in the specified area. These requirements together form the boundary constraints of the adversarial sample. The smooth contour generation algorithm based on the Bessel function can flexibly characterize the contour of the adversarial patch and can easily change the shape of the adversarial patch during the iteration process; in one embodiment, a third-order Bessel function is used, such as Figure 4 As shown in Figure a, given four control points, the third-order Bessel function can form a smooth curve in the area containing the four control points. The curve passes through the first and fourth control points, but not the second and third control points. The specific formula is: i (t) = f B (P i ,D,E,P i+1 )=(1-t) 3 P i +3t(1-t) 2 +3t 2 (1-t)D+t 3 P i+1 t∈[0,1]; where P i represents the i-th anchor point, P i+1represents the i+1th anchor point, and D and E represent two control points.

[0069] exist Figure 4 In figure a, the Bezier curve is at the first control point P i The tangent direction is determined by the first control point P i Point to the second control point E, the curve is at the fourth control point P i+1 The tangent direction of the third control point D points to the fourth control point P i+1 , therefore, the second control point and the third control point control the direction and shape of the curve. In this embodiment, given a single-modal image as input, the anchor points of the initial adversarial sample are calculated in the adversarial sample initialization stage, and the two adjacent anchor points are used as the starting point and end point of a Bezier curve. Two control points are defined between the two anchor points to control the change of the curve, and the four control points required to draw the Bezier curve are obtained. A Bezier curve is obtained using the Bezier function; it can be understood that a set of Bezier curves can form the contour boundary line of the initial adversarial sample. Figure 4 As shown in Figure b, a Bezier curve is drawn between every two adjacent anchor points. Connecting all Bezier curve segments can obtain a closed curve contour. In addition, in order to ensure smooth connection of the curve, when calculating the control points matching the anchor points, keep B′ i-1 (P i )=B′ i (P i ), that is, the slopes of the two curves passing through the same anchor point are equal at the anchor point. Based on this condition, the smoothness of the contour curve at the connection point is ensured, thereby obtaining a smooth curve contour.

[0070] In addition, to ensure the effectiveness of the initial adversarial sample, the positions of the anchor points and control points are further constrained. Specifically, the constraints involve: the distance between the anchor point and the center point, the smoothness of the curve segment connection points, and the avoidance of contour self-intersection.

[0071] For example, in terms of the distance between the anchor point and the center point, Figure 5 As shown in Figure a, based on the radius r * The inner circle of the circle and the large circle with radius r constrain the feasible domain of the anchor point, that is, the range of the anchor point is inside the ring formed by the two circles; specifically, the distance function can be used Determine anchor point P i The position when r * ≤d i ≤r, then the anchor point P i The position of is reasonable, otherwise it is necessary to re-select the point. In the distance function, (x o ,y o ) represents the position coordinates of the center point, (x i ,yi ) represents the anchor point P i The location coordinates of .

[0072] In addition, this application uses segmented Bezier curves to form a closed adversarial sample outline to avoid the following Figure 5 In the case of non-smooth connection as shown in Figure b, the position of the control points also needs to be constrained. In terms of the smoothness of the curve segment connection points, the formula B i (t) = f B (P i ,D,E,P i+1 )=(1-t) 3 P i +3t(1-t) 2 E+3t 2 (1-t)D+t 3 P i+1 Take the derivative on both sides of t∈[0,1] to get the tangent representation of the Bessel function at the two endpoints; then when calculating the control point, ensure that B′ i (1) = B′ i+1 (0), which ensures that the slopes of the two Bezier curves at the connection point are equal, thereby ensuring the smoothness of the adversarial sample contour.

[0073] Since the present application first calculates the anchor points required to connect the contours, and then calculates two control points between two adjacent anchor points to control the shape of the Bezier curve, it is feasible to avoid self-intersection of the contours by limiting the control points. Figure 5 As shown in Figure c, when the control point moves from B to B', the Bezier curve will self-intersect. To solve this problem, this application uses the angle to determine the position of the control point. For example, when the control point is between two anchor points, the first line connecting the control point and the center point and the second line connecting the anchor point and the center point are first determined. The angle formed by the first and second lines must be smaller than the angle formed by the lines connecting the two anchor points and the center point. If this is not satisfied, it means that the control point will cause the Bezier curve to self-intersect, and the position of the control point should be recalculated.

[0074] Through the above embodiments, it can be found that the adversarial sample generation method of the multimodal fusion model disclosed in the present application uses the Bessel function and the interactive information of the multimodal images to generate the initial adversarial sample for the multimodal fusion model by the modal alternation selection algorithm, and under the guidance of the genetic algorithm, iteratively evolves to generate samples with better attack effects as the final optimal adversarial samples. The effect diagram of the optimal adversarial sample generated by the adversarial sample generation method of the multimodal fusion model of the present application is shown as follows: Figure 2 shown.

[0075] Figure 3FIG. 1 is a flow chart of a method for generating adversarial samples for a multimodal fusion model according to another embodiment of the present invention. Figure 3 As shown in FIG, the adversarial sample generation method specifically includes the following steps: step S01, image quality judgment: judging whether the initial fused image after fusion needs to be attacked; step S02, adversarial sample initialization: determining the center point of each initial adversarial sample according to the size of the model input image; step S03, modal alternation optimization: injecting the obtained initial adversarial sample into the single modal image to measure its attack effect on the single modal image; step S04, sample interaction and mutation: performing information interaction and mutation in the selected matrix to obtain sub-adversarial samples; step S05, boundary constraint: ensuring the effectiveness and smoothness of the adversarial sample; step S06, redundancy post-processing: selecting the optimal adversarial sample with the best attack effect from a large number of sub-adversarial samples.

[0076] In the above embodiment, the smooth contour generation algorithm based on the Bessel function forms a flexible and controllable contour representation based on the anchor point, and the shape of the adversarial sample contour can be easily modified by modifying the anchor point coordinates. In addition, when determining the adversarial sample contour, a constraint algorithm related to the position condition is adopted, considering the distance between the anchor point and the center point, the smoothness of the connection point, and the smoothness of the overall contour to comprehensively form a reasonable adversarial sample contour. This method first attacks on a single modal image, verifies the attack effect on the fused image, and finally selects samples that perform well on a single modality as the master sample, derives child adversarial samples from the master sample, and then obtains the final adversarial sample. This method accelerates learning.

[0077] Through the above embodiments, it can be found that the adversarial sample generation method of the multimodal fusion model disclosed in the present application utilizes texture information and adopts pixel perturbation or patch perturbation to perform adversarial attacks, which performs well under high visibility conditions. In addition, when obtaining adversarial samples, the adversarial sample generation method attacks the multimodal fusion model, that is, it attacks the information of multiple modal images at the same time, and uses Bessel functions to form a controllable and variable contour shape, and the representation method is simple and flexible. Compared with the existing technology, the present application can make full use of multimodal information to achieve end-to-end training, and can work in more complex environments, greatly broadening its application environment.

[0078] In addition, the adversarial sample generation method of the present application selects adversarial samples with excellent performance under the condition of attacking single-modal images, uses gene interaction and mutation in the selected adversarial samples, and adopts new sub-adversarial samples to attack multimodal fusion images, making good use of the interactive information between multiple modalities, so that the multimodal fusion model after adversarial training based on the adversarial samples has better robustness and security.

[0079] Corresponding to the above method, the present invention also provides an adversarial sample generation system for a multimodal fusion model, which includes a computer device, wherein the computer device includes a processor and a memory, wherein the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the system implements the steps of the method described above.

[0080] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.

[0081] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0082] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0083] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0084] The above merely illustrates the preferred embodiments of the present application, and is not used to limit the present application. The embodiments of the present application can be variously changed and modified by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall fall within the scope of protection of the present application.

Claims

1. A method for generating adversarial samples for a multimodal fusion model, characterized in that: The method comprises the following steps: Acquiring a multimodal image, and inputting the multimodal image into a multimodal fusion model to obtain an initial fused image, wherein the multimodal image includes multiple single-modal images; Determining a plurality of initial adversarial examples based on the multimodal image; Injecting each of the initial adversarial samples into each of the unimodal images to obtain each unimodal perturbation image, obtaining a semi-perturbation fusion image corresponding to each of the unimodal perturbation images through a multimodal fusion model based on each of the unimodal perturbation images and each of the unimodal images, and determining an optimal initial adversarial sample corresponding to each of the unimodal images based on each of the semi-perturbation fusion images and the initial fusion image; generating a plurality of sub-adversarial samples based on each of the optimal initial adversarial samples, and determining an optimal adversarial sample of the multimodal fusion model from the plurality of sub-adversarial samples; Determining the optimal initial adversarial sample corresponding to each of the single-modal images based on each of the semi-perturbed fused images and the initial fused image includes: Calculating the confidence difference between each of the semi-perturbed fused images and the initial fused image; Determining the initial adversarial sample size corresponding to each of the semi-perturbed fusion images and the degree of deviation of the perturbation region contour before and after the attack; Determining the optimal initial adversarial sample corresponding to each of the single-modal images based on the confidence difference, the size of the initial adversarial sample, and the degree of deviation of the perturbation region contour; Generating a plurality of sub-adversarial samples based on each of the optimal initial adversarial samples, and determining an optimal adversarial sample of the multimodal fusion model from the plurality of sub-adversarial samples, comprising: Select, mutate, and crossover based on each of the optimal initial adversarial samples to obtain multiple sub-adversarial samples; Injecting each sub-adversarial sample into the multimodal image to obtain a multimodal perturbation image corresponding to each sub-adversarial sample; Inputting the multimodal perturbation image into the multimodal fusion model to obtain a fully perturbation fusion image; An optimal adversarial sample of the multimodal fusion model is determined based on each of the fully perturbed fusion images and the initial fusion image.

2. The method for generating adversarial samples for a multimodal fusion model according to claim 1, wherein: Based on each of the single-modal perturbation images and each of the single-modal images, a semi-perturbation fusion image corresponding to each of the single-modal perturbation images is obtained through a multimodal fusion model, including: The first unimodal perturbation image and other second unimodal images in the multimodal image except the first unimodal image corresponding to the first unimodal perturbation image are input into the multimodal fusion model to obtain a semi-perturbation fusion image corresponding to the first unimodal perturbation image.

3. The method for generating adversarial samples for a multimodal fusion model according to claim 2, wherein: The calculation formula for the optimal initial adversarial sample is: ; in, represents the optimal initial adversarial sample, M represents the initial adversarial sample, represents the first single modality image, represents the second single modality image, represents the confidence difference, Indicates the degree of displacement of the disturbance area contour, represents the initial adversarial sample size, and Represents the coefficient.

4. The method for generating adversarial samples for a multimodal fusion model according to claim 1, wherein: Determining a plurality of initial adversarial examples based on the multimodal image includes: Determining, based on the size of the multimodal image, a plurality of sample center point coordinates and anchor point position coordinates corresponding to each of the sample center point coordinates; Based on the anchor point position coordinates corresponding to the center point coordinates of each sample, a sample contour curve corresponding to the center point coordinates of each sample is obtained by using a Bessel function; Each initial adversarial sample is determined based on each of the sample contour curves.

5. The method for generating adversarial samples for a multimodal fusion model according to claim 4, wherein: The slopes of the two curves passing through the same anchor point in each of the sample profile curves are the same.

6. The method for generating adversarial samples for a multimodal fusion model according to claim 5, wherein: The Bezier curve is: in, represents the i-th anchor point, represents the i+1th anchor point, and D and E represent two control points.

7. A system for generating adversarial samples for a multimodal fusion model, comprising a processor, a memory, and a computer program stored in the memory, characterized in that: The processor is configured to execute the computer program. When the computer program is executed, the system implements the steps of the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Multi-modal general adversarial sample generation method based on image fusion

    CN118015417A

  • Method and apparatus with multi-modal feature fusion

    US20230154170A1