A Method and System for Generating Cross-Modal Video Adversarial Examples

By performing feature extraction and image block processing on video samples, efficient and hidden video confrontation samples are generated, which solves the problems of low generation efficiency and poor concealment in the prior art, and improves the effect of cross-modal data retrieval.

CN115496966BActive Publication Date: 2025-07-29广州市省信软件有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211167515.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-07-29
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

When the prior art generates video confrontation samples across modalities, the generation efficiency is low and the disturbance is poor, making it difficult to meet users' needs for cross-modal data retrieval.

Method used

Convert clean video samples into a series of picture frames, perform feature extraction and keyframe determination, divide image blocks and calculate gradient scores, add perturbations to generate adversarial frames, and update perturbations through the Adam optimizer until the similarity is minimal, and replace keyframes to generate video adversarial samples.

Benefits of technology

The generation efficiency and concealment of video adversarial samples are improved, and the cross-modal transferability of adversarial samples is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496966B_ABST
    Figure CN115496966B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for generating cross-modal video adversarial samples, relating to the technical field of deep learning, including: obtaining clean video samples and converting them into a series of picture frames; extracting features from each picture frame to obtain corresponding feature vectors; determining key frames in the series of picture frames according to the feature vectors; dividing each key frame to obtain a series of image patch pictures and calculating the gradient score of each image patch picture; selecting the image patch picture with the largest gradient score as the local picture; adding perturbations to the local picture to obtain an adversarial frame and calculating the similarity between the local picture and the adversarial frame; updating the perturbations until the similarity reaches the minimum value, and taking the corresponding adversarial frame as the picture adversarial sample; using the picture adversarial sample to replace the corresponding key frame to obtain the video adversarial sample. The method and system for generating video adversarial samples in the present invention have high generation efficiency and strong concealment, and improve the cross-modal transferability of adversarial samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and more specifically, to a method and system for generating cross-modal video adversarial samples. Background Art

[0002] With the development of deep learning-related technologies, deep neural networks have made great progress in computer vision and natural language processing tasks; however, the state-of-the-art deep neural networks are vulnerable to adversarial samples. The form of adversarial attacks is that the attacker adds tiny perturbations that are difficult to distinguish manually to the original input data according to the algorithm to form adversarial samples. There are various methods for generating adversarial samples for adversarial attacks on images. The Fast Gradient Sign Method (FGSM) proposed by Goodfellow is a relatively classic method for generating adversarial samples. Aiming at the defect that the perturbation scale of the FGSM method is relatively large, the Iterative Fast Gradient Sign Method (I-FGSM) with a more restricted scale is proposed. By iteratively performing multiple small perturbations along the direction of increasing gradient and recalculating the gradient direction after each small step, more accurate perturbations can be constructed compared to FGSM, but the cost is an increase in the computational amount; in addition, there are also the Jacobian based Saliency Map Attack (JSMA), DeepFool, C&W, the Projected Gradient Descent (PGD) algorithm, etc. Compared with images, video recognition is a task with a wider distribution range and more application scenarios. The rapid growth of multi-modal data makes it difficult for users to effectively search for information of interest, thus giving rise to various retrieval and search technologies. In recent years, the popularization of mobile devices and emerging social networking sites has led to an increasing demand for cross-modal data retrieval from users. Therefore, the task of implementing adversarial attacks from images to videos is also very meaningful. The related work mainly focuses on improving the transferability of adversarial samples for video recognition models. Many existing works are based on gradient optimization. Since video samples are much more complex than images, it will cause the curse of dimensionality in videos, resulting in problems such as low generation efficiency and large perturbation concealment.

[0003] The prior art discloses a method for generating video adversarial samples based on target tracking and motion estimation, including: constructing an evaluation index for the continuity of video adversarial samples; using an object recognition algorithm to perform associative matching on the objects between two consecutive frames in a video sequence; obtaining relative motion data of the objects according to the associative matching results; calculating the model parameters of an adversarial sample generation model; optimizing the model parameters; using each frame image in the video sequence as the input of the final adversarial sample generation model corresponding to this frame image, and using the evaluation index for the continuity of video adversarial samples as a constraint condition to generate adversarial samples for each frame image. This method uses each frame image in the video sequence as the input of the final adversarial sample generation model corresponding to this frame image, with high iterative calculation complexity, low generation efficiency, and large perturbation concealment. Summary of the Invention

[0004] To overcome the defects of low generation efficiency and large perturbation concealment when generating video samples across modalities in the above prior art, the present invention provides a method and system for generating video adversarial samples across modalities.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] The present invention provides a method for generating video adversarial samples across modalities, including:

[0007] S1: Obtain a clean video sample and convert the clean video sample into a series of picture frames;

[0008] S2: Extract features from each picture frame to obtain corresponding feature vectors;

[0009] S3: Determine key frames in the series of picture frames according to the feature vectors;

[0010] S4: Divide each key frame to obtain a series of image block pictures, and calculate the gradient score of each image block picture; select the image block picture with the largest gradient score as the local picture;

[0011] S5: Add perturbations to the local picture to obtain an adversarial frame, and calculate the similarity between the local picture and the adversarial frame;

[0012] S6: Update the perturbations until the similarity reaches the minimum value, and use the corresponding adversarial frame as a picture adversarial sample;

[0013] S7: Use the picture adversarial sample to replace the corresponding key frame to obtain a video adversarial sample.

[0014] Preferably, in step S2, before extracting features from each picture frame, the color picture frame needs to be grayscale processed to obtain a grayscale picture frame.

[0015] Preferably, in the step S2, the specific method for extracting features from each picture frame to obtain the corresponding feature vector is as follows:

[0016] Using the texture-based feature extraction method, the gray-scaled picture frame is divided into N gray image blocks. For each gray image block, it is further divided into K*K small gray image blocks, and the mean and variance of each gray image block are calculated:

[0017]

[0018]

[0019] In the formula, μ l represents the mean of the l-th gray image block, represents the variance of the l-th gray image block, l = 1, 2,..., N; x l (u, v) represents the brightness value of the small gray image block in the u-th row and v-th column of the l-th gray image block, u, v = 0, 1,..., K - 1;

[0020] Combining the mean and standard deviation of each gray image block to obtain the feature vector of the picture frame:

[0021] A t = [μ1, σ1, μ2, σ2,..., μ N , σ N

[0022] In the formula, A t represents the feature vector of the t-th picture frame, t = 1, 2,..., M, M represents the number of picture frames; σ N represents the standard deviation of the N-th gray image block.

[0023] Preferably, in the step S3, the specific method for determining the key frames from a series of picture frames according to the feature vectors is as follows:

[0024] S3.1: Select several picture frames from a series of picture frames as the data centers, calculate the similarity between the data centers and the remaining picture frames according to the feature vectors, and group the picture frames according to the similarity;

[0025] S3.2: Update the data centers multiple times, group the picture frames until the grouping no longer changes, obtain the final grouping, and use the picture frames corresponding to the data centers in each group of the final grouping as the key frames.

[0026] Preferably, in the step S3.1, the specific method for selecting several picture frames from a series of picture frames as the data centers, calculating the similarity between the data centers and the remaining picture frames according to the feature vectors, and grouping the picture frames according to the similarity is as follows: ​

[0027] Select m picture frames from a series of picture frames as data centers, and calculate the similarity between the data centers and the remaining picture frames according to the feature vectors:

[0028]

[0029] In the formula, similarity j,k represents the similarity between the j-th picture frame and the k-th picture frame, and A j represents the feature vector of the j-th picture frame in the data center, and A k represents the feature vector of the k-th picture frame in the remaining picture frames; n represents the dimension of the feature vector, represents the i-th dimension of the feature vector of the j-th picture frame in the data center;

[0030] Use the above formula to calculate the similarity of each of the remaining picture frames relative to each picture frame selected as the data center, and divide each of the remaining picture frames into the group where the data center with the smallest similarity value is located to obtain p groups, where p ≤ m.

[0031] Preferably, in step S4, the specific method for dividing each key frame to obtain a series of image block pictures and calculating the gradient score of each image block picture is as follows:

[0032] Denote the key frame as U and divide it into a image blocks; perform zero padding on each image block to make the size of the image block the same as that of the key frame to obtain an image block picture:

[0033]

[0034] In the formula, represents the a-th image block picture, G(*) represents the segmentation operation, and P(*) represents the padding operation;

[0035]

[0036] In the formula, score a represents the gradient score of the a-th image block picture, represents the gradient of the a-th image block picture, represents the score calculation operation.

[0037] Preferably, in step S5, the specific method for adding perturbations to the local picture to obtain an adversarial frame and calculating the similarity between the local picture and the adversarial frame is as follows:

[0038] S5.1: Extract the feature vector of the local picture to obtain the local picture feature vector;

[0039] S5.2: Add perturbations to the local image to obtain adversarial frames and adversarial feature vectors;

[0040] S5.3: Calculate the similarity between the local image and the adversarial frame based on the local image feature vector and the adversarial feature vector.

[0041] Preferably, in the step S5.3, the specific method for calculating the similarity between the local image and the adversarial frame based on the local image feature vector and the adversarial feature vector is as follows:

[0042]

[0043] In the formula, represents the similarity between the a-th local image and the adversarial frame, x a represents the a-th local image feature vector, (x a +ε) represents the a-th adversarial feature vector, and ε a represents the perturbation added to the a-th local image.

[0044] Preferably, in the step S6, the Adam optimizer is used to update the perturbation:

[0045]

[0046] In the formula, represents the perturbation added to the a-th local image in the (t + 1)-th iteration, Adam represents the Adam optimizer, represents the perturbation added to the a-th local image in the t-th iteration, and θ represents the Adam optimizer parameter.

[0047] The present invention also provides a cross-modal video adversarial sample generation system for implementing the above cross-modal video adversarial sample generation method. The system includes:

[0048] A video acquisition and processing module for acquiring a clean video sample and converting the clean video sample into a series of picture frames;

[0049] A feature extraction module for extracting features from each picture frame to obtain corresponding feature vectors;

[0050] A key frame selection module for determining key frames in a series of picture frames according to the feature vectors;

[0051] A local image generation module for dividing each key frame to obtain a series of image block pictures, and calculating the gradient score of each image block picture; selecting the image block picture with the largest gradient score as the local image;

[0052] An adversarial frame generation module, which is used to add perturbations to local pictures to obtain adversarial frames and calculate the similarity between the local pictures and the adversarial frames;

[0053] A perturbation update module, which is used to update the perturbations until the similarity reaches the minimum value, and use the corresponding adversarial frame as the picture adversarial sample;

[0054] A video adversarial sample generation module, which uses the picture adversarial sample to replace the corresponding key frame to obtain the video adversarial sample.

[0055] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0056] The present invention first obtains a clean video sample and converts the modality into a series of picture frames; then extracts features from each picture frame, determines the key frames according to the obtained feature vectors; divides the key frames to obtain a series of image block pictures, and screens out local pictures from them; adds perturbations to the local images to obtain adversarial frames; by updating the perturbations, the similarity between the local pictures and the adversarial frames reaches the minimum value, and the corresponding adversarial frame is used as the picture adversarial sample; finally, the picture adversarial sample is used to replace the corresponding key frame to obtain the video adversarial sample. The present invention generates perturbations for the local images of the key frames, and the generation efficiency of the video adversarial sample is high and the concealment is good, improving the cross-modal transferability of the adversarial sample. Description of the Drawings

[0057] Figure 1 It is a flowchart of a method for generating cross-modal video adversarial samples described in Embodiment 1.

[0058] Figure 2 It is a structural schematic diagram of a system for generating cross-modal video adversarial samples described in Embodiment 3. Detailed Embodiments

[0059] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0060] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the dimensions of the actual product;

[0061] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0062] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0063] Embodiment 1

[0064] This embodiment provides a method for generating cross-modal video adversarial samples, as Figure 1 shown, including:

[0065] S1: Obtain a clean video sample and convert the clean video sample into a series of picture frames;

[0066] S2: Extract features from each picture frame to obtain the corresponding feature vectors;

[0067] S3: Determine key frames from the series of picture frames according to the feature vectors;

[0068] S4: Divide each key frame to obtain a series of image patch pictures, and calculate the gradient score of each image patch picture; Select the image patch picture with the maximum gradient score as the local picture;

[0069] S5: Add perturbations to the local picture to obtain adversarial frames, and calculate the similarity between the local picture and the adversarial frames;

[0070] S6: Update the perturbations until the similarity reaches the minimum value, and use the corresponding adversarial frames as picture adversarial samples;

[0071] S7: Replace the corresponding key frames with the picture adversarial samples to obtain video adversarial samples.

[0072] In the specific implementation process, this embodiment first obtains a clean video sample and converts it into a series of picture frames through modality conversion; then extracts features from each picture frame, and determines key frames according to the obtained feature vectors; divides the key frames to obtain a series of image patch pictures, and selects local pictures from them; adds perturbations to the local images to obtain adversarial frames; by updating the perturbations, the similarity between the local picture and the adversarial frames reaches the minimum value, and the corresponding adversarial frames are used as picture adversarial samples; finally, the corresponding key frames are replaced with the picture adversarial samples to obtain video adversarial samples. This embodiment has high generation efficiency and strong concealment for generating video adversarial samples, and improves the cross-modal transferability of adversarial samples.

[0073] Embodiment 2

[0074] The embodiment provides a method for generating cross-modal video adversarial samples, including:

[0075] S1: Obtain a clean video sample and convert the clean video sample into a series of picture frames;

[0076] S2: Perform grayscale processing on the color picture frames to obtain grayscale picture frames; Extract features from each grayscale picture frame to obtain the corresponding feature vectors; Specifically:

[0077] Using a texture-based feature extraction method, divide the grayscale picture frames into N grayscale image patches, and for each grayscale image patch, further divide it into K*K small grayscale image patches, and calculate the mean and variance of each grayscale image patch:

[0078]

[0079]

[0080] where μ l represents the mean value of the l-th grayscale image block, represents the variance of the l-th grayscale image block, l = 1, 2, …, N; x l (u, v) represents the luminance value of the small grayscale image block at the u-th row and v-th column in the l-th grayscale image block, u, v = 0, 1, …, K - 1;

[0081] Combining the mean value and standard deviation of each grayscale image block to obtain the feature vector of the picture frame:

[0082] A t = [μ1, σ1, μ2, σ2, …, μ N , σ N

[0083] where A t represents the feature vector of the t-th picture frame, t = 1, 2, …, M, M represents the number of picture frames; σ N represents the standard deviation of the N-th grayscale image block.

[0084] S3: Determine key frames in a series of picture frames according to the feature vector; specifically:

[0085] S3.1: Select several picture frames in a series of picture frames as data centers, calculate the similarity between the data centers and the remaining picture frames according to the feature vector, and group the picture frames according to the similarity; specifically:

[0086] Select m picture frames in a series of picture frames as data centers, and calculate the similarity between the data centers and the remaining picture frames according to the feature vector:

[0087]

[0088] where similarity j,k represents the similarity between the j-th picture frame and the k-th picture frame, A j represents the feature vector of the j-th picture frame in the data center, A k represents the feature vector of the k-th picture frame in the remaining picture frames; n represents the dimension of the feature vector, represents the i-th dimension of the feature vector of the j-th picture frame in the data center;

[0089] ​Calculate the similarity of each remaining picture frame relative to each picture frame selected as the data center using the above formula, and divide each remaining picture frame into the group where the data center with the smallest similarity value is located, obtaining p groups, where p ≤ m;

[0090] S3.2: Update the data center multiple times, group the picture frames until the grouping no longer changes, obtain the final grouping, and use the picture frames corresponding to the data centers in each group in the final grouping as key frames.

[0091] S4: Divide each key frame to obtain a series of image patch pictures, and calculate the gradient score of each image patch picture; select the image patch picture with the largest gradient score as the local picture; specifically:

[0092] Denote the key frame as U and divide it into a image patches; perform zero-padding on each image patch to make the size of the image patch consistent with the size of the key frame, obtaining the image patch pictures:

[0093]

[0094] In the formula, represents the a-th image patch picture, G(*) represents the segmentation operation, and P(*) represents the padding operation;

[0095]

[0096] In the formula, score a represents the gradient score of the a-th image patch picture, represents the gradient of the a-th image patch picture, represents the score calculation operation.

[0097] S5: Add perturbations to the local picture to obtain an adversarial frame, and calculate the similarity between the local picture and the adversarial frame; specifically:

[0098] S5.1: Extract features from the local picture to obtain the local picture feature vector;

[0099] S5.2: Add perturbations to the local picture to obtain the adversarial frame and the adversarial feature vector;

[0100] S5.3: Calculate the similarity between the local picture and the adversarial frame according to the local picture feature vector and the adversarial feature vector; specifically:

[0101]

[0102] In the formula, represents the similarity between the a-th local picture and the adversarial frame, x a represents the a-th local picture feature vector, (x aThe (a)-th adversarial feature vector is denoted as (+ε), where ε a represents the perturbation added to the a-th local image.

[0103] S6: Update the perturbation using the Adam optimizer until the similarity reaches the minimum value, and use the corresponding adversarial frame as the image adversarial sample;

[0104] The old method for updating the perturbation using the Adam optimizer is:

[0105]

[0106] In the formula, represents the perturbation added to the a-th local image at the (t + 1)-th iteration, Adam represents the Adam optimizer, represents the perturbation added to the a-th local image at the t-th iteration, and θ represents the parameters of the Adam optimizer.

[0107] S7: Replace the corresponding key frame with the image adversarial sample to obtain the video adversarial sample.

[0108] Embodiment 3

[0109] This embodiment provides a cross-modal video adversarial sample generation system for implementing the cross-modal video adversarial sample generation method described in Embodiment 1 or 2. As Figure 2 shown, the system includes:

[0110] A video acquisition and processing module, configured to acquire a clean video sample and convert the clean video sample into a series of picture frames;

[0111] A feature extraction module, configured to extract features from each picture frame to obtain the corresponding feature vector;

[0112] A key frame selection module, configured to determine key frames from a series of picture frames according to the feature vectors;

[0113] A local image generation module, configured to divide each key frame to obtain a series of patch images, and calculate the gradient score of each patch image; select the patch image with the largest gradient score as the local image;

[0114] An adversarial frame generation module, configured to add a perturbation to the local image to obtain an adversarial frame, and calculate the similarity between the local image and the adversarial frame;

[0115] A perturbation update module, configured to update the perturbation until the similarity reaches the minimum value, and use the corresponding adversarial frame as the image adversarial sample;

[0116] A video adversarial sample generation module, which uses the image adversarial sample to replace the corresponding key frame to obtain the video adversarial sample.

[0117] Like or similar reference numerals correspond to like or similar components;

[0118] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;

[0119] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for generating cross-modal video adversarial samples, characterized in that, Including: S1: Obtain a clean video sample and convert the clean video sample into a series of picture frames; S2: Extract features from each picture frame to obtain corresponding feature vectors; S3: Determine key frames in the series of picture frames according to the feature vectors; S4: Divide each key frame to obtain a series of image patch pictures, and calculate the gradient score of each image patch picture; Select the image patch picture with the largest gradient score as the local picture; The specific method for dividing each key frame to obtain a series of image patch pictures and calculating the gradient score of each image patch picture is: Denote the key frame as U and divide it into a image patches; Zero-pad each image patch to make the size of the image patch consistent with the size of the key frame to obtain an image patch picture: Where, represents the a-th image block, G(*) represents the segmentation operation, and P(*) represents the filling operation; In the formula, score a Represents the gradient score of the a-th image block, Represents the gradient of the a-th image block, Represents a fraction calculation operation; S5: Add perturbations to the local picture to obtain an adversarial frame, and calculate the similarity between the local picture and the adversarial frame. The specific method is: S5.1: Extract features from the local picture to obtain a local picture feature vector; S5.2: Add perturbations to the local picture to obtain an adversarial frame and an adversarial feature vector; S5.3: Calculate the similarity between the local picture and the adversarial frame according to the local picture feature vector and the adversarial feature vector. The specific method is: In the formula, represents the similarity between the a-th local image and the adversarial frame, x a represents the feature vector of the a-th local image, (x a +ε) represents the a-th adversarial feature vector, and ε a represents the perturbation added to the a-th local image; S6: Update the perturbations until the similarity reaches the minimum value, and use the corresponding adversarial frame as the picture adversarial sample; Update the perturbations using the Adam optimizer: wherein, represents the perturbation added to the a-th local image at the (t + 1)-th iteration, and Adam represents the Adam optimizer, represents the perturbation added to the a-th local image at the t-th iteration, and θ represents the parameters of the Adam optimizer; S7: Replace the corresponding key frame with the picture adversarial sample to obtain the video adversarial sample.

2. The method for generating cross-modal video adversarial samples according to claim 1, wherein In step S2, before extracting features from each picture frame, the color picture frame needs to be grayscaled to obtain the grayscaled picture frame.

3. The method for generating cross-modal video adversarial samples according to claim 2, wherein In step S2, the specific method for extracting features from each picture frame to obtain corresponding feature vectors is: Using a texture-based feature extraction method, divide the grayscaled picture frame into N grayscale image patches. For each grayscale image patch, further divide it into K*K small grayscale image patches, and calculate the mean and variance of each grayscale image patch: where μl represents the mean of the l-th grayscale image block, represents the variance of the l-th grayscale image block, l = 1, 2, …, N; x l (u, v) represents the luminance value of the small grayscale image block at the u-th row and v-th column in the l-th grayscale image block, u, v = 0, 1, …, K - 1; Combine the mean and standard deviation of each grayscale image patch to obtain the feature vector of the picture frame: A t =[μ1,σ1,μ2,σ2,…,μ N ,s N ] Where A t represents the feature vector of the t-th image frame, t=1,2,…,M, where M represents the number of image frames; σ N Represents the standard deviation of the Nth grayscale image block.

4. The method for generating cross-modal video adversarial samples according to claim 1, wherein In step S3, the specific method for determining key frames in the series of picture frames according to the feature vectors is: S3.1: Select several picture frames in the series of picture frames as data centers, calculate the similarity between the data centers and the remaining picture frames according to the feature vectors, and group the picture frames according to the similarity; S3.2: Update the data centers multiple times and group the picture frames until the grouping no longer changes to obtain the final grouping. Use the picture frames corresponding to the data centers in each group of the final grouping as key frames.

5. The method for generating cross-modal video adversarial samples according to claim 4, characterized in that: In step S3.1, the specific method for selecting several picture frames in the series of picture frames as data centers, calculating the similarity between the data centers and the remaining picture frames according to the feature vectors, and grouping the picture frames according to the similarity is: Select m picture frames in the series of picture frames as data centers, and calculate the similarity between the data centers and the remaining picture frames according to the feature vectors: Where similatity j,k Indicates the similarity between the j-th picture frame and the k-th picture frame, A j Represents the feature vector of the jth image frame in the data center, A k represents the feature vector of the kth picture frame in the remaining picture frames; n represents the dimension of the feature vector, The i-th dimension of the feature vector representing the j-th image frame in the data center; Calculate the similarity of each remaining picture frame relative to each picture frame selected as the data center using the above formula, and divide each remaining picture frame into the group where the data center with the smallest similarity value is located to obtain p groups, where p ≤ m.

6. A system for generating cross-modal video adversarial samples, used to implement the method for generating cross-modal video adversarial samples according to any one of claims 1 to 5, characterized in that: The system includes: A video acquisition and processing module, configured to acquire a clean video sample and convert the clean video sample into a series of picture frames; A feature extraction module, configured to extract features from each picture frame to obtain a corresponding feature vector; A key frame selection module, configured to determine key frames in a series of picture frames according to the feature vectors; A local picture generation module, configured to divide each key frame to obtain a series of image block pictures and calculate the gradient score of each image block picture; select the image block picture with the largest gradient score as the local picture; An adversarial frame generation module, configured to add perturbations to the local picture to obtain an adversarial frame and calculate the similarity between the local picture and the adversarial frame; A perturbation update module, configured to update the perturbation until the similarity reaches the minimum value, and use the corresponding adversarial frame as a picture adversarial sample; A video adversarial sample generation module, configured to replace the corresponding key frame with the picture adversarial sample to obtain a video adversarial sample.

Citation Information

Patent Citations

  • Cross-modal high-concealment confrontation sample generation method and system oriented to underwater sound intelligent camouflage

    CN115081510A

  • Learning to rank with cross-modal graph convolutions

    EP3896581A1