Text video retrieval system confrontation sample generation method and related device

By building a greedy and cautious attack sample set, optimizing the anti-perturbation to maximize the similarity or similarity ratio, the generated adversarial samples realize hidden and effective attacks in the text video retrieval system, improving the security and defense capabilities of the system, and solving the problems of difficulty in deploying attack methods and unrealistic targets in the existing technology.

CN120429465APending Publication Date: 2025-08-05XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510531235.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The adversarial sample generation method of the existing text video retrieval system is difficult to implement in real scenarios, the attack method is difficult to deploy, and the attack target does not meet actual needs, resulting in limited improvement in system defense.

Method used

By constructing a greedy and cautious attack sample set, the encoder of the text video retrieval system extracts features, optimizes the anti-perturbation to maximize similarity or similarity ratio, generates greedy and cautious adversarial samples, simulates real user query behavior, and adjusts the perturbation video to promote or push specific query samples on multiple query texts.

Benefits of technology

The generated adversarial samples are insensitive on the ordinary user side, with significant attack effects, meeting the needs of actual attack scenarios, and improving the security and defense capabilities of the text video retrieval system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429465A_ABST
    Figure CN120429465A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of text video retrieval, and discloses a text video retrieval system confrontation sample generation method and related device.The text video retrieval system confrontation sample generation method comprises the steps that a text video retrieval system is inquired based on text descriptions of candidate videos, a recommended video list is obtained, text descriptions of all videos in the recommended video list are obtained, and a plurality of text samples are obtained; constructing a greedy attack sample set according to the plurality of text samples, and extracting text features of the greedy attack sample set by adopting a text encoder which is the same as the text video retrieval system; and repeating the first optimization step to preset repetition times to obtain final greedy adversarial disturbance, and obtaining a greedy disturbance video of the candidate video based on the final greedy adversarial disturbance as a greedy attack adversarial sample. The confrontation sample of the text video retrieval system with aggressiveness conforming to a real application scene and excellent attack ability is obtained, so that an optimization basis is provided for security protection of the text video retrieval system, and the ability of the text video retrieval system to resist the confrontation attack is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of text video retrieval, and relates to a method for generating adversarial samples in a text video retrieval system and related devices. Background Art

[0002] The task of text-video retrieval involves, given a query text, searching a database for all videos similar to the query text and returning one or more videos similar to the query text. With the continuous advancement of information retrieval technology, the search modality of information retrieval has evolved from traditional single-modal retrieval, such as image-to-image and text-to-text, to cross-modal retrieval technologies, such as text-to-audio and text-to-video, which are attracting increasing attention. As one of the primary methods for information retrieval, text-video retrieval technology is also widely used. Furthermore, with the continuous improvement of large-scale model multimodal alignment capabilities, text-video retrieval technology will continue to be applied and implemented, potentially replacing text retrieval in the information retrieval field in the future. The principle of text-video retrieval technology is to leverage the encoding capabilities of large-scale language-vision models to encode inputs in both text and video modalities and extract their features. Subsequently, through specialized multimodal alignment techniques, such as image-text similarity scores and cross-modal alignment modules, the corresponding text and video pairs are aligned, achieving alignment across different modalities for the same sample. In other words, similar to unimodal retrieval tasks, cross-modal retrieval also relies on training models to learn an embedding space in which similar homomodal and heteromodal samples are close together, while dissimilar homomodal and heteromodal samples are farther apart. However, with the discovery of adversarial attacks, the fragility and robustness of large language and vision models have gradually become apparent. Adversarial attacks typically involve attackers adding tiny perturbations to visual samples that are imperceptible to the human eye or replacing individual words in text samples, causing deep neural networks to make significant errors in their judgment of the sample.

[0003] Existing adversarial attacks against cross-modal retrieval systems mainly perturb specific samples to change their distribution position in the embedding space learned by the model, making them further away from similar samples, thereby reducing the recall rate of the model. Existing research can be divided into multimodal perturbation adversarial attacks and unimodal perturbation adversarial attacks based on the perturbation modality. Taking the text-video retrieval system as an example, a multimodal adversarial attack is to simultaneously perturb the query text and candidate videos to reduce the image-text similarity scores between the two, making it impossible to recall the correct video samples, thereby reducing the recall rate of the model. A unimodal adversarial attack is to perturb the query text to move it away from the candidate videos being queried, thereby making it impossible to recall the correct video samples, thereby reducing the recall rate of the model. Currently, there is little research on adversarial attacks against text-video retrieval systems. Existing work mainly uses the above-mentioned unimodal adversarial attacks on query text to achieve ranking of candidate videos and reduce the recall rate.

[0004] However, existing adversarial sample generation methods for text-to-video retrieval systems are difficult to implement in real-world text-to-video query scenarios, and their attack targets are inconsistent with real-world attack scenarios. Furthermore, these attack methods are difficult to implement and their targets are unrealistic. Specifically, single-modal attacks targeting query text require the attacker to tamper with or replace the user's query text content to achieve the attack. This makes such attacks difficult to deploy in real-world user retrieval scenarios, resulting in difficult-to-implement attack methods and limited usefulness of such adversarial samples. Furthermore, such attacks aim to lower the ranking of candidate videos in retrieval results, preventing the correct samples from being properly recalled. However, for text-to-video retrieval systems, in order to gain video clicks, thereby disseminating information and generating illicit profits, real-world attackers are more likely to rank incorrect samples higher, significantly increasing the probability of perturbed videos appearing in retrieval results and using adversarial tactics to achieve video promotion. Therefore, adversarial samples generated by existing methods fail to reflect the attacker's true intentions, and their attack objectives are unrealistic. Consequently, it is impossible to improve the performance and defense capabilities of text-to-video retrieval systems based on such adversarial samples. Summary of the Invention

[0005] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a method for generating adversarial samples for a text video retrieval system and related devices.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] In a first aspect, the present invention provides a method for generating adversarial samples for a text video retrieval system, comprising: querying a text video retrieval system based on a text description of a candidate video, obtaining a list of recommended videos and obtaining text descriptions of each video in the recommended video list, and obtaining a number of text samples; constructing a greedy attack sample set based on the number of text samples, and extracting text features of the greedy attack sample set using the same text encoder as the text video retrieval system; repeating the first optimization step to a preset number of repetitions to obtain a final greedy adversarial perturbation, and obtaining a greedy perturbation video of the candidate video based on the final greedy adversarial perturbation as a greedy attack adversarial sample; wherein the first optimization step comprises: obtaining a greedy perturbation video of the candidate video based on the current greedy adversarial perturbation, and extracting video features of the greedy perturbation video using the same video encoder as the text video retrieval system; optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set.

[0008] Optionally, when optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set, the greedy attack optimization objective function is adopted:

[0009]

[0010] st||δ g || p <∈

[0011] in, Video features for greedy perturbation videos Text features of greedy attack sample sets The similarity between them, Sim(.) is the similarity function; δ g To fight against disturbances for greed, ‖·‖ p for l p norm, ∈ is the disturbance limit parameter.

[0012] Optionally, when optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set, the greedy adversarial perturbation is optimized in the following manner: obtaining the gradient of the greedy attack optimization objective function, and projecting the gradient of the greedy attack optimization objective function to a hypersphere space of size ∈, obtaining the projected gradient and optimizing the greedy adversarial perturbation based on the projected gradient.

[0013] Optionally, it also includes: constructing a cautious attack sample set based on a number of text samples; wherein the cautious attack sample set includes a target text sample set and a risk text sample set, which are text samples corresponding to the first preset number of videos ranked before and the second preset number of videos ranked after in the recommended video list respectively; using the same text encoder as the text video retrieval system to extract text features of the target text sample set and text features of the risk text sample set; repeating the second optimization step to a preset number of repetitions to obtain the final cautious adversarial perturbation, and obtaining a cautious perturbation video of the candidate video based on the final cautious adversarial perturbation as a cautious attack adversarial sample: wherein the second optimization step includes: obtaining a cautious perturbation video of the candidate video based on the current cautious adversarial perturbation, and using the same video encoder as the text video retrieval system to extract video features of the cautious perturbation video; optimizing the cautious adversarial perturbation with the maximization of the ratio of the first cautious similarity to the second cautious similarity as the optimization target; wherein the first cautious similarity is the similarity between the video features of the cautious perturbation video and the text features of the target text sample set, and the second cautious similarity is the similarity between the video features of the cautious perturbation video and the text features of the risk text sample set.

[0014] Optionally, when optimizing the cautious counter-perturbation with maximizing the ratio of the first cautious similarity to the second cautious similarity as the optimization goal, a cautious attack optimization objective function is adopted:

[0015]

[0016] st‖δ S ‖ p <∈

[0017] in, For the first cautious similarity, is the second cautious similarity, Sim(.) is the similarity function, To carefully perturb the video features of the video, is the text feature of the target text sample set, is the text feature of the risk text sample set; δ s To be cautious against disturbances, p for l p norm, ∈ is the disturbance limit parameter.

[0018] Optionally, when optimizing the cautious adversarial perturbation with maximizing the ratio of the first cautious similarity to the second cautious similarity as the optimization target, the cautious adversarial perturbation is optimized in the following manner: obtaining the gradient of the cautious attack optimization objective function, and projecting the gradient of the cautious attack optimization objective function to a hypersphere space of size ∈, obtaining the projected gradient and optimizing the cautious adversarial perturbation based on the projected gradient.

[0019] Optionally, when constructing the greedy attack sample set based on the plurality of text samples, the text samples of different candidate videos are shared; when constructing the cautious attack sample set based on the plurality of text samples, the text samples of different candidate videos are not shared.

[0020] According to a second aspect of the present invention, a text video retrieval system adversarial sample generation system is provided, comprising: a query module for querying the text video retrieval system based on the text description of a candidate video, obtaining a list of recommended videos and obtaining the text description of each video in the recommended video list, and obtaining a number of text samples; a feature module for constructing a greedy attack sample set based on the number of text samples, and extracting text features of the greedy attack sample set using the same text encoder as the text video retrieval system; a sample generation module for repeating the first optimization step to a preset number of repetitions to obtain a final greedy adversarial perturbation, and obtaining a greedy perturbation video of the candidate video based on the final greedy adversarial perturbation as a greedy attack adversarial sample; wherein the first optimization step comprises: obtaining a greedy perturbation video of the candidate video based on the current greedy adversarial perturbation, and extracting video features of the greedy perturbation video using the same video encoder as the text video retrieval system; and optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set.

[0021] In a third aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the adversarial sample generation method of the above-mentioned text-video retrieval system are implemented.

[0022] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the adversarial sample generation method of the above-mentioned text video retrieval system are implemented.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] The present invention proposes an innovative adversarial sample generation method for a text-based video retrieval system, not only proposing an innovative attack sample construction strategy but also deeply optimizing the feasibility and pertinence of adversarial attacks in practical applications. Specifically, the method first queries the system using the text descriptions of candidate videos and obtains a list of recommended videos and their text descriptions. This step cleverly simulates the behavioral patterns of real users querying videos, providing a rich and authentic text sample foundation for subsequent adversarial sample generation. Next, by constructing a greedy attack sample set and extracting text features, the present invention sets a clear target direction for optimizing the adversarial perturbation. During the optimization process, the present invention repeatedly executes the first optimization step, continuously adjusting the greedy adversarial perturbation to maximize the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set. This greedy strategy not only ensures that the adversarial sample achieves good generalization effects on multiple query texts, but also achieves a balance between attack efficiency and attack effectiveness by controlling the preset number of repetitions. More importantly, the present invention proposes a practical solution to the problems of difficult-to-implement attack methods and unrealistic attack targets in existing adversarial sample generation methods. By injecting disturbed videos, adversarial attacks are rendered almost imperceptible to ordinary users, thereby achieving covert and effective adversarial sample generation. At the same time, the goal of traditional adversarial attacks is changed from simply pushing away specific query samples to pushing toward specific query samples. This shift not only better meets the needs of actual attack scenarios, but also provides a new approach and method for adversarial sample generation. Ultimately, the present invention obtains adversarial samples for text-video retrieval systems whose aggressiveness matches real-world application scenarios and whose attack capabilities are excellent. This provides an optimized foundation for the security protection of text-video retrieval systems, enhances the ability of text-video retrieval systems to resist adversarial attacks, and ensures the safe operation of text-video retrieval systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flow chart of the adversarial sample generation method for the text video retrieval system according to an embodiment of the present invention.

[0026] Figure 2 Schematic diagram of the adversarial sample generation principle of the text video retrieval system according to an embodiment of the present invention.

[0027] Figure 3 This is a structural block diagram of the adversarial sample generation system of the text video retrieval system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] The present invention is described in further detail below with reference to the accompanying drawings:

[0031] See also Figure 1 In one embodiment of the present invention, a method for generating adversarial samples for a text-video retrieval system is provided, which can generate adversarial samples for a text-video retrieval system that conform to real application scenarios and have excellent attack capabilities, thereby providing a data basis for improving the ability of the text-video retrieval system to resist adversarial attacks.

[0032] Specifically, the adversarial sample generation method of the text video retrieval system of the present invention includes the following steps:

[0033] S1: Query the text video retrieval system based on the text description of the candidate video to obtain a list of recommended videos and obtain the text description of each video in the recommended video list to obtain several text samples;

[0034] S2: Construct a greedy attack sample set based on several text samples, and use the same text encoder as the text video retrieval system to extract text features of the greedy attack sample set;

[0035] S3: Repeat the first optimization step to a preset number of repetitions to obtain the final greedy adversarial perturbation, and obtain the greedy perturbation video of the candidate video based on the final greedy adversarial perturbation as the greedy attack adversarial sample.

[0036] Among them, the first optimization step includes: obtaining a greedy perturbation video of the candidate video based on the current greedy adversarial perturbation, and using the same video encoder as the text video retrieval system to extract the video features of the greedy perturbation video; optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set.

[0037] The present invention proposes an innovative adversarial sample generation method for a text-based video retrieval system, not only proposing an innovative attack sample construction strategy but also deeply optimizing the feasibility and pertinence of adversarial attacks in practical applications. Specifically, the method first queries the system using the text descriptions of candidate videos and obtains a list of recommended videos and their text descriptions. This step cleverly simulates the behavioral patterns of real users querying videos, providing a rich and authentic text sample foundation for subsequent adversarial sample generation. Next, by constructing a greedy attack sample set and extracting text features, the present invention sets a clear target direction for optimizing the adversarial perturbation. During the optimization process, the present invention repeatedly executes the first optimization step, continuously adjusting the greedy adversarial perturbation to maximize the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set. This greedy strategy not only ensures that the adversarial sample achieves good generalization effects on multiple query texts, but also achieves a balance between attack efficiency and attack effectiveness by controlling the preset number of repetitions. More importantly, the present invention proposes a practical solution to the problems of difficult-to-implement attack methods and unrealistic attack targets in existing adversarial sample generation methods. By injecting disturbed videos, adversarial attacks are rendered almost imperceptible to ordinary users, thereby achieving covert and effective adversarial sample generation. At the same time, the goal of traditional adversarial attacks is changed from simply pushing away specific query samples to pushing toward specific query samples. This shift not only better meets the needs of actual attack scenarios, but also provides a new approach and method for adversarial sample generation. Ultimately, the present invention obtains adversarial samples for text-video retrieval systems whose aggressiveness matches real-world application scenarios and whose attack capabilities are excellent. This provides an optimized foundation for the security protection of text-video retrieval systems, enhances the ability of text-video retrieval systems to resist adversarial attacks, and ensures the safe operation of text-video retrieval systems.

[0038] Exemplarily, the specific process of querying the text video retrieval system based on the text description of the candidate video includes: using the black box query method, using the text description corresponding to the candidate video as the query text, querying the text video retrieval system, obtaining the recommended video list returned by the text video retrieval system, and then using the text descriptions of all videos in the recommended video list as text samples.

[0039] Optionally, based on the required number of text samples, the above black box query steps can be repeated by replacing the candidate videos until the required number of text samples is reached.

[0040] Explanatory,greedy attack aims to make the attacked candidate video close to as many query texts as possible at the same time,that is, to improve the ranking in as many text queries as possible,that is, to achieve the promotion of the attacked candidate video.

[0041] For example, when constructing a greedy attack sample set based on a number of text samples, a certain number of text samples can be randomly selected from the plurality of text samples according to the quantity requirement. As diverse text samples as possible can be selected to approximate the true distribution of the text dataset for the text-to-video retrieval system. Furthermore, by optimizing the perturbation, the greedy attack adversarial samples are made as close as possible to the entire greedy attack sample set to ensure adversarial performance.

[0042] Optionally, when constructing the greedy attack sample set, a greedy attack sample test set can be constructed at the same time to test the effect of the final greedy attack adversarial sample. For example, the number of text samples is set to 300, and 100 text samples are randomly selected as the greedy attack sample set, denoted as T train The remaining 200 text samples are used as the greedy attack sample test set, denoted as T test .

[0043] For example, when constructing the greedy attack sample set, text samples of different candidate videos are shared with each other, that is, all candidate videos can be queried and extracted in a unified manner.

[0044] For example, when the text features of the greedy attack sample set are extracted using the same text encoder as the text video retrieval system, the same text encoder E as the text encoder of the text video retrieval system is used. T As the basis of feature extraction, the text samples of the greedy attack sample set are input into the text encoder E T , get the text features of the greedy attack sample set, recorded as Among them, L is the size of the greedy attack sample set, M is the maximum number of words, and D is the feature dimension.

[0045] For example, when the same video encoder as the text video retrieval system is used to extract the video features of the greedy perturbation video, the same video encoder E as the video encoder of the text video retrieval system is used. I As the basis of feature extraction, for a given greedy perturbation video V+δ g , uniformly extract N frames as key frames, namely V i ={v0,v1,…,v N}; Then the video V i,raw Input to video encoder E I , we get the video features of the greedy perturbation video, recorded as Among them, δ g Fighting disturbance for greed.

[0046] In one possible implementation, see Figure 2 When optimizing the greedy adversarial perturbation with the goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set, the greedy attack optimization objective function is adopted:

[0047]

[0048] st||δ g || p <∈

[0049] in, Video features for greedy perturbation videos Text features of greedy attack sample sets The similarity between them is obtained by matrix multiplication, Sim(.) is the similarity function; δ g To fight against disturbances for greed, ‖·‖ p l p The norm, ∈, is a perturbation limitation parameter to ensure the invisibility of video perturbations.

[0050] In a possible implementation, when optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set, the greedy adversarial perturbation is optimized in the following manner: obtaining the gradient of the greedy attack optimization objective function, and projecting the gradient of the greedy attack optimization objective function to a hypersphere space of size ∈, obtaining the projected gradient and optimizing the greedy adversarial perturbation based on the projected gradient.

[0051] Explanatory, by projecting the gradient of the greedy attack optimization objective function to the hypersphere space of size ∈, to ensure that the greedy attack perturbation optimization satisfies l p The constraint that the norm is less than ∈.

[0052] In one possible implementation, see again Figure 2 The adversarial sample generation method for the text video retrieval system further includes: constructing a cautious attack sample set based on a plurality of text samples; wherein the cautious attack sample set includes a target text sample set and a risk text sample set, which are text samples corresponding to the first preset number of videos ranked first in the recommended video list and the second preset number of videos ranked last, respectively; using the same text encoder as the text video retrieval system to extract text features of the target text sample set and the risk text sample set; repeating the second optimization step to a preset number of repetitions to obtain a final cautious adversarial perturbation, and obtaining a cautious perturbation video of the candidate video based on the final cautious adversarial perturbation as a cautious attack adversarial sample: wherein the second optimization step includes: obtaining a cautious perturbation video of the candidate video based on the current cautious adversarial perturbation, and using the same video encoder as the text video retrieval system to extract video features of the cautious perturbation video; optimizing the cautious adversarial perturbation with maximizing the ratio of a first cautious similarity to a second cautious similarity as the optimization target; wherein the first cautious similarity is the similarity between the video features of the cautious perturbation video and the text features of the target text sample set, and the second cautious similarity is the similarity between the video features of the cautious perturbation video and the text features of the risk text sample set.

[0053] Explanatory, cautious attacks aim to make the attacked candidate videos close to query texts with similar semantics, while staying away from query texts with different semantics, so as to avoid the attacked candidate videos appearing in irrelevant text queries, thereby arousing user suspicion and achieving more accurate video ranking improvement.

[0054] At this time, the attacker selects text samples that are close to the text description of the candidate video as the target text sample set T rel,train , and select text samples that are completely different from the text description of the candidate video as the risk text sample set T ir,train , in order to optimize the cautious adversarial perturbation so that the final cautious attack adversarial sample is as close as possible to the target text sample set T rel,train , while staying as far away from the risk text sample set T as possible ir,train .

[0055] Exemplarily, the number of target text sample sets and risk text sample sets is set to 10, that is, the first preset number and the second preset number are both 10. Optionally, a target text sample test set and a risk text sample test set can be constructed at the same time. In this case, the first preset number and the second preset number are both set to 20, and 10 target text samples are randomly selected as the target text sample set, and the remaining 10 target text samples are used as the target text sample test set; similarly, 10 risk text samples are randomly selected as the risk text sample set, and the remaining 10 risk text samples are used as the risk text sample test set.

[0056] For example, when constructing the target text sample set and the risk text sample set, text samples of different candidate videos are not shared with each other, that is, independent query and extraction are performed on each candidate video.

[0057] Explanatory, when using the same text encoder as the text video retrieval system to extract the text features of the target text sample set and the text features of the risk text sample set, the above-mentioned process of using the same text encoder as the text video retrieval system to extract the text features of the greedy attack sample set can be referred to, and will not be repeated here.

[0058] Explanatory note, when using the same video encoder as the text video retrieval system to extract video features of the cautious perturbation video, the process of extracting video features of the greedy perturbation video using the same video encoder as the text video retrieval system can be referred to, and will not be repeated here.

[0059] In a possible implementation, when optimizing the cautious countermeasure disturbance with maximizing the ratio of the first cautious similarity to the second cautious similarity as the optimization goal, a cautious attack optimization objective function is adopted:

[0060]

[0061] st‖δ s ‖ p <∈

[0062] in, For the first cautious similarity, is the second cautious similarity, Sim(.) is the similarity function, To carefully perturb the video features of the video, is the text feature of the target text sample set, is the text feature of the risk text sample set; δ s To be cautious against disturbances, p for l p The norm, ∈, is a perturbation limitation parameter to ensure the invisibility of video perturbations.

[0063] In a possible implementation, when optimizing the cautious adversarial perturbation with the optimization objective of maximizing the ratio of the first cautious similarity to the second cautious similarity, the cautious adversarial perturbation is optimized in the following manner: obtaining the gradient of the cautious attack optimization objective function, and projecting the gradient of the cautious attack optimization objective function to a hypersphere space of size ∈, obtaining the projected gradient and optimizing the cautious adversarial perturbation based on the projected gradient.

[0064] Explanatory, by projecting the gradient of the cautious attack optimization objective function into the hypersphere space of size ∈, we ensure that the cautious attack perturbation optimization satisfies l pThe constraint that the norm is less than ∈.

[0065] Explanatory, the adversarial sample generation method of the text video retrieval system of the present invention is based on the attack idea of optimizing perturbations by using image-text similarity, and designs two adversarial attack methods, greedy attack and cautious attack. By selecting different text samples as perturbation optimization objects, it realizes video promotion attacks on ordinary users using candidate videos, solving the problems that existing attack methods are difficult to implement and the attack targets are unreasonable.

[0066] In one possible implementation, this method for generating adversarial samples for a text-video retrieval system was experimentally demonstrated against greedy and cautious attacks on three Transformer models: Singularity, DRL, and Cap4Video; and on three datasets: MSR-VTT, DiDeMo, and ActivityNet, in three scenarios: white-box, gray-box, and black-box. The adversarial samples for the text-video retrieval system generated by this method effectively improved video ranking and video recall on the test set texts under different datasets, different models, and different attack scenarios. This validated the adversarial robustness vulnerabilities of the text-video retrieval system and optimized the text-video retrieval system accordingly, achieving good results. Furthermore, cautious attacks not only achieved better ranking and higher recall on relevant texts, but also ensured their concealment in risk testing, preventing the perturbed videos from appearing under irrelevant query texts, thus verifying the practical significance of this method.

[0067] The following are device embodiments of the present invention, which can be used to perform the method embodiments of the present invention. For details not disclosed in the device embodiments, please refer to the method embodiments of the present invention.

[0068] See also Figure 3 In another embodiment of the present invention, a text-video retrieval system adversarial sample generation system is provided, which can be used to implement the above-mentioned text-video retrieval system adversarial sample generation method. Specifically, the text-video retrieval system adversarial sample generation system includes a query module, a feature module and a sample generation module.

[0069] Among them, the query module is used to query the text video retrieval system based on the text description of the candidate video, obtain a recommended video list and obtain the text description of each video in the recommended video list, and obtain a number of text samples; the feature module is used to construct a greedy attack sample set based on a number of text samples, and use the same text encoder as the text video retrieval system to extract the text features of the greedy attack sample set; the sample generation module is used to repeat the first optimization step to a preset number of repetitions to obtain the final greedy adversarial perturbation, and obtain the greedy perturbation video of the candidate video based on the final greedy adversarial perturbation as the greedy attack adversarial sample; wherein the first optimization step includes: obtaining a greedy perturbation video of the candidate video based on the current greedy adversarial perturbation, and using the same video encoder as the text video retrieval system to extract the video features of the greedy perturbation video; optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set.

[0070] In a possible implementation, the adversarial sample generation system of the text-video retrieval system further includes a cautious sample set construction module, a cautious feature module, and a cautious sample generation module.

[0071] Among them, the cautious sample set construction module is used to construct a cautious attack sample set based on a number of text samples; wherein the cautious attack sample set includes a target text sample set and a risk text sample set, which are text samples corresponding to the first preset number of videos ranked in the recommended video list and the second preset number of videos ranked in the back, respectively; the cautious feature module is used to use the same text encoder as the text video retrieval system to extract text features of the target text sample set and the text features of the risk text sample set; the cautious sample generation module is used to repeat the second optimization step to a preset number of repetitions to obtain the final cautious adversarial perturbation, and obtain a cautious perturbation video of the candidate video based on the final cautious adversarial perturbation as a cautious attack adversarial sample: wherein the second optimization step includes: obtaining a cautious perturbation video of the candidate video based on the current cautious adversarial perturbation, and extracting video features of the cautious perturbation video using the same video encoder as the text video retrieval system; optimizing the cautious adversarial perturbation with maximizing the ratio of the first cautious similarity to the second cautious similarity as the optimization target; wherein the first cautious similarity is the similarity between the video features of the cautious perturbation video and the text features of the target text sample set, and the second cautious similarity is the similarity between the video features of the cautious perturbation video and the text features of the risk text sample set.

[0072] All relevant contents of each step involved in the embodiment of the adversarial sample generation method of the aforementioned text video retrieval system can be referred to the functional description of the functional module corresponding to the adversarial sample generation system of the text video retrieval system in the embodiment of the present invention, and will not be repeated here.

[0073] The module division in the embodiments of the present invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in various embodiments of the present invention may be integrated into a single processor, exist physically as separate modules, or two or more modules may be integrated into a single module. The integrated modules may be implemented in either hardware or software functional modules.

[0074] In another embodiment of the present invention, a computer device is provided, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the adversarial sample generation method of the text video retrieval system.

[0075] In another embodiment of the present invention, the present invention further provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space that stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the adversarial sample generation method for the text video retrieval system in the above embodiment.

[0076] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0077] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0078] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for generating adversarial samples for a text-video retrieval system, characterized in that: include: Based on the text description of the candidate video, a text video retrieval system is queried to obtain a list of recommended videos and a text description of each video in the recommended video list to obtain a number of text samples; A greedy attack sample set is constructed based on several text samples, and the text features of the greedy attack sample set are extracted using the same text encoder as the text video retrieval system. Repeat the first optimization step to a preset number of repetitions to obtain the final greedy adversarial perturbation, and obtain a greedy perturbation video of the candidate video based on the final greedy adversarial perturbation as a greedy attack adversarial sample; wherein the first optimization step includes: obtaining a greedy perturbation video of the candidate video based on the current greedy adversarial perturbation, and using the same video encoder as the text video retrieval system to extract video features of the greedy perturbation video; optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set.

2. The adversarial sample generation method for a text-video retrieval system according to claim 1, characterized in that: When optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set, the greedy attack optimization objective function is adopted: s.t.||δ g || p <∈ in, Video features for greedy perturbation videos Text features of greedy attack sample sets The similarity between them, Sim(.) is the similarity function; δ g To fight against disturbances for greed, ‖·‖ p l p norm, ∈ is the disturbance limit parameter.

3. The adversarial sample generation method for a text-video retrieval system according to claim 2, characterized in that: When optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set, the greedy adversarial perturbation is optimized in the following manner: Obtain the gradient of the greedy attack optimization objective function, project the gradient of the greedy attack optimization objective function to the hypersphere space of size ∈, obtain the projected gradient and optimize the greedy adversarial perturbation according to the projected gradient.

4. The adversarial sample generation method for a text-video retrieval system according to claim 1, characterized in that: Also includes: Constructing a cautious attack sample set based on a plurality of text samples; wherein the cautious attack sample set includes a target text sample set and a risky text sample set, which are text samples corresponding to the first preset number of videos ranked in the recommended video list and the second preset number of videos ranked in the recommended video list respectively; The same text encoder as the text video retrieval system is used to extract text features of the target text sample set and the text features of the risk text sample set; Repeat the second optimization step to a preset number of repetitions to obtain the final cautious adversarial perturbation, and obtain the cautious perturbation video of the candidate video based on the final cautious adversarial perturbation as the cautious attack adversarial sample: Among them, the second optimization step includes: obtaining a cautious perturbation video of the candidate video based on the current cautious adversarial perturbation, and using the same video encoder as the text video retrieval system to extract the video features of the cautious perturbation video; optimizing the cautious adversarial perturbation with the ratio of the first cautious similarity and the second cautious similarity as the optimization target; wherein the first cautious similarity is the similarity between the video features of the cautious perturbation video and the text features of the target text sample set, and the second cautious similarity is the similarity between the video features of the cautious perturbation video and the text features of the risky text sample set.

5. The adversarial sample generation method for a text-video retrieval system according to claim 4, characterized in that: When optimizing the cautious counter-perturbation with the ratio of maximizing the first cautious similarity to the second cautious similarity as the optimization goal, the cautious attack optimization objective function is adopted: s.t.‖δ s ‖ p <∈ in, For the first cautious similarity, is the second cautious similarity, Sim(.) is the similarity function, To carefully perturb the video features of the video, is the text feature of the target text sample set, is the text feature of the risk text sample set; δ s To be cautious against disturbances, p l p norm, ∈ is the disturbance limit parameter.

6. The method for generating adversarial samples for a text-video retrieval system according to claim 5, characterized in that: When optimizing the cautious counter-perturbation with the maximization of the ratio of the first cautious similarity to the second cautious similarity as the optimization target, the cautious counter-perturbation is optimized in the following manner: Obtain the gradient of the cautious attack optimization objective function, project the gradient of the cautious attack optimization objective function to the hypersphere space of size ∈, obtain the projected gradient and optimize the cautious adversarial perturbation based on the projected gradient.

7. The method for generating adversarial samples for a text-video retrieval system according to claim 4, wherein: When constructing the greedy attack sample set based on the plurality of text samples, the text samples of different candidate videos are shared with each other; when constructing the cautious attack sample set based on the plurality of text samples, the text samples of different candidate videos are not shared with each other.

8. A text-video retrieval system adversarial sample generation system, characterized by: include: A query module is used to query the text video retrieval system based on the text description of the candidate video, obtain a recommended video list and obtain the text description of each video in the recommended video list, and obtain a number of text samples; The feature module is used to construct a greedy attack sample set based on a number of text samples and extract text features of the greedy attack sample set using the same text encoder as the text video retrieval system; The sample generation module is used to repeat the first optimization step to a preset number of repetitions to obtain the final greedy adversarial perturbation, and obtain a greedy perturbation video of the candidate video based on the final greedy adversarial perturbation as a greedy attack adversarial sample; wherein the first optimization step includes: obtaining a greedy perturbation video of the candidate video based on the current greedy adversarial perturbation, and using the same video encoder as the text video retrieval system to extract video features of the greedy perturbation video; optimizing the greedy adversarial perturbation with the optimization goal of maximizing the similarity between the video features of the greedy perturbation video and the text features of the greedy attack sample set.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the adversarial sample generation method for a text video retrieval system as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for generating adversarial samples for a text-video retrieval system as claimed in any one of claims 1 to 7 are implemented.