Method and device for generating adversarial samples

By using vector pools to retrieve similar vectors in the adversarial sample generation method, the problem of insufficient effectiveness and authenticity of adversarial samples in the prior art is solved, and more effective adversarial sample generation is achieved, which improves the robustness of the model.

CN112990383BActive Publication Date: 2025-05-23ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110510166.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-11
Publication Date
2025-05-23
Estimated Expiration
2041-05-11

AI Technical Summary

Technical Problem

The prior art has poor effectiveness when generating adversarial samples, which cannot effectively improve the robustness of the model. The generated adversarial samples cannot meet the authenticity requirements and cannot appear in actual attack behavior.

Method used

By obtaining the original sample, at least two original vectors are obtained, the vector to be disturbed is selected and the adversarial perturbation is added, and the perturbation vector is obtained. Search vectors similar to perturbation vectors in a preset vector pool, and use these similar vectors to generate adversarial samples.

Benefits of technology

The generation of adversarial samples is achieved more efficiently, so that the generated adversarial samples meet both adversarial requirements and authenticity requirements, and can appear in subsequent attack behaviors, thereby better enhancing the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112990383B_ABST
    Figure CN112990383B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a method and device for generating adversarial samples. In the method, firstly, an original sample is obtained; at least two original vectors are obtained according to the original sample; a vector to be disturbed is selected from the at least two original vectors; an adversarial disturbance is added to the vector to be disturbed to obtain a disturbance vector; a vector similar to the disturbance vector is retrieved from a pre-set vector pool; wherein the vector pool includes each historical original vector obtained according to each historical original sample; and an adversarial sample is obtained according to the retrieved similar vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to electronic information technology, and more particularly, to methods and devices for generating adversarial samples. Background Art

[0002] Adversarial samples refer to input samples that can cause machine learning algorithms to output incorrect results after minor adjustments. For example, in image recognition, the original image is an image of the number "7". After adding noise perturbation, the recognition model incorrectly classifies the perturbed image as "2". This image with added noise perturbation is an adversarial sample.

[0003] Adversarial samples can be used for enhanced training of models to improve the robustness of the models. Therefore, a more effective method for generating adversarial samples is needed. Summary of the invention

[0004] One or more embodiments of this specification describe a method and device for generating adversarial samples, which can generate adversarial samples more effectively.

[0005] According to a first aspect, a method for generating an adversarial sample is provided, comprising:

[0006] Get the original sample;

[0007] According to the original sample, at least two original vectors are obtained;

[0008] Selecting a vector to be disturbed from the at least two original vectors;

[0009] Add adversarial disturbance to the disturbance vector to obtain the disturbance vector;

[0010] Retrieving a vector similar to the disturbance vector from a pre-set vector pool; wherein the vector pool includes each historical original vector obtained according to each historical original sample;

[0011] Based on the retrieved similar vectors, adversarial samples are obtained.

[0012] Wherein, the sample corresponds to a specified business type;

[0013] The historical original vectors included in the vector pool are obtained according to the historical original samples corresponding to the specified service type.

[0014] The adding of adversarial disturbance to the disturbance vector includes:

[0015] Get the disturbance value corresponding to the vector to be disturbed in one dimension;

[0016] The value of the vector to be disturbed in this dimension is increased or decreased by a predetermined multiple of the disturbance value corresponding to this dimension.

[0017] Wherein, obtaining the disturbance value corresponding to the vector to be disturbed in one dimension includes at least one of the following:

[0018] Randomly generate a disturbance value corresponding to one dimension of the vector to be disturbed;

[0019] Obtain a direction vector of the gradient function used by the model, and use the value of the direction vector in one dimension as the perturbation value corresponding to the vector to be perturbed in the dimension; wherein the model is a model trained by the original sample and the adversarial sample.

[0020] The step of retrieving a vector similar to the disturbance vector from a preset vector pool includes:

[0021] The simHash algorithm, the KNN algorithm or the KDTree algorithm is used to retrieve a vector similar to the disturbance vector from a pre-set vector pool.

[0022] The vector similar to the disturbance vector satisfies at least one of the following:

[0023] The Euclidean distance to the disturbance vector is less than a preset distance value; the preset distance value is a positive integer;

[0024] The cosine distance from the disturbance vector is greater than the preset angle value;

[0025] The Jaccard similarity coefficient with the disturbance vector is greater than the preset coefficient value;

[0026] is not equal to the vector to be disturbed.

[0027] The step of obtaining an adversarial sample based on the retrieved similar vectors includes:

[0028] Using the similar vectors retrieved this time, similar samples are obtained;

[0029] Inputting the similar samples into the model to obtain a first recognition result;

[0030] Determining whether a difference between the first recognition result and a second recognition result obtained when the original sample is input into the model satisfies a confrontation requirement;

[0031] If not, return to the step of adding the counter-disturbance to the vector to be disturbed to the step of judging until the judgment result is yes;

[0032] If yes, the vector retrieved this time is determined as the adversarial vector;

[0033] Generate adversarial examples using adversarial vectors.

[0034] in,

[0035] The sample is text data;

[0036] The granularity of the text data corresponding to the vector is: character, n-gram segment, word or sentence.

[0037] According to a second aspect, a device for generating an adversarial sample is provided, comprising:

[0038] An input module, configured to obtain raw samples;

[0039] A vector conversion module, configured to obtain at least two original vectors according to the original samples;

[0040] A disturbance processing module is configured to select a vector to be disturbed from the at least two original vectors; add an adversarial disturbance to the vector to be disturbed to obtain a disturbance vector;

[0041] The adversarial sample determination module is configured to retrieve a vector similar to the perturbation vector in a pre-set vector pool; wherein the vector pool includes each historical original vector obtained according to each historical original sample; and obtain the adversarial sample according to the retrieved similar vector.

[0042] Wherein, the disturbance processing module is configured to execute:

[0043] Get the disturbance value corresponding to the vector to be disturbed in one dimension;

[0044] The value of the vector to be disturbed in this dimension is increased or decreased by a predetermined multiple of the disturbance value corresponding to this dimension.

[0045] The disturbance processing module is configured to perform at least one of the following:

[0046] Randomly generate a disturbance value corresponding to one dimension of the vector to be disturbed;

[0047] Obtain a direction vector of the gradient function used by the model, and use the value of the direction vector in one dimension as the perturbation value corresponding to the vector to be perturbed in the dimension; wherein the model is a model trained by the original sample and the adversarial sample.

[0048] The adversarial sample determination module is configured to execute:

[0049] Using the similar vectors retrieved this time, similar samples are obtained;

[0050] Inputting the similar samples into the model to obtain a first recognition result;

[0051] Determining whether a difference between the first recognition result and a second recognition result obtained when the original sample is input into the model satisfies a confrontation requirement;

[0052] If not, return to the step of adding the counter-disturbance to the vector to be disturbed to the step of judging until the judgment result is yes;

[0053] If yes, the vector retrieved this time is determined as the adversarial vector;

[0054] Generate adversarial examples using adversarial vectors.

[0055] According to a third aspect, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any embodiment of the present specification is implemented.

[0056] The adversarial sample generation method and device provided in the embodiments of the present specification are pre-set with a vector pool, which includes various historical original vectors, and each historical original vector is obtained based on various original samples that actually existed in history. Therefore, the vector in the vector pool can correspond to an actually existing text data, such as an actually existing vocabulary, and will not correspond to a non-existent text data. In this way, after the original vector is perturbed to obtain a perturbation vector, the adversarial sample is not directly obtained by using the perturbation vector, but a vector similar to the perturbation vector is retrieved from the vector pool. The retrieved vector meets the adversarial requirement because it is similar to the perturbation vector. At the same time, the retrieved vector can correspond to an actually existing text data, and therefore also meets the authenticity requirement, that is, it may appear in subsequent attack behaviors. Using such retrieved similar vectors can more effectively obtain adversarial samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0058] Figure 1 It is a flowchart of a method for generating adversarial samples in one embodiment of this specification.

[0059] Figure 2 It is a schematic diagram of a method for generating adversarial samples in one embodiment of this specification.

[0060] Figure 3 It is a schematic diagram of the structure of an adversarial sample generation device in one embodiment of this specification. DETAILED DESCRIPTION

[0061] In the prior art, when generating adversarial samples, a gradient perturbation adversarial solution is usually adopted. That is, in order to train a model, the parameters are adjusted in the direction of the gradient descent of the model's loss function, so as to continuously optimize the model. Correspondingly, adversarial samples are fine-tuned in the direction of the gradient descent of the loss function, so as to achieve the purpose of adversarial by causing a greater impact on the model results with a small perturbation.

[0062] However, when the prior art uses the above-mentioned gradient perturbation method to generate adversarial samples, the effectiveness of the generated adversarial samples is poor, and thus it is impossible to perform better enhanced training on the model.

[0063] Taking the training sample as text data as an example, the shortcomings of the existing technical methods are explained. For example, because the text space is different from the continuous pixel space of the image, an original sample, such as a word "transfer" in an article, is discrete and discontinuous in the vector space after being converted to vector A. After the vector A converted from the original sample is gradient perturbed, the perturbed vector A' may not be restored to the text space, that is, the perturbation vector A' cannot correspond to a real word. It can be seen that the adversarial samples generated in this way cannot meet the authenticity requirements and will not appear in actual attack behaviors, so it is impossible to better enhance the model training.

[0064] In order to solve the problems of existing technologies, it is necessary to make the generated adversarial samples meet the requirements of authenticity and may appear in actual attack behaviors.

[0065] The specific implementation of the above concept is described below.

[0066] Figure 1 The flowchart of the method for generating adversarial samples in one embodiment of the present specification is shown. The execution subject of the method is the device for generating adversarial samples. It can be understood that the method can also be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 1 and Figure 2 , the method comprising:

[0067] Step 101: Obtain original samples.

[0068] Step 103: Obtain at least two original vectors according to the original samples.

[0069] Step 105: Select a vector to be disturbed from the at least two original vectors.

[0070] Step 107: Add adversarial disturbance to the disturbance vector to obtain a disturbance vector.

[0071] Step 109: searching a vector similar to the disturbance vector in a pre-set vector pool; wherein the vector pool includes each historical original vector obtained according to each historical original sample.

[0072] Step 111: Obtain an adversarial sample based on the retrieved similar vectors.

[0073] It can be seen that in the above Figure 1 In the process, a vector pool is pre-set, which includes various historical original vectors, and each historical original vector is obtained based on various original samples that actually existed in history. Therefore, the vector in the vector pool can correspond to a real sample data, such as an actually existing word, and will not correspond to a non-existent sample data. In this way, after the original vector is perturbed to obtain a perturbation vector, the perturbation vector is not directly used to obtain the adversarial sample, but a vector similar to the perturbation vector is retrieved from the vector pool. The retrieved vector meets the adversarial requirement because it is similar to the perturbation vector. At the same time, the retrieved vector can correspond to a real sample data such as a word or sentence in a text, so it also meets the authenticity requirement, that is, it may appear in subsequent attack behaviors. The use of such retrieved similar vectors can meet both the adversarial requirement and the authenticity requirement, and the adversarial sample can be obtained more effectively.

[0074] Below Figure 1 Each step is described below.

[0075] First, in step 101, an original sample is obtained.

[0076] In the field of machine learning technology, samples are usually used to train models. The original sample is the sample that needs to be input into the trained model, and the model will output a recognition result for the original sample, such as the first recognition result.

[0077] In one embodiment of this specification, Figure 1 The adversarial sample generation method shown can be applied to model training in the image field, so that the original sample in this step 101 is image data.

[0078] In another embodiment of this specification, Figure 1 The adversarial sample generation method shown can be applied to model training in the text field. In this way, the original sample in step 101 is text data, such as a paragraph of text or an article.

[0079] Next, in step 103, at least two original vectors are obtained according to the original samples.

[0080] Taking the original sample as text data as an example, the process of step 103 may include:

[0081] Step 1031: Preprocess the original text.

[0082] Preprocessing may include cleaning and tokenizing the original text, and the granularity of tokenization may be word level, character level, or n-gram level.

[0083] Step 1033: Map the preprocessed text to the vector space to obtain at least two original vectors corresponding to the original text.

[0084] For example, word vectors are obtained through self-supervised training using methods such as word2vec and glove. If the granularity of the text data corresponding to the original vector is the word, the obtained word vector is the original vector in this step. If the granularity of the text data corresponding to the original vector is the sentence or article level, then a text feature extraction model such as TextCNN, LSTM, Transformer, BERT and other models are further used in combination with specific text tasks such as text classification to obtain a sentence or article-level representation vector, and the obtained sentence or article-level representation vector is used as the original vector in this step.

[0085] Next, in step 105, a vector to be disturbed is selected from at least two original vectors.

[0086] Usually, when performing perturbations, not all original vectors corresponding to the original samples are perturbed, but a small number of original vectors are selected for perturbation, thereby achieving the adversarial purpose of changing the model recognition results with tiny perturbations.

[0087] For example, taking the original sample as image data, when each pixel in the image is mapped to the vector space, the original vectors corresponding to each pixel are obtained. For example, there are 3 million original vectors. In order to meet the purpose of confrontation, 10,000 original vectors can be selected from the 3 million original vectors as the vectors to be disturbed.

[0088] For example, taking the original sample as text, after mapping the text to the vector space, we get 10,000 original vectors corresponding to 10,000 words. To meet the adversarial purpose, we can select 50 original vectors corresponding to the words from the 10,000 original vectors as perturbation vectors.

[0089] Next, in step 107, an adversarial disturbance is added to the vector to be disturbed to obtain a disturbance vector.

[0090] The specific process of step 107 includes:

[0091] Step 1071: Obtain disturbance values ​​corresponding to at least one dimension of the vector to be disturbed.

[0092] Step 1073: For each dimension of the at least one dimension, increase or decrease the value of the vector to be disturbed in the dimension by a predetermined multiple of the disturbance value corresponding to the dimension.

[0093] It can be understood that the vector to be disturbed selected from the original vector is a multi-dimensional data, such as the vector to be disturbed A{3,1,5}, which includes the value 3 on dimension 1, the value 1 on dimension 2, and the value 5 on dimension 3. Therefore, the way to add disturbance can be to change the values ​​on one or more dimensions, for example, the disturbance value corresponding to dimension 1 is 1, and the disturbance value corresponding to dimension 3 is 2. The value 3 on dimension 1 is added with one times the disturbance value 1, that is, 3+1*1=4, and the value 5 on dimension 3 is subtracted with 0.25 times the disturbance value 2, that is, 5-0.25*2=4.5, thereby obtaining the disturbance vector A'{4,1,4.5} of the vector to be disturbed.

[0094] Of course, in order to simplify the processing, each dimension may correspond to the same disturbance value, and the values ​​on each dimension may be increased by the disturbance value or decreased by the disturbance value at the same time.

[0095] In one embodiment of the present specification, in step 1071, when obtaining the disturbance value corresponding to the vector to be disturbed in a certain dimension, the specific implementation method includes:

[0096] Method 1: Random generation.

[0097] In this first method, a disturbance value corresponding to a dimension can be randomly generated.

[0098] Method 2: Use direction vector.

[0099] In the second method, because the gradient function used by the model itself has a direction vector, the dimension of the direction vector is the same as the dimension of the vector to be disturbed. Therefore, the value of the direction vector in one dimension can be used as the perturbation value corresponding to the vector to be disturbed in this dimension; wherein, the model is a model trained with original samples and adversarial samples.

[0100] Next, in step 109, a vector similar to the disturbance vector is retrieved from a pre-set vector pool; wherein the vector pool includes each historical original vector obtained according to each historical original sample.

[0101] For the explanation of step 109 , the vector pool is first described.

[0102] The vector pool includes various historical original vectors obtained based on various historical original samples. Historical original samples are real samples that have been generated or appeared in actual business processes in history. For example, a paper was previously used as an original sample to train a text similarity recognition model. The article will be converted into multiple vectors in the historical process, and the multiple vectors will be added to the vector pool. This will be continuously executed, and new original vectors will be continuously added to the vector pool. It can be seen that, firstly, the vector pool includes a large number of original vectors accumulated in the historical process, which is large in number and convenient for subsequent retrieval. Secondly, each original vector in the vector pool is generated based on the original samples in the actual historical business process. Therefore, each original vector can be restored to the sample space. For example, each original vector can restore a real existing word.

[0103] In the embodiments of the present specification, only one vector pool may be established for different business types, so that the historical original vectors in the vector pool may correspond to different business types. For example, for business scenario 1 for text similarity recognition, all the original vectors generated historically in business scenario 1 are added to the vector pool, and for business scenario 2 for risk recognition of the transfer business, all the original vectors generated historically in business scenario 2 are added to the vector pool, so that the number and types of vectors in the vector pool are richer.

[0104] In another embodiment of the present specification, a vector pool can be established for each different business type. In this way, each historical original vector included in a vector pool is obtained based on each historical original sample of the corresponding specified business type. For example, for the business scenario 1 for text similarity identification, all the original vectors generated historically in the business scenario 1 are added to the vector pool 1, and for the business scenario 2 for risk identification of the transfer business, all the original vectors generated historically in the business scenario 2 are added to the vector pool 2. Each vector pool only stores the original vectors generated in the historical process of this business type. In this way, it can be further ensured that in the subsequent process, for the business type corresponding to the sample, only the vectors of the business type that have appeared historically in the vector pool of the business type are retrieved, thereby obtaining a more interpretable and similar vector, that is, it meets the similarity requirements under a specific business type, so that the retrieved results are more accurate.

[0105] In step 109, a simHash algorithm, a KNN algorithm or a KDTree algorithm may be used to retrieve a vector similar to the disturbance vector in the vector pool.

[0106] In step 109, the vector similar to the disturbance vector needs to meet at least one of the following requirements:

[0107] The first type: the Euclidean distance to the disturbance vector is less than a preset distance value; the preset distance value is a positive integer.

[0108] The smaller the Euclidean distance between two vectors, the more similar the two vectors are.

[0109] The second type: the cosine distance from the disturbance vector is greater than the preset angle value.

[0110] The larger the cosine distance between two vectors, the more similar the two vectors are.

[0111] The third type: the Jaccard similarity coefficient with the disturbance vector is greater than the preset coefficient value.

[0112] The larger the Jaccard similarity coefficient between two vectors, the more similar the two vectors are.

[0113] If the Euclidean distance between a vector in the vector pool and the currently retrieved disturbance vector is 0 (or the cosine distance is 360 degrees or the Jaccard similarity coefficient is 1), it means that the retrieved vector is exactly equal to the disturbance vector, and the disturbance vector obtained in the previous step can be restored to a real sample data such as a real word. If the Euclidean distance between a vector in the vector pool and the currently retrieved disturbance vector is not equal to 0 but only satisfies a preset distance value such as less than 0.2 (or the cosine distance is not equal to 360 degrees but only satisfies a preset angle value such as greater than 180 degrees; or the Jaccard similarity coefficient is not equal to 1 but only satisfies a preset coefficient value such as greater than 0.8), it means that although the disturbance vector obtained in the previous step may not be restored to a real sample data such as a real word, because the disturbance vector is similar to the retrieved vector, the retrieved vector (which can be restored to the real sample data) can be used to represent the disturbance vector.

[0114] The fourth type: not equal to the vector to be disturbed before the disturbance is added to the disturbance vector.

[0115] For example. Based on the original sample, a vector to be disturbed is obtained, such as A{3,1,5}. After the vector to be disturbed is perturbed, the perturbation vector A'{4,1,4.5} is obtained. In order to avoid A'{4,1,4.5} being unable to be restored to the sample space, a vector similar to the perturbation vector A' is retrieved in the vector pool, for example, a similar vector A"{4,1,4} may be retrieved. However, the vector pool may already include the historical vector A{3,1,5}, and according to the retrieval, A{3,1,5} in the vector pool may be used as a vector similar to the perturbation vector A'{4,1,4.5}. However, according to the purpose of generating adversarial samples, the retrieved similar vector cannot be the same as the vector to be disturbed, otherwise, the sample obtained using the retrieved similar vector will be the same as the original sample and cannot be used as an adversarial sample. Therefore, in order to avoid this situation, the vector similar to the perturbation vector cannot be equal to the vector to be disturbed before the perturbation vector is perturbed.

[0116] Next, in step 111, adversarial samples are obtained based on the retrieved similar vectors.

[0117] The process of this step 111 may include:

[0118] Step 1111: Use the vector retrieved this time to obtain similar samples.

[0119] Taking the original sample as image data as an example, for example, in step 105, the original vectors of each pixel in the corresponding image are obtained, for example, there are 3 million original vectors, and 10,000 original vectors are selected as 10,000 vectors to be disturbed. Then in this step 1111, the 10,000 original vectors are replaced by 10,000 vectors retrieved from the vector pool for the 10,000 original vectors, and together with other vectors that are not selected as vectors to be disturbed, an image is restored. The image has differences and similarities with the image as the original sample. The image restored here is the similar sample.

[0120] Taking the original sample as text data as an example, for example, in step 105, a text obtains 10,000 words, corresponding to 10,000 original vectors, and 50 original vectors corresponding to 50 words are selected from the 10,000 original vectors as 50 perturbation vectors. Then in this step 1111, the 50 original vectors are replaced by 50 vectors retrieved from the vector pool for the 50 original vectors, and together with other vectors that are not selected as the vectors to be perturbed, a text is restored. The text has differences and similarities with the text as the original sample. The text restored here is the similar sample.

[0121] Step 1113: input the similar sample into the model to obtain a first recognition result;

[0122] Step 1115: Determine whether the difference between the first recognition result and the second recognition result obtained when the original sample is input into the model meets the confrontation requirement, if yes, execute step 1117. Otherwise, return to execute the judgment steps from step 107 to step 1115 until the judgment result is yes.

[0123] Step 1117: Determine the vector retrieved this time as the adversarial vector.

[0124] Step 1119: Generate adversarial samples using adversarial vectors.

[0125] To be an adversarial sample, it needs to satisfy the recognition result of the model for the adversarial sample and the recognition result of the model for the original sample is different, and this difference needs to meet the adversarial requirements, such as the difference in the recognition result output by the model, such as the score, which must be greater than a set value, or one recognition result is yes and the other is no, then it can be used as an adversarial sample. For the similar vectors retrieved in the above steps, it is necessary to further verify whether they meet the adversarial requirements. Therefore, through the processing of this step 1115, vectors that meet the adversarial requirements can be screened out.

[0126] The text space is different from the continuous pixel space of the image. The discreteness of the text space makes it difficult to correspond the continuous perturbation of the representation space to the perturbation of the text space or the corresponding perturbation amplitude is too large. Therefore, the method provided in the embodiment of this specification has a better effect on the generation of text-type adversarial samples.

[0127] In one embodiment of the present specification, the granularity of the text data corresponding to the vector may be: character, n-gram segment, word or sentence.

[0128] After generating adversarial samples, the adversarial samples can be used to enhance the model training.

[0129] In one embodiment of this specification, a device for generating adversarial samples is proposed, see Figure 3 , the device 300 comprises:

[0130] An input module 301 is configured to obtain an original sample;

[0131] The vector conversion module 302 is configured to obtain at least two original vectors according to the original samples;

[0132] The disturbance processing module 303 is configured to select a vector to be disturbed from the at least two original vectors; add an adversarial disturbance to the vector to be disturbed to obtain a disturbance vector;

[0133] The adversarial sample determination module 304 is configured to retrieve a vector similar to the perturbation vector from a pre-set vector pool; wherein the vector pool includes each historical original vector obtained according to each historical original sample; and obtain an adversarial sample according to the retrieved similar vector.

[0134] In one embodiment of the device provided in this specification, the sample is applied to a specified service type;

[0135] The historical original vectors included in the vector pool are obtained according to the historical original samples corresponding to the specified service type.

[0136] In one embodiment of the device provided in this specification, the disturbance processing module 303 is configured to execute:

[0137] Get the disturbance value corresponding to the vector to be disturbed in one dimension;

[0138] The value of the vector to be disturbed in this dimension is increased or decreased by a predetermined multiple of the disturbance value corresponding to this dimension.

[0139] In one embodiment of the device provided in this specification, the disturbance processing module 303 is configured to perform at least one of the following:

[0140] Randomly generate a disturbance value corresponding to one dimension of the vector to be disturbed;

[0141] Obtain a direction vector of the gradient function used by the model, and use the value of the direction vector in one dimension as the perturbation value corresponding to the vector to be perturbed in the dimension; wherein the model is a model trained by the original sample and the adversarial sample.

[0142] In one embodiment of the device provided in this specification, the adversarial sample determination module 304 is configured to execute: using the simHash algorithm, the KNN algorithm or the KDTree algorithm to retrieve a vector similar to the perturbation vector in a pre-set vector pool.

[0143] In one embodiment of the device provided in this specification, the vector similar to the disturbance vector satisfies at least one of the following:

[0144] The Euclidean distance to the disturbance vector is less than a preset distance value; the preset distance value is a positive integer;

[0145] The cosine distance from the disturbance vector is greater than the preset angle value;

[0146] The Jaccard similarity coefficient with the disturbance vector is greater than the preset coefficient value;

[0147] is not equal to the vector to be disturbed.

[0148] In one embodiment of the apparatus provided in this specification, the adversarial sample determination module 304 is configured to execute:

[0149] Using the similar vectors retrieved this time, similar samples are obtained;

[0150] Inputting the similar samples into the model to obtain a first recognition result;

[0151] Determining whether a difference between the first recognition result and a second recognition result obtained when the original sample is input into the model satisfies a confrontation requirement;

[0152] If not, return to the step of adding the counter-disturbance to the disturbed vector to the step of judging until the judgment result is yes;

[0153] If yes, the vector retrieved this time is determined as the adversarial vector;

[0154] Generate adversarial examples using adversarial vectors.

[0155] One embodiment of the present specification provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute a method in any one of the embodiments of the present specification.

[0156] An embodiment of the present specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method in any embodiment of the specification is implemented.

[0157] It is understood that the structures illustrated in the embodiments of this specification do not constitute a specific limitation on the adversarial sample generation device. In other embodiments of the specification, the adversarial sample generation device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0158] Since the information interaction, execution process, etc. between the modules in the above-mentioned devices and systems are based on the same concept as the method embodiments of this specification, the specific contents can be found in the description of the method embodiments of this specification and will not be repeated here.

[0159] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0160] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, widgets, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0161] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. Methods for generating adversarial samples, include: Get the original sample; According to the original sample, at least two original vectors are obtained; Selecting a vector to be disturbed from the at least two original vectors; Add adversarial disturbance to the disturbance vector to obtain the disturbance vector; Retrieving a vector similar to the disturbance vector in a pre-set vector pool; wherein the vector similar to the disturbance vector satisfies that it is not equal to the vector to be disturbed; wherein the vector pool includes various historical original vectors; the various historical original vectors are obtained based on various original samples actually existing in history, and can correspond to the real sample data; According to the retrieved similar vectors, adversarial samples are obtained; The original sample is image data or text data.

2. The method according to claim 1, in, The sample corresponds to a specified business type; The historical original vectors included in the vector pool are obtained according to the historical original samples corresponding to the specified service type.

3. The method according to claim 1, in, The adding of adversarial disturbance to the disturbance vector comprises: Get the disturbance value corresponding to the vector to be disturbed in one dimension; The value of the vector to be disturbed in this dimension is increased or decreased by a predetermined multiple of the disturbance value corresponding to this dimension.

4. The method according to claim 3, in, Obtaining a disturbance value corresponding to the vector to be disturbed in one dimension includes at least one of the following: Randomly generate a disturbance value corresponding to one dimension of the vector to be disturbed; Obtain a direction vector of the gradient function used by the model, and use the value of the direction vector in one dimension as the perturbation value corresponding to the vector to be perturbed in the dimension; wherein the model is a model trained by the original sample and the adversarial sample.

5. The method according to claim 1, in, The step of retrieving a vector similar to the disturbance vector from a preset vector pool comprises: The simHash algorithm, the KNN algorithm or the KDTree algorithm is used to retrieve a vector similar to the disturbance vector from a pre-set vector pool.

6. The method according to claim 1, in, The vector similar to the disturbance vector satisfies at least one of the following: The Euclidean distance to the disturbance vector is less than a preset distance value; the preset distance value is a positive integer; The cosine distance from the disturbance vector is greater than the preset angle value; The Jaccard similarity coefficient with the disturbance vector is greater than a preset coefficient value.

7. The method according to claim 1, in, The step of obtaining an adversarial sample according to the retrieved similar vectors includes: Using the similar vectors retrieved this time, similar samples are obtained; Inputting the similar samples into the model to obtain a first recognition result; Determining whether a difference between the first recognition result and a second recognition result obtained when the original sample is input into the model satisfies a confrontation requirement; If not, return to the step of adding the counter-disturbance to the vector to be disturbed to the step of judging until the judgment result is yes; If yes, the vector retrieved this time is determined as the adversarial vector; Generate adversarial examples using adversarial vectors.

8. The method according to any one of claims 1 to 7, in, The sample is text data; The granularity of the text data corresponding to the vector is: character, n-gram segment, word or sentence.

9. Device for generating adversarial samples, include: An input module, configured to obtain raw samples; A vector conversion module, configured to obtain at least two original vectors according to the original samples; A disturbance processing module is configured to select a vector to be disturbed from the at least two original vectors; add an adversarial disturbance to the vector to be disturbed to obtain a disturbance vector; The adversarial sample determination module is configured to retrieve a vector similar to the perturbation vector in a pre-set vector pool; wherein the vector similar to the perturbation vector satisfies that it is not equal to the vector to be perturbed; wherein the vector pool includes various historical original vectors, and the various historical original vectors are obtained according to various original samples actually existing in history, and can correspond to real sample data; according to the retrieved similar vectors, an adversarial sample is obtained; The original sample is image data or text data.

10. The device according to claim 9, in, The disturbance processing module is configured to perform: Get the disturbance value corresponding to the vector to be disturbed in one dimension; The value of the vector to be disturbed in this dimension is increased or decreased by a predetermined multiple of the disturbance value corresponding to this dimension.

11. The device according to claim 10, in, The disturbance processing module is configured to perform at least one of the following: Randomly generate a disturbance value corresponding to one dimension of the vector to be disturbed; Obtain a direction vector of the gradient function used by the model, and use the value of the direction vector in one dimension as the perturbation value corresponding to the vector to be perturbed in the dimension; wherein the model is a model trained by the original sample and the adversarial sample.

12. The device according to any one of claims 9 to 11, in, The adversarial sample determination module is configured to execute: Using the similar vectors retrieved this time, similar samples are obtained; Inputting the similar samples into the model to obtain a first recognition result; Determining whether a difference between the first recognition result and a second recognition result obtained when the original sample is input into the model satisfies a confrontation requirement; If not, return to the step of adding the counter-disturbance to the vector to be disturbed to the step of judging until the judgment result is yes; If yes, the vector retrieved this time is determined as the adversarial vector; Generate adversarial examples using adversarial vectors.

13. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.

14. A computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Method and device for enhancing text samples

    CN111783451A

  • Adversarial sample generation method based on image retrieval model

    CN112199543A