Method and apparatus for repairing adversarial text

By slight perturbing of text and inputting classification models, the predicted classification results are determined to repair the original prediction results, and the problem of insufficient accuracy in combat text recognition in the prior art is solved, effectively repairing semantic attacks and improving the accuracy of classification models.

CN114254105BActive Publication Date: 2025-05-27HUAWEI TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011014720.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-24
Publication Date
2025-05-27
Estimated Expiration
2040-09-24

AI Technical Summary

Technical Problem

With the accuracy limit, existing adversarial text recognition methods are difficult to effectively identify and repair adversarial text caused by semantic attacks, resulting in vulnerability to classification models and the risk of denial of service attacks.

Method used

Generate perturbation text by slight perturbation of the adversarial text to be repaired and input it into the classification model to determine the predicted classification results of the adversarial text, thereby repairing its original prediction results and improving the accuracy of the classification model.

Benefits of technology

This method can effectively repair adversarial text caused by semantic attacks, improve the accuracy of classification models, reduce the risk of denial of service attacks, and provide repair text to improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114254105B_ABST
    Figure CN114254105B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for repairing adversarial texts in the field of artificial intelligence. The method includes: scrambling the adversarial text to be repaired to generate one or more perturbed texts, where the semantics of each perturbed text in the one or more perturbed texts is similar to or the same as the semantics of the adversarial text; inputting the one or more perturbed texts into a first classification model to obtain first classification results corresponding to the one or more perturbed texts; and determining a predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts, where the predicted classification result is different from the classification result of the adversarial text. Based on the discovery that if a slight perturbation is made to the adversarial text, it is possible to repair the original prediction result of the adversarial text, the applicant designed the above repair solution. By scrambling the adversarial text to obtain one or more perturbed texts, the possibility of obtaining a repaired text corresponding to the adversarial text is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more particularly, to a method and apparatus for identifying adversarial samples. Background Art

[0002] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Although artificial intelligence has achieved great success in many fields, it has been found through research that classification models based on artificial intelligence technology are extremely vulnerable to adversarial samples. In many cases, classification models with different structures obtained through training will misclassify the same adversarial sample.

[0003] Currently, a variety of techniques are known for identifying whether an input sample is an adversarial sample or a non-adversarial sample. However, due to the limitations of the accuracy of the identification method, the identified adversarial samples are only suspected adversarial samples. Therefore, after identifying these suspected adversarial samples, it is not appropriate to simply reject these samples. For example, when a user edits content on a public platform, even if the edited content is identified as an adversarial sample, we should not directly reject the user's operation, but should automatically remove or modify the harmful part of the edited content. In addition, always rejecting adversarial samples can easily lead to a denial-of-service attack.

[0004] Currently, especially in the field of natural language processing (NLP), existing repair techniques are all aimed at repairing adversarial texts with word-level perturbations to achieve the purpose of defense. However, as more and more adversarial attacks are targeted at the semantics of the input text, there is an urgent need for a repair technique that can address semantic attacks. Summary of the Invention

[0005] This application provides a method and apparatus for repairing adversarial texts to repair adversarial texts, which is beneficial to improving the accuracy of classification by a classification model.

[0006] In a first aspect, a method for repairing adversarial text is provided, including: scrambling the adversarial text to be repaired to generate one or more perturbed texts, where the semantics of each perturbed text in the one or more perturbed texts are similar to or the same as the semantics of the adversarial text; inputting the one or more perturbed texts into a first classification model to obtain first classification results corresponding to the one or more perturbed texts; and determining a predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts, where the predicted classification result is different from the label of the adversarial text.

[0007] Optionally, the above first classification results may include the label to which the input text output by the first classification model belongs. The above first classification results may also include the probability corresponding to the label to which the input text output by the prediction classification model belongs, or both. Similarly, the above predicted classification results may include the label to which the adversarial text belongs. The above predicted classification results may also include the probability corresponding to the label to which the adversarial text belongs, or both.

[0008] In the embodiments of the present application, the applicant, based on the discovery that if slight perturbations are made to the adversarial text, it is possible to repair the original prediction result of the adversarial text, obtains one or more perturbed texts by scrambling the adversarial text to increase the possibility of obtaining a repaired text corresponding to the adversarial text, and determines the predicted classification result of the adversarial text from the first classification results corresponding to the one or more perturbed texts to repair the prediction result of the adversarial text, which is beneficial to improving the accuracy of the classification model.

[0009] On the other hand, the semantics of each perturbed text in the one or more perturbed texts are similar to or the same as the semantics of the adversarial text, so that the solution of the embodiments of the present application can, to a certain extent, repair the adversarial attack on the semantics of the input text, which is beneficial to improving the accuracy of the classification model.

[0010] In a possible implementation manner, the method further includes: determining a repaired text of the adversarial text based on the predicted classification result, where the predicted classification result is the label of the repaired text output by the first classification model.

[0011] In the embodiments of the present application, in addition to returning the predicted classification result of the adversarial text to the user, a repaired text of the adversarial text may also be fed back to the user to prompt the reason why the input text is recognized as an adversarial text, which is beneficial to improving the user experience.

[0012] In a possible implementation, before determining the predicted classification result of the adversarial text based on the first classification result corresponding to the one or more perturbed texts, the method further includes: inputting the one or more perturbed texts into a second classification model to obtain a second classification result corresponding to each of the one or more perturbed texts, where the second classification model and the first classification model are different models with the same function; determining the predicted classification result of the adversarial text based on the first classification result corresponding to the one or more perturbed texts includes: if the first classification result is the same as the second classification result, determining the predicted classification result of the adversarial text based on the first classification result corresponding to the one or more perturbed texts.

[0013] The above-mentioned first classification result being the same as the second classification result can be understood as that the first label in the first classification result is the same as the second label in the second classification result, and / or the first output vector in the first classification result is similar to the second output vector in the second classification result.

[0014] In the embodiments of the present application, when the first classification result is the same as the second classification result, it indicates that the perturbed text corresponding to the label may be a non-adversarial text. At this time, determining the predicted classification result of the adversarial text based on the first classification result corresponding to the perturbed text is beneficial to improving the accuracy of determining the predicted classification result of the adversarial text.

[0015] In a possible implementation, at least some of the one or more perturbed texts are non-adversarial texts.

[0016] In the embodiments of the present application, using the text identified as a non-adversarial text as a perturbed text is beneficial to improving the accuracy of repairing the adversarial text.

[0017] In a possible implementation, the first classification results corresponding to the multiple perturbed texts are multiple labels, and the predicted classification result includes a predicted classification label. Determining the predicted classification result of the adversarial text based on the first classification result corresponding to the one or more perturbed texts includes: for the i-th label c among the multiple labels i generating a first hypothesis and a second hypothesis, where the first hypothesis is that the predicted label corresponding to the adversarial text is c i , and the second hypothesis is that the predicted label corresponding to the adversarial text is not c i , where i = 1,..., n, and n represents the total number of the multiple labels; performing a hypothesis test on the first hypothesis and the second hypothesis to obtain a test result, where the test result is used to indicate whether the predicted label of the adversarial text is c i ; determining the predicted label of the adversarial text based on the test result.

[0018] In the embodiments of the present application, by means of hypothesis testing, multiple labels are tested, which is beneficial to improving the accuracy of determining the predicted labels of adversarial texts.

[0019] In a possible implementation manner, scrambling the adversarial text to be repaired to generate one or more perturbed texts includes: scrambling the adversarial text to be repaired to generate the one or more perturbed texts based on at least one of random perturbation processing, text error processing, and semantic equivalence adversarial SEAs.

[0020] In the embodiments of the present application, by scrambling the adversarial text to be repaired based on at least one of random perturbation processing, text error processing, and semantic equivalence adversarial SEAs, it is beneficial to increase the probability of obtaining a repaired text.

[0021] In a second aspect, a device for repairing adversarial texts is provided, and the device includes units for executing each of the first aspect or any possible implementation manner of the first aspect.

[0022] In a third aspect, a device for repairing adversarial samples is provided, and the device has the functions of the device in the method design of the first aspect described above. These functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more units corresponding to the above functions.

[0023] In a fourth aspect, a computing device is provided, including an input / output interface, a processor, and a memory. The processor is used to control the input / output interface to transmit and receive signals or information, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the computing device executes the method in the first aspect described above.

[0024] In a fifth aspect, a computer-readable medium is provided, and the computer-readable medium stores program code. When the computer program code runs on a computer, the computer executes the methods in the above aspects.

[0025] In a sixth aspect, a computer program product is provided, and the computer program product includes: computer program code. When the computer program code runs on a computer, the computer executes the methods in the above aspects.

[0026] In a seventh aspect, a chip system is provided. The chip system includes a processor for a computing device to implement the functions involved in the above aspects. For example, to generate, receive, send, or process the data and / or information involved in the above methods. In a possible design, the chip system further includes a memory for storing the necessary program instructions and data of the computing device. The chip system may be composed of chips or may include chips and other discrete devices. Description of the Drawings

[0027] Figure 1 is a schematic diagram of a natural language processing system applicable to an embodiment of the present application.

[0028] Figure 2 is a flowchart of a method for repairing adversarial text according to an embodiment of the present application.

[0029] Figure 3 is a flowchart of a method for identifying adversarial text according to an embodiment of the present application.

[0030] Figure 4 is a schematic flowchart of a method for repairing adversarial text according to another embodiment of the present application.

[0031] Figure 5 is a schematic diagram of a system architecture applicable to training a target classification model according to an embodiment of the present application.

[0032] Figure 6 Introduce a schematic diagram of a system architecture that can implement another embodiment of the present application.

[0033] Figure 7 is a schematic diagram of a device for repairing adversarial samples according to an embodiment of the present application.

[0034] Figure 8 is a schematic block diagram of a computing device according to another embodiment of the present application. Detailed Embodiments

[0035] Next, the technical solutions in the present application will be described with reference to the drawings.

[0036] For ease of understanding of the present application, the following combines Figure 1 and Figure 2 to introduce a natural language processing system applicable to an embodiment of the present application.

[0037] Figure 1 is a schematic diagram of a natural language processing system applicable to an embodiment of the present application. Figure 1 The natural language processing system 100 shown includes a user device 110 and a data processing device 120.

[0038] The user device 110 can be an intelligent terminal such as a user's mobile phone, personal computer, or information processing center. The user device is the initiating end of natural language data processing and is the initiator of requests such as language Q&A or queries. Usually, the user can initiate requests through the user device.

[0039] The data processing device 120 can be a device or server with data processing capabilities. For example, cloud servers, network servers, application servers, and management servers, etc. The data processing device can receive query statements / voices / texts, etc. from the intelligent terminal through an interaction interface, and then perform language data processing in ways such as machine learning, deep learning, searching, reasoning, and decision-making through the memory for storing data and the processor for data processing.

[0040] Among them, the memory can be a general term, including local storage and a database for storing historical data. The database can be on the data processing device or on other network servers.

[0041] For example, the user can send input text to the data processing device 120 through the user device 110. Then, the data processing device 120 sends the received input text to the processor. Through the classification model stored in the processor, the input text is classified, and the label of the input text is fed back to the user device 110 through the interaction interface. Among them, the memory can be used to store intermediate data generated during the classification process or data such as labels.

[0042] As mentioned above, currently, there are known various techniques for identifying whether the input text is adversarial text or non-adversarial text. However, due to the limitation of the accuracy rate of the identification method, the identified adversarial text is only suspected adversarial text. Therefore, after identifying these suspected adversarial texts, it is not appropriate to completely reject the processing of these texts. For example, when a user edits content on a public platform, even if the edited content is identified as adversarial text, we should not directly reject the user's operation, but should automatically remove or modify the harmful parts of the edited content. In addition, always rejecting adversarial text can easily lead to a denial-of-service attack.

[0043] Traditional repair techniques are all aimed at repairing adversarial texts with word-level perturbations to achieve the purpose of defense. However, as more and more adversarial attacks are targeted at the semantics of the input text, there is an urgent need for a repair technique that can target semantic attacks.

[0044] Generally, attacks against the semantics of the input text usually perturb the input text into adversarial texts with the same or similar semantics. Although such perturbations make minor changes to the text and are not easily noticed by people, for the classification model, they will cause the classification model to output incorrect labels. The applicant found that if the adversarial text is slightly perturbed, it may be possible to repair the original prediction result of the adversarial text. Therefore, based on this discovery, the applicant designed a set of repair schemes for adversarial texts. The following combines Figure 2 to introduce the method for repairing adversarial texts in the embodiments of the present application.

[0045] Figure 2 is a flowchart of the method for repairing adversarial texts in the embodiments of the present application. It should be understood that Figure 2 the method shown can be executed by Figure 1 the data processing device 120 shown, or can be executed by other devices with computing capabilities. The embodiments of the present application do not make any limitations in this regard. Figure 2 The method shown includes step 210 and step 230.

[0046] 210, scramble the adversarial text to be repaired to generate one or more perturbed texts, wherein the semantics of each perturbed text in the one or more perturbed texts is similar to or the same as the semantics of the adversarial text.

[0047] The above-mentioned adversarial text refers to the input text formed by deliberately adding subtle interferences to the input text, which will cause the classification model to output an incorrect label with high confidence.

[0048] Optionally, in order to improve the accuracy of the predicted classification results of the adversarial samples, at least some of the one or more perturbed texts are non-adversarial texts. The recognition scheme for adversarial texts executed on the perturbed texts can be the recognition method introduced in Table 5 later, or other recognition methods for adversarial texts can also be used. The embodiments of the present application do not make any limitations in this regard.

[0049] There are many scrambling methods for generating perturbed texts with the same or similar semantics as the adversarial text. The embodiments of the present application do not make any limitations in this regard. The following will take random perturbation processing, text error (TextBugger) processing, and semantically equivalent adversary (SEAs) processing as examples for introduction.

[0050] First, random perturbation (RP) processing aims to randomly select a part of the words in the adversarial text and replace the selected words with their synonyms on the premise of retaining the semantics of the adversarial text.

[0051] Suppose the adversarial text x consists of n words. That is to say, the adversarial example x can be formed by concatenating n words, which is expressed as x = [w 1 , w 2 , ……, w n . For the selected i-th word w i in the adversarial example, k synonyms can be selected from the sorted list of synonyms [w i , w i1 , w i2 , …, w iL corresponding to w g . In this way, when g words in the adversarial text need to be modified, k

[0052] perturbation combinations can be obtained, where i, n, g, and k are positive integers, and k ≤ L. i It should be understood that there are many ways to generate the above-mentioned sorted list of synonyms, and the embodiments of this application do not limit this. For example, L synonyms can be selected from the embedding space corresponding to w i in ascending order of the distance from w

[0053] to other words, and used as the words in the above-mentioned sorted list of synonyms.

[0054] Table 1

[0055]

[0056] Step 1: Configure the variable "n" to store the total number of words included in the adversarial text x.

[0057] Step 2: Configure the variable "g" to store the total number of words to be replaced (i.e., the perturbed text above).

[0058] Step 3: Randomly select g words to be perturbed from the adversarial text x and store them in the set "combs".

[0059] Step 4: Use the synonym w i1 in the sorted list of synonyms [w i2 , …, w iL to replace the word w i1 in the set "combs", and assign the perturbed text obtained by the replacement to the variable "x'". i

[0060] Step 5: Return the perturbed text x'.

[0061] II. Text Bugger Handling. Generally, TEXTBUGGER includes five error generation methods: insertion, deletion, swapping, Substitute-C (Sub-C), and Substitute-W (Sub-W). Among them, insertion means inserting a space into a word in the adversarial text. Deletion means deleting any character in a word of the adversarial text except the first and last characters. Swapping means randomly swapping two adjacent letters in a word of the adversarial text without changing the first or last letter of the word. Sub-C means substituting a character in a word of the adversarial text with a visually similar character or an adjacent character on the keyboard. Visually similar characters can be, for example, "0" and "o", "1" and "l", "@" and "a"; adjacent characters on the keyboard can be, for example, "n" and "m". Sub-W means selecting a target word in the context-aware word vector space and replacing the word in the adversarial text with the target word.

[0062] When determining the label corresponding to the perturbed text using the classification method introduced below, after the perturbed text is input into the first classification model and the second classification model, if the similarity between the classification result output by the first classification model and the classification result output by the second classification model is relatively low, for example, the Kullback-Leibler Divergence (KL divergence) is large, it will be recognized as adversarial text. Therefore, to avoid the processed perturbed text being recognized as adversarial text, we can select important sentences from the adversarial text, determine important words from the important sentences, and perturb the important words.

[0063] Suppose the adversarial sample can be formed by concatenating m sentences, that is, x = [s 1 , s 2 , ……, s m . The important sentences can be obtained by inputting the sentence s i into the first classification model f 1 (·) and the second classification model f 2 (·) respectively, obtaining two classification results f 1 (s i ) and f 2 (s i ), and taking the sentences with the KL divergence kl(f 1 (s i ), f 2 (s i )) greater than the preset value as important sentences.

[0064] Suppose the selected important sentence is s i , we can determine the word w jsentence s i and without w j sentence s i Change Δ of KLD between KLD to obtain word w j importance, that is

[0065] Δ KLD = kl(f 1 (s i ), f 2 (s i )) - kl(f 1 (s i \w j ), f 2 (s i \w j )),(1)

[0066] where s i \w j represents sentence s without w j sentence s i s i represents sentence s with w j sentence s i .

[0067] The following combines the steps of the TextBugger algorithm shown in Table 2 to introduce the process of generating perturbed text in the embodiments of the present application. The algorithm shown in Table 2 includes 12 steps.

[0068] Table 2

[0069]

[0070] Step 1: Configure the array variable "C s " to store the importance scores of each sentence in the adversarial text x.

[0071] Steps 2 to 3: Calculate the KL divergence KLD(s i ) of the i-th sentence in the adversarial text x, and assign KLD(s i ) to the i-th element "C s (i)" in the array variable.

[0072] Step 4: Sort the sentences in the adversarial sample x according to the importance scores of the sentences recorded in the array variable "C s ", and assign the sorted sentences to the array variable "S ordered ".

[0073] Steps 5 to 9: For each sentence s in the array variable "S ordered " iPerform the following operations: Configure the array variable "C w " to store the importance scores of each word in the i-th sentence s i ; and calculate the importance score of the j-th word w i in the i-th sentence s j based on the above equation (1), and then assign the importance score of the j-th word w j to the j-th element C w in the array variable C w ; finally, sort each word in the i-th sentence s w according to the importance scores of the words recorded in the array variable "C i ", and assign the sorted words to the array variable "W ordered ".

[0074] Step 10: Select g words according to S ordered and W ordered and store them in the set variable "combs".

[0075] Step 11: Replace each word in the set variable "combs" with its synonym, and assign the perturbed text obtained by the replacement to the variable "x'".

[0076] Step 12: Return the perturbed text x'.

[0077] III. SEAs are designed to provide semantically-preserving perturbations. Optionally, neural machine translation (NMT) can be used to obtain the above semantically-preserving perturbed text. Formally, NMT can be represented as a function T(s, d, x): X s → X d , where s represents the source language, d represents the target language, and x represents the adversarial text. The process of generating the perturbed text is to translate the adversarial text into another language and then translate it back, that is, the perturbed text x' can be represented as x' = T(d, s, T(s, d, x)).

[0078] It should be noted that multiple interferences can be generated by changing the target language d to improve the diversity of the perturbed text. This is beneficial to increasing the possibility of the original prediction result of the adversarial sample. Of course, the adversarial text x can also be translated into multiple languages to generate more perturbations. For example, two different target languages d 1 and d 2 can be selected for translation. In this way, the perturbed text x' can be represented as x' = T(d 2 , s, T(d 1 , d 2 , T(s, d 1, x)).

[0079] The process of generating perturbed text in the embodiments of the present application is introduced below in combination with the SEAs algorithm steps shown in Table 3. The algorithm shown in Table 3 includes 5 steps.

[0080] Table 3

[0081]

[0082] Step 1: Select a target language from the set of target languages therein.

[0083] Step 2: Use NMT to translate the sentences in the adversarial text from the source language s to the target language d, and store them in the variable T 1 therein.

[0084] Step 3: Use NMT to translate the text obtained in Step 2 from the source language d to the target language s, and store it in the variable T 2 therein.

[0085] Step 4: Assign the text obtained in Step 2 to the variable x'.

[0086] Step 5: Return the scrambled text T 2 (d, s, x').

[0087] 220. Input one or more perturbed texts into the first classification model to obtain the first classification results corresponding to the one or more perturbed texts.

[0088] The above first classification results may include the labels to which the perturbed texts output by the first classification model belong. The above first classification results may also include the probabilities corresponding to the labels to which the perturbed texts output by the first classification model belong, or both.

[0089] The above first classification results corresponding to the one or more perturbed texts may include one classification result or multiple classification results, which are not limited in the embodiments of the present application.

[0090] The above first classification model may also be referred to as a classifier, which is used to map the input text to one of the given categories, so that it can be applied to data prediction. Among them, the classification model is a general term for the models that classify input samples in data mining. The classification model may include a classification model based on a decision tree, a classification model based on logistic regression, a classification model based on naive Bayes, and a classification model based on a neural network.

[0091] 230. Determine the predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts, where the predicted classification result is different from the classification result of the adversarial text.

[0092] The above prediction classification result may include the label to which the adversarial text belongs. The above prediction classification result may also include the probability corresponding to the label to which the adversarial text belongs, or both.

[0093] The above prediction classification result can be understood as the classification result of the original text corresponding to the adversarial text predicted by the method of the embodiments of the present application. Or rather, the original prediction classification result of the adversarial text predicted by the method of the embodiments of the present application. The above prediction classification result is different from the classification result of the adversarial text, and may include that the prediction classification label is different from the classification label of the adversarial text.

[0094] The label of the above adversarial text can be the label of the adversarial text output by the first classification model, or the label of the adversarial text predicted by other classification models. The embodiments of the present application do not make any limitations thereto.

[0095] There are many ways to determine the above prediction classification result. For example, since the label of the adversarial text output by the first classification model is an incorrect label, then a label different from the above incorrect label can be selected from the first classification result as the prediction classification result. This method will have a high accuracy in binary classification, because the label corresponding to the adversarial text is usually incorrect, so the prediction classification result corresponding to the adversarial text is the label other than the incorrect label among the two labels. Another example is that in the case where there are multiple perturbed texts, the label corresponding to the majority of the perturbed texts in the first classification result can be selected as the prediction classification result.

[0096] However, when the above-described scheme for determining the prediction classification result is applied to a multi-class classification scenario, the accuracy of the determined prediction classification result may not be satisfactory. Therefore, in order to improve the accuracy of determining the prediction classification result, the present application also provides a scheme for determining the prediction classification result based on hypothesis testing.

[0097] That is, the first classification results corresponding to the above multiple perturbed texts include multiple labels, and the prediction classification result includes a prediction classification label. The above step 230 includes: for the i-th label c among the multiple labels i Generate a first hypothesis and a second hypothesis. The first hypothesis is that the predicted label corresponding to the adversarial text is c i and the second hypothesis is that the predicted label corresponding to the adversarial text is not c i where i = 1,..., n, and n represents the total number of the multiple labels; perform hypothesis testing on the first hypothesis and the second hypothesis to obtain a test result, and the test result is used to indicate whether the predicted label of the adversarial text is c i Determine the predicted label of the adversarial text based on the test result.

[0098] The above first hypothesis can be used as the "null hypothesis" and is denoted as H0 (c i ): P(f(x) = c i ) ≥ ρ, the above second hypothesis can be used as an alternative hypothesis and is denoted as H 1 (c i ): P(f(x) = c i ) < ρ, where P(f(x) = c i ) represents the probability that the predicted label of the adversarial text x is c i , and ρ represents a preset probability threshold.

[0099] Optionally, the probability that the predicted label of the above adversarial text x is c i can be calculated by the following formula

[0100]

[0101] where X * represents the above-mentioned multiple scrambled texts, y represents the scrambled text whose label corresponding to the first classification model among the multiple scrambled texts is c i , |·| represents the number of texts calculated. In other words, |y ∈ X * ∧ f 1 (y) = c i | represents the number of texts of the scrambled text y among the multiple scrambled texts, where the label obtained by inputting the scrambled text y into the first classification model is c i , |X * | represents the number of texts of the multiple scrambled texts.

[0102] Since there are multiple labels, a set of hypotheses can be configured for each label among the multiple labels, and a hypothesis test is performed for each set of hypotheses. The specific methods for performing hypothesis tests include the Fixed-size Sampling Test (FSST) and the Sequential probability ratio test (SPRT). For FSST, it aims to perform hypothesis tests on a fixed number of texts. That is, a set X * containing a large number of perturbed texts needs to be generated first, and then according to the above formula (2), for each label c i among the multiple labels, calculate P(f(x) = c i ), and then compare the test result with the probability threshold ρ.

[0103] It should be noted that in order to determine the error range, it is necessary to determine the minimum number of perturbed texts required for FSST. In actual situations, in order to improve the accuracy, the number of perturbed texts usually required to execute FSST is very large. As the number of perturbed texts increases, the computational overhead occupied during the execution of FSST becomes larger. If the adversarial text repair scheme based on FSST is executed online, there may be a problem of insufficient computing power. Therefore, SPRT can also be used for hypothesis testing, and the number of perturbed texts required can be dynamically determined in this testing scheme, which is usually faster than FSST. The hypothesis testing process based on SPRT is introduced below in conjunction with Table 4. Algorithm 5 shown in Table 4 includes 11 steps.

[0104] Table 4

[0105]

[0106] Step 1: Configure the variable "k" to record the number of perturbed texts in the set X of perturbed texts. * of the perturbed texts.

[0107] Step 2: Configure the variable "z" to record the number of the perturbed text y in the set X of perturbed texts, where the perturbed text y is the perturbed text in the set X of perturbed texts * with the label c * of the perturbed texts. i

[0108] Step 3: Configure the parameters α, β, σ, ρ as the parameters for hypothesis testing, where α represents the probability of rejecting the null hypothesis when the null hypothesis is true; β represents the probability of accepting the null hypothesis when the alternative hypothesis is true; ρ represents the preset probability threshold; σ represents the preset value.

[0109] It should be noted that by setting the value of σ, the difference between "p 0 " and "p 1 " can be controlled. When the value of σ is large, the difference between "p 0 " and "p 1 " is large. When the value of σ is small, the difference between "p 0 " and "p 1 " is small.

[0110] Step 4: Assign the result of ρ - σ to the variable "p 0 ".

[0111] Step 5: Assign the result of ρ + σ to the variable "p 1 ".

[0112] Step 6: Assign the likelihood ratio Pr(z, k, p 0 , p 1)Assign to the variable "sprt_ratio", where the likelihood ratio Pr(z,k,p 0 ,p 1 ) satisfies the formula

[0113] Steps 7 to 8: If then return to accept the first hypothesis and reject the second hypothesis.

[0114] Steps 9 to 10: If then return to accept the second hypothesis and reject the first hypothesis.

[0115] Step 11: If neither nor is satisfied, then return "uncertain".

[0116] It should be noted that if the likelihood ratio Pr(z,k,p 0 ,p 1 ), and do not satisfy the situation in Step 7 nor the situation in Step 9, it can be determined that the above test result is uncertain, which means that more perturbed text is needed.

[0117] It should be noted that in order to reduce the computational amount, it is not necessary to perform hypothesis tests on all of the above n labels. During the process of performing hypothesis tests on the labels, once it is determined that the predicted label of the adversarial text for the hypothesis test is c i holds, the hypothesis tests on the remaining labels can be stopped. Of course, it is also possible to perform tests on all of the multiple labels, and then select one from the multiple labels corresponding to the predicted label of the adversarial text for which the test result is c i as the final predicted label of the adversarial text. The embodiments of the present application do not make any limitations in this regard.

[0118] In order to improve the user experience, in addition to returning the predicted classification result of the adversarial text to the user, the solution of the embodiments of the present application can also return the repaired text of the adversarial text to the user to prompt the reason why the input text is recognized as an adversarial text.

[0119] That is, the above method further includes 240, determining the repaired text of the adversarial text based on the predicted classification result, where the predicted classification result is the classification result of the repaired text output by the first classification model.

[0120] That is to say, after determining the predicted classification result of the adversarial text based on the solution introduced above, the perturbed text corresponding to the predicted classification result can be used as the repaired text of the adversarial text.

[0121] It should be noted that in the embodiments of the present application, only the repaired text of the adversarial text can be returned to the user, or only the predicted classification result of the adversarial text can be returned to the user, or both the repaired text and the predicted classification result of the adversarial text can be returned to the user. The embodiments of the present application do not limit this.

[0122] As described above in conjunction with Figures 1 to 2 the method for repairing adversarial text in the embodiments of the present application, the method for identifying adversarial text applicable to the embodiments of the present application will be introduced below in conjunction with Figure 1 It should be noted that the adversarial text recognition scheme introduced below can be applied to the field of natural language processing, or can be applied to other fields (for example, the field of image processing). In addition, the repair scheme of the embodiments of the present application can also be used in combination with other adversarial text recognition methods. The embodiments of the present application do not limit this.

[0123] Figure 3 is a flowchart of the method for identifying adversarial text in the embodiments of the present application. Figure 3 The method shown can be executed by any device with computing functions. The embodiments of the present application do not limit this. Figure 3 The method shown includes: step 310 to step 330.

[0124] 310. Obtain the input text to be recognized.

[0125] 320. Input the input text into the first classification model and the second classification model respectively to obtain the first classification result of the input text and the second classification result of the input text, where the first classification model and the second classification model are different models with the same function, the first classification result is obtained by the first classification model classifying the input text, and the second classification result is obtained by the second classification model classifying the input text.

[0126] Optionally, the above first classification result may include the label to which the input text belongs output by the first classification model, and the above first classification result may also include the probability corresponding to the label to which the input text belongs output by the first classification model, or both. Similarly, the above second classification result may include the label to which the input text belongs output by the second classification model, and the above second classification result may also include the probability corresponding to the label to which the input text belongs output by the second classification model, or both.

[0127] The above first classification model and the second classification model are models with the same function. It can be understood that the above two classification models are used to determine the classification to which the input text belongs from the same multiple classifications, or in other words, the given classifications corresponding to the above two classification models are the same.

[0128] The above-mentioned first classification model and second classification model are different, which may include that the weights in the first classification model are different from those in the second classification model. The above-mentioned first classification model and second classification model are different, and may also include that the above-mentioned first classification model and second classification model are different models trained based on different subsets in the training dataset. The above-mentioned first classification model and second classification model are different, and may also include that the above-mentioned first classification model and second classification model are models with different structures. For example, when both the first classification model and the second classification model are neural network models, that the first classification model and the second classification model are models with different structures can be understood as that the number of hidden layers in the first classification model is different from the number of hidden layers in the second classification model. The above-mentioned first classification model and second classification model are different, and may also include that the above-mentioned first classification model and second classification model are models with different types. For example, the first classification model is a classification model based on a decision tree, and the second classification model is a classification model based on a neural network.

[0129] In the embodiments of the present application, when the first classification model and the second classification model are models with different types, it is beneficial to improve the accuracy of identifying adversarial texts, and it avoids the situation where when the types of the first classification model and the second classification model are the same, the generalization capabilities of the first classification model and the second classification model are similar, resulting in both classification models being unable to identify adversarial texts.

[0130] Optionally, the above-mentioned first classification model and second classification model may be trained based on different training texts, or the above-mentioned first classification model and second classification model may also be obtained through model mutation. The embodiments of the present application do not limit this.

[0131] 330. Based on the first classification result and the second classification result, identify the input text as an adversarial text or a non-adversarial text.

[0132] Optionally, the above-mentioned step 330 includes: if the first classification result and the second classification result are the same, determine that the input text is a non-adversarial text; if the first classification result and the second classification result are different, determine that the input text is an adversarial text.

[0133] However, the above-mentioned solution of only identifying the input text as a non-adversarial text or an adversarial text based on the first classification result and the second classification result is not accurate. In some cases, when the first classification result and the second classification result are the same, there is still a possibility that the input text is an adversarial text. Therefore, in the identification solution provided in the embodiments of the present application, when the first classification result and the second classification result are the same, the input text can be further identified based on the first output vector output by the first classification model for the input text and the second output vector output by the second classification model for the input text.

[0134] That is, the above first classification result includes a first label and a first output vector, and the above second classification result includes a second label and a second output vector. The above method further includes: inputting the input text into the first classification model and the second classification model respectively to obtain the first output vector corresponding to the input text and the second output vector corresponding to the input text. The first output vector is used to indicate the confidence of the input text belonging to each classification in the first classification model, and the second output vector is used to indicate the confidence of the input text belonging to each classification in the second classification model. The above step 330 includes: when the first label and the second label are the same, based on the first output vector and the second output vector, identifying whether the input text is an adversarial text or a non-adversarial text.

[0135] The above first classification model and second classification model can be multi-classification models. Therefore, whether it is the first classification model or the second classification model, while outputting the label corresponding to the input text, it can also output an output vector (also known as a "probability vector") to indicate the confidence of the input text in each classification corresponding to the classification model.

[0136] Generally, when the first label and the second label are the same, if the first output vector and the second output vector are the same or not very different, the input text can be identified as a non-adversarial text. If the first output vector and the second output vector are quite different, the input text can be identified as an adversarial text.

[0137] Suppose the input text is represented as x, the first classification model is represented as f 1 (·), and the second classification model is represented as f 2 (·). Then the first output vector output by the first classification model for the input text can be represented as f 1 (x) = [p 0 , p 1 , …, p K , and the second output vector output by the second classification model for the input text can be represented as f 2 (x) = [q 0 , q 1 , …, q K , where K represents the total number of multiple classifications corresponding to the first classification model and the second classification model, i represents the i-th classification among the multiple classifications corresponding to the first classification model and the second classification model, i = 1, ……, K, p i represents the confidence that the input text predicted by the first classification model is classified as the i-th classification, and q i represents the confidence that the input text predicted by the second classification model is classified as the i-th classification.

[0138] Optionally, the difference between the first output vector and the second output vector can be represented by the similarity between the first output vector and the second output vector. That is, the above method further includes: obtaining the similarity between the first output vector and the second output vector; when the first label and the second label are the same, based on the first output vector and the second output vector, identifying the input text as an adversarial text or a non-adversarial text, including: when the first label and the second label are the same, if the similarity is higher than a preset first similarity threshold, determining that the input text is a non-adversarial text; when the first label and the second label are the same, if the similarity is lower than a preset second similarity threshold, determining that the input text is an adversarial text.

[0139] Wherein, the first similarity threshold can be the same as the second similarity threshold, the first similarity threshold can be different from the second similarity threshold, and the first similarity threshold is greater than the similarity threshold. The embodiments of the present application do not make any limitations in this regard.

[0140] Of course, the difference between the first output vector and the second output vector can also be represented based on the difference between each component in the first output vector and each component in the second output vector. The embodiments of the present application do not make any limitations in this regard.

[0141] The similarity between the first output vector and the second output vector can be calculated by KL divergence, or by other ways of calculating similarity. The embodiments of the present application do not make any limitations in this regard.

[0142] Below, taking the calculation of the similarity between the first output vector and the second output vector by KL divergence as an example, a method for identifying adversarial texts is introduced.

[0143] The first output vector f 1 (x) and the KL divergence between the second output vector f 2 (x) can be expressed as Wherein, kl(f 1 (x), f 2 (x)) represents the KL divergence between the first output vector f 1 (x) and the second output vector f 2 (x). When kl(f 1 (x), f 2 (x)) < ε, it indicates that the first output vector f 1 (x) and the second output vector f 2 (x) are similar, then the input text can be identified as a non-adversarial text. When kl(f 1 (x), f 2 (x)) ≥ ε, it indicates that the first output vector f 1(x) and the second output vector f 2 If (x) is not similar, the input text can be identified as adversarial text, where ε represents a preset similarity threshold.

[0144] The method for identifying the input text based on the KL divergence between the first label, the second label, and the first output vector f 1 (x) and the second output vector f 2 (x) can be represented by the algorithm shown in Table 5.

[0145] Table 5

[0146]

[0147] Among them, step 1 shows that the first label output by the first classification model f 1 (·) for the input text x is c 1 .

[0148] Step 2 shows that the second label output by the second classification model f 2 (·) for the input text x is c 2 .

[0149] Steps 3 and 4 show that if the first label is c 1 and the second label is c 2 are the same, and kl(f 1 (x), f 2 (x)) < ε, then return "false", indicating that the input text is non-adversarial text.

[0150] Step 5 If the first label is c 1 and the second label is c 2 are not the same, and / or kl(f 1 (x), f 2 (x)) < ε is not satisfied, then return "true", indicating that the input text is adversarial text.

[0151] The above only takes the first classification model and the second classification model to identify the input text as an example for introduction. To improve the accuracy of identifying adversarial texts, the solution of the present application can also be used in scenarios of more than two classification models. For example, multiple classification models can be divided into multiple groups of models. The above first classification model and second classification model belong to the first group of models among the multiple groups of models. The combinations of classification models included in different groups of models among the multiple groups of models are different. Then, the above method further includes: obtaining a first prediction result and at least one second prediction result, where the first prediction result is used to indicate the prediction result of the first group of models predicting the input text, and the at least one second prediction result is used to indicate the prediction results of other groups of models predicting the input text. The prediction results include that the input text is an adversarial text or the input text is a non-adversarial text. The other groups of models are the model groups other than the first group of models among the multiple groups of models; based on the first prediction result and the at least one second prediction result, identify whether the input text is an adversarial text or a non-adversarial text.

[0152] The above identifying whether the input text is an adversarial text or a non-adversarial text based on the first prediction result and at least one second prediction result may include identifying whether the input text is an adversarial text or a non-adversarial text by means of majority voting based on the first prediction result and at least one second prediction result. Of course, in addition to the majority voting method, other methods can also be used to identify whether the input text is an adversarial text or a non-adversarial text. For example, when one of the first prediction result and at least one second prediction result indicates that the input text is an adversarial text, it can be identified that the input text is an adversarial text. The embodiments of the present application do not make specific limitations in this regard.

[0153] The above multiple classification models are models with the same function, that is, the above multiple classification models are used to determine the classification to which the input text belongs from the same multiple classifications, or in other words, the given classifications corresponding to the above multiple classification models are the same.

[0154] Among the multiple classification models described above, different classification models are different models, which may include different weights of different classification models among the multiple classification models. Among the multiple classification models described above, different classification models are different models, and may also include that different classification models among the multiple classification models are trained based on different subsets in the training dataset. Among the multiple classification models described above, different classification models are different models, and may also include that different classification models among the multiple classification models have different structures. For example, when different classification models among the multiple classification models are all neural network models, different classification models among the multiple classification models with different structures can be understood as having different numbers of hidden layers among different classification models among the multiple classification models. Among the multiple classification models described above, different classification models are different models, and may also include that different classification models among the multiple classification models are models of different types. For example, among the multiple classification models, there are classification models based on decision trees and classification models based on neural networks.

[0155] In the embodiments of the present application, when the multiple classification models are models of different types, it is beneficial to improve the accuracy of identifying adversarial texts, and it avoids the situation where when the types of the multiple classification models are the same, the generalization capabilities of the multiple classification models are similar, resulting in the failure of all the multiple classification models to identify adversarial texts.

[0156] Optionally, the multiple classification models described above may be trained based on different training texts, or the multiple classification models may also be obtained through model mutation. The embodiments of the present application do not make any limitations in this regard.

[0157] For example, if the multiple classification models described above include Classification Model 1, Classification Model 2, and Classification Model 3, then the above 3 classification models can be divided into 3 groups, that is, Model Group 1 includes Classification Model 1 and Classification Model 2; Model Group 2 includes Classification Model 1 and Classification Model 3; Model Group 3 includes Classification Model 2 and Classification Model 3.

[0158] Then, according to the scheme described above, the prediction result 1 of Model Group 1 for the input text, the prediction result 2 of Model Group 2 for the input text, and the prediction result 3 of Model Group 3 for the input text can be obtained respectively. Then, based on the above 3 prediction results, the input text can be identified as a non-adversarial text or an adversarial text by means of majority voting. That is, if the majority of the above 3 prediction results indicate that the input text is a non-adversarial text, then the input text is a non-adversarial text; if the majority of the above 3 prediction results indicate that the input text is an adversarial text, then the input text is an adversarial text.

[0159] As described above in conjunction with Figure 3A solution for identifying adversarial texts is introduced. The above solution can also be used in the solution for determining whether a perturbed text is an adversarial text in the embodiments of the present application. Before step 230, the above method further includes: inputting one or more perturbed texts into a second classification model to obtain a second classification result corresponding to each of the one or more perturbed texts, where the second classification model and the first classification model are different models with the same function; determining a predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts, including: if the first classification result is the same as the second classification result, determining the predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts.

[0160] Of course, it is also possible to determine whether one or more perturbed texts are adversarial texts based on the solution combining tags and output vectors introduced above. After determining that the perturbed text generated above is an adversarial text, the perturbed text can be deleted from the "one or more perturbed texts" so that each perturbed text in the "one or more perturbed texts" above is a non-adversarial text.

[0161] For ease of understanding, the following describes a solution for repairing adversarial texts in another embodiment of the present application in combination with Table 6, a solution for identifying adversarial texts, and a solution for repairing adversarial texts. The method shown in Table 6 includes 20 steps.

[0162] Table 6

[0163]

[0164] Steps 1 and 20: Call Algorithm 1 to identify the input text x. If the input text x is non-adversarial, execute step 20 and return the input text x. If the input text x is identified as an adversarial text, execute steps 2 to 19.

[0165] Step 2: Configure a set X of perturbed texts with an initial value of empty * .

[0166] Step 3: Configure a set C with an initial value of empty to store possible labels, where the possible labels can be understood as the labels corresponding to the above one or more perturbed texts.

[0167] Step 4: Configure a set D with an initial value of empty to store rejected labels, where the rejected labels can be understood as the labels that the predicted classification result of the adversarial text cannot take.

[0168] Steps 5 to 6: Input the adversarial text x into the first classification model f 1 (·) and the second classification model f 2 (·). If the obtained labels are the same, that is, f 1 (x) = f2 (x), then the obtained label f 1 (x) or f 2 (x) is stored in the set D.

[0169] The purpose of the loop algorithm from step 7 to step 19 introduced below is to repair the adversarial text x.

[0170] Step 7: If the result returned by Algorithm 1 is "true", that is, the input text is an adversarial text, then execute step 8.

[0171] Step 8: Scramble the adversarial text x to obtain a perturbed text, and store the perturbed text in the variable y.

[0172] Steps 9 to 10: Call Algorithm 1 to identify the perturbed text y. If the perturbed text y is identified as an adversarial text, then continue to execute the loop from step 8 to step 10 until a non-adversarial perturbed text y is obtained, and then execute step 11.

[0173] Step 11: If the perturbed text y is non-adversarial, then assign the label f 1 (y) of the perturbed text y to the variable c.

[0174] Steps 12 to 13: If the label f 1 (y) belongs to neither set C nor set D, then store the label f 1 (y) in set C.

[0175] Step 14: For each label c that belongs to set C and does not belong to set D i Perform the operations of steps 15 - 19.

[0176] Step 15: Call Algorithm 5 (see Table 4), and store the test result in the variable result.

[0177] Steps 16 to 17: If the test result is acceptance, then return that the perturbed text x' belongs to the set of perturbed texts X * , and the label corresponding to the perturbed text x' is c i .

[0178] Steps 18 to 19: If the test result is rejection, then store the label c i in set D.

[0179] It should be noted that if the test result is rejection, then the label c iIt is added to the set D so that this label will not be tested again, and then the next iteration is continued. That is to say, in order to reduce the computational overhead, the solution of this application conducts hypothesis testing in a lazy way, by maintaining the set D to record a set of rejected labels so as to avoid testing the rejected labels in subsequent iterations.

[0180] For ease of understanding, hereinafter, taking the identification of adversarial texts by two classification models and the repair of adversarial texts as an example, in combination with Figure 4 introduce the repair process of the adversarial texts in the embodiments of this application. Figure 4 is a schematic flowchart of a method for repairing adversarial texts according to another embodiment of this application. Figure 4 The method shown includes steps 410 to step 470.

[0181] 410, Input the input text into the first classification model and the second classification model respectively to obtain the first classification result and the second classification result.

[0182] Among them, the first classification result includes the label #1 and the first output vector of the input text output by the first classification model, and the second classification result includes the label #2 and the second output vector of the input text output by the second classification model.

[0183] 420, Based on the first classification result and the second classification result, identify whether the input text is an original text or an adversarial text.

[0184] Specifically, if the label #1 and the label #2 are different, the input text can be directly identified as an adversarial text. If the label #1 and the label #2 are the same, based on the similarity between the first output vector and the second output vector, determine whether the input text is an original text or an adversarial text. That is, if the similarity between the first output vector and the second output vector is higher than the similarity threshold, it can be determined that the input text is an original text; if the similarity between the first output vector and the second output vector is lower than the similarity threshold, it can be determined that the input text is an adversarial text.

[0185] 430, If the input text is an original text, output the classification result (including the label and / or the output vector) of the input text.

[0186] 440, If the input text is an adversarial text, scramble the input text to obtain multiple perturbed texts.

[0187] 450, Input the multiple perturbed texts into the first classification model and the second classification model for classification to obtain the classification results corresponding to the multiple perturbed texts, where the classification results corresponding to the multiple perturbed texts are the classification results output by the first classification model and the second classification model respectively for each of the multiple perturbed texts.

[0188] It should be noted that the above-mentioned multiple perturbed texts can be understood as being recognized by the first classification model and the second classification model and being recognized as the original text, that is, the classification results output by the first classification model and the second classification model for each perturbed text are the same or similar. In the embodiments of the present application, for those perturbed texts recognized as adversarial texts, they can be discarded.

[0189] 460. Based on the predicted classification results corresponding to the multiple perturbed texts, determine the predicted classification result of the adversarial text.

[0190] The specific determination method can refer to the introduction related to step 230 above. For the sake of brevity, it will not be elaborated here.

[0191] 470. Based on the predicted classification result, determine the repaired text of the adversarial text.

[0192] As described above, the solution of the present application is applied to the field of natural language processing to identify the input text. Generally, when identifying the input text, since the input text is usually sequential data, we can utilize the advantage of the recurrent neural network (RNN) in processing sequential data and select the RNN as the above-mentioned first classification model. Optionally, in order to improve the recognition speed, we can also select a classification model based on the convolutional neural network (CNN) as the above-mentioned second classification model. In addition, since the structures of the RNN and the CNN are quite different, it is beneficial to improve the accuracy of identifying adversarial texts and avoid the situation where when the types of the first classification model and the second classification model are the same, the generalization capabilities of the first classification model and the second classification model are similar, resulting in both classification models being unable to identify adversarial texts.

[0193] For the sake of easy understanding, the following combines Figure 5 and Figure 6 to introduce the training processes of the first classification model and the second classification model in the embodiments of the present application and the devices that can implement the training processes. It should be noted that different training data sets can be used in the training processes of the above two classification models, or the same training data set can be used. The embodiments of the present application do not make any limitations in this regard. In addition, since the training processes of the two classification models are similar, for the sake of brevity, the following takes the training of one of the classification models (referred to as the "target classification model") as an example for introduction.

[0194] Figure 5 is a schematic diagram of the system architecture applicable to the training of the target classification model in the embodiments of the present application. See the appendix Figure 5, an embodiment of the present application provides a system architecture 500. The data acquisition device 560 is used to acquire training texts and store them in the database 530, and the training device 530 generates a target classification model / rule 501 based on the training texts maintained in the database 530. How the training device 530 obtains the target classification model / rule 501 based on the training texts will be described in more detail below. The target classification model / rule 501 can classify the input texts.

[0195] Taking the target classification model as a deep neural network as an example, the work of each layer in the deep neural network can be described by the mathematical expression as follows: From a physical perspective, the work of each layer in the deep neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / dimensionality reduction; 2. Magnification / shrinkage; 3. Rotation; 4. Translation; 5. "Bending". Among them, the operations of 1, 2, and 3 are completed by , the operation of 4 is completed by +b, and the operation of 5 is implemented by a(). The reason for using the word "space" here is that the object to be classified is not a single thing, but a class of things. Space refers to the set of all individuals of this class of things. Among them, W is the weight vector, and each value in the vector represents the weight value of a neuron in this layer of the neural network. The vector W determines the space transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space. The purpose of training the deep neural network, that is, ultimately obtaining the weight matrix of all layers of the trained neural network (the weight matrix formed by many layers of vectors W). Therefore, the training process of the neural network is essentially to learn the way to control the space transformation, and more specifically, to learn the weight matrix.

[0196] Since it is desired that the output of the deep neural network be as close as possible to the value that is truly desired to be predicted, the weight vectors of each layer of the neural network can be updated by comparing the prediction result of the current network with the truly desired target result and then according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, parameters are pre-configured for each layer in the deep neural network). For example, if the prediction result of the network is too high, the weight vector is adjusted to make it predict lower, and continuous adjustment is made until the neural network can predict the truly desired target result. Therefore, it is necessary to pre-define "how to compare the difference between the prediction result and the target result", which is the loss function or the objective function. They are important equations for measuring the difference between the prediction result and the target result. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0197] The target classification model / rule obtained by the training device 520 can be applied to different systems or devices. In the appendix Figure 5 In it, the execution device 510 is configured with an I / O interface 512 to interact with external devices, and the "user" can input text to the I / O interface 512 through the client device 540.

[0198] The execution device 510 can call texts, codes, etc. in the data storage system 550, and can also store texts, instructions, etc. in the data storage system 550.

[0199] The calculation module 511 classifies the input text using the target classification model / rule 501 to determine the category corresponding to the input text.

[0200] Finally, the I / O interface 513 returns the prediction result to the client device 540 for the user.

[0201] Optionally, the training device 530 can generate corresponding target classification models / rules 501 based on different texts for different targets to provide better results for the user.

[0202] In the appendix Figure 5In the case shown, the user can manually specify the text in the input execution device 510. For example, the user can operate in the interface provided by the I / O interface 513. In another case, the client device 540 can automatically input text to the I / O interface 513 and obtain the result. If the automatic input of text by the client device 540 requires the authorization of the user, the user can set the corresponding permissions in the client device 540. The user can view the result output by the execution device 510 in the client device 540, and the specific presentation form can be specific ways such as display, sound, action, etc. The client device 540 can also be used as a data collection end to store the collected data as text in the database 530.

[0203] It should be noted that Figure 5 is only a schematic diagram of a system architecture provided by an embodiment of the present invention. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 5 the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510.

[0204] It should be noted that Figure 5 only taking the training process of the deep neural network as an example for introduction, the target classification model of the embodiments of the present application can also be other classification models such as decision trees. The training process can refer to the existing classification model training process. For the sake of brevity, it will not be elaborated here.

[0205] The following combines Figure 6 to introduce another system architecture that can implement the embodiments of the present application. Figure 6 is a schematic diagram of another system architecture applicable to the training target classification model of the embodiments of the present application.

[0206] The system architecture 600 includes an execution device 610, which is implemented by one or more servers. Optionally, in cooperation with other computing devices, such as devices for data storage, routers, load balancers, etc.; the execution device 610 can be arranged on one physical site or distributed on multiple physical sites. The execution device 610 can use the data in the data storage system 660 or call the program code in the data storage system 660 to classify the input text.

[0207] The user can operate their respective user devices (such as the local device 601 and the local device 602) to interact with the execution device 610. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a smart camera, a smart car, or other types of cellular phones, media consumption devices, wearable devices, set-top boxes, game consoles, etc.

[0208] The local device of each user can interact with the execution device 610 through a communication network using any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0209] In another implementation, one or more aspects of the execution device 610 can be implemented by each local device. For example, the local device 601 can provide local data or feedback calculation results for the execution device 610.

[0210] It should be noted that all functions of the execution device 610 can also be implemented by the local device. For example, the local device 601 implements the functions of the execution device 610 and provides services for its own users, or provides services for the users of the local device 601.

[0211] As described above in connection with Figures 1 to 6 the repair method and model training process of the embodiments of the present application are introduced. Below in connection with Figures 7 to 8 the devices of the embodiments of the present application are introduced. It should be understood that Figures 7 to 8 the devices shown can implement each step in the above method. For the sake of brevity, they will not be described in detail here.

[0212] Figure 7 is a schematic diagram of the adversarial sample repair device of the embodiments of the present application. Figure 7 The device 700 shown includes: an acquisition unit 710 and a processing unit 720. Among them, the acquisition unit 710 is used to acquire the adversarial text to be repaired, and the processing unit 720 is used to perform one or more of the following steps: scrambling the adversarial text to be repaired to generate one or more perturbed texts, where the semantics of each perturbed text in the one or more perturbed texts is similar to or the same as the semantics of the adversarial text; inputting the one or more perturbed texts into a first classification model to obtain first classification results corresponding to the one or more perturbed texts; and determining a predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts, where the predicted classification result is different from the label of the adversarial text.

[0213] Optionally, as an embodiment, the processing unit 720 is further used to: determine a repaired text of the adversarial text based on the predicted classification result, where the predicted classification result is the label of the repaired text output by the first classification model.

[0214] Optionally, as an embodiment, the processing unit 720 is further configured to: input the one or more perturbed texts into a second classification model to obtain a second classification result corresponding to each perturbed text in the one or more perturbed texts, where the second classification model and the first classification model are different models with the same function; if the first classification result is the same as the second classification result, determine a predicted classification result of the adversarial text based on the first classification result corresponding to the one or more perturbed texts.

[0215] Optionally, as an embodiment, at least some of the one or more perturbed texts are non-adversarial texts.

[0216] Optionally, as an embodiment, the first classification results corresponding to the multiple perturbed texts include multiple labels, and the processing unit 720 is further configured to: for the i-th label c in the multiple labels i generate a first hypothesis and a second hypothesis, where the first hypothesis is that the predicted label corresponding to the adversarial text is c i and the second hypothesis is that the predicted label corresponding to the adversarial text is not c i , where i = 1,..., n, and n represents the total number of the multiple labels; perform a hypothesis test on the first hypothesis and the second hypothesis to obtain a test result, where the test result is used to indicate whether the predicted label of the adversarial text is c i ; determine the predicted label of the adversarial text based on the test result.

[0217] Optionally, as an embodiment, the processing unit 720 is further configured to: generate the one or more perturbed texts by scrambling the adversarial text to be repaired based on at least one of random perturbation processing, text error processing, and semantic equivalence adversarial SEAs.

[0218] In an alternative embodiment, the processing unit 720 may be a processor 820, the obtaining unit 710 may be a communication interface 830, and the communication device may further include a memory 810, specifically as Figure 8 shown.

[0219] Figure 8 is a schematic block diagram of a computing device according to another embodiment of the present application. Figure 8The computing device 800 shown may include: a memory 810, a processor 820, and a communication interface 830. Among them, the memory 810, the processor 820, and the communication interface 830 are connected through an internal connection path. The memory 810 is used to store instructions, and the processor 820 is used to execute the instructions stored in the memory 820 to control the input / output interface 830 to receive / send at least some parameters of the second channel model. Optionally, the memory 810 can be coupled to the processor 820 through an interface or integrated with the processor 820.

[0220] It should be noted that the above communication interface 830 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the communication device 800 and other devices or communication networks. The above communication interface 830 may further include an input / output interface.

[0221] In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 820 or the instructions in the form of software. The method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware processor, or executed and completed by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 810, and the processor 820 reads the information in the memory 810 and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0222] It should be understood that in the embodiments of the present application, the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0223] It should also be understood that in the embodiments of the present application, the memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the processor may also include a non-volatile random access memory. For example, the processor may also store information about the device type.

[0224] It should be understood that in the embodiments of the present application, "repair" means to repair the incorrect classification output of the classification model for adversarial text into a correct classification output, and it is possible to repair the adversarial text into the original text to a certain extent.

[0225] It should also be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document generally indicates that the associated objects before and after are in an "or" relationship.

[0226] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0227] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0228] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0229] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0230] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0231] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0232] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0233] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for repairing adversarial text, characterized in that, comprising: scrambling the adversarial text to be repaired to generate one or more perturbed texts, wherein the semantics of each perturbed text in the one or more perturbed texts is similar to or the same as the semantics of the adversarial text; inputting the one or more perturbed texts into a first classification model to obtain first classification results corresponding to the one or more perturbed texts; determining a predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts; determining a repaired text of the adversarial text based on the predicted classification result, wherein the predicted classification result is different from the predicted result of the adversarial text, and the predicted classification result is the classification result of the repaired text output by the first classification model.

2. The method according to claim 1, characterized in that, before determining the predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts, the method further comprises: inputting the one or more perturbed texts into a second classification model to obtain second classification results corresponding to each perturbed text in the one or more perturbed texts, wherein the second classification model and the first classification model are different models with the same function; the determining the predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts comprises: if the first classification result is the same as the second classification result, determining the predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts.

3. The method according to claim 1, characterized in that, at least some of the one or more perturbed texts are non-adversarial texts.

4. The method according to any one of claims 1-3, characterized in that, the first classification results corresponding to the multiple perturbed texts include multiple labels, and the predicted classification result includes a predicted label, the determining the predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts comprises: For the i-th label c among the multiple labels i generate a first hypothesis and a second hypothesis, where the first hypothesis is that the predicted label corresponding to the adversarial text is c i , and the second hypothesis is that the predicted label corresponding to the adversarial text is not c i , where i = 1, ……, n, and n represents the total number of the multiple labels; Perform hypothesis testing on the first hypothesis and the second hypothesis to obtain a test result, where the test result is used to indicate whether the predicted label of the adversarial text is c i ; determining the predicted label of the adversarial text based on the test result.

5. The method according to claim 4, characterized in that, the scrambling the adversarial text to be repaired to generate one or more perturbed texts comprises: scrambling the adversarial text to be repaired to generate the one or more perturbed texts based on at least one of random perturbation processing, text error processing, and semantic equivalence adversarial SEAs.

6. An apparatus for repairing adversarial text, characterized in that, comprising: a processing unit, configured to scramble the adversarial text to be repaired to generate one or more perturbed texts, wherein the semantics of each perturbed text in the one or more perturbed texts is similar to or the same as the semantics of the adversarial text; the processing unit is further configured to input the one or more perturbed texts into a first classification model to obtain first classification results corresponding to the one or more perturbed texts; the processing unit is further configured to determine a predicted classification result of the adversarial text based on the first classification results corresponding to the one or more perturbed texts; The processing unit is further configured to determine a repaired text of the adversarial text based on the predicted classification result, where the predicted classification result is different from the label of the adversarial text, and the predicted classification result is the classification result of the repaired text output by the first classification model.

7. The apparatus according to claim 6, wherein, the processing unit is further configured to: input the one or more perturbed texts into a second classification model to obtain a second classification result corresponding to each of the one or more perturbed texts, where the second classification model and the first classification model are different models with the same function; if the first classification result is the same as the second classification result, determine the predicted classification result of the adversarial text based on the first classification result corresponding to the one or more perturbed texts.

8. The apparatus according to claim 6, wherein, at least some of the one or more perturbed texts are non-adversarial texts.

9. The apparatus according to any one of claims 6-8, wherein, the first classification results corresponding to the multiple perturbed texts include multiple labels, and the predicted classification result includes a predicted label. The processing unit is further configured to: For the i-th label c among the multiple labels i generate a first hypothesis and a second hypothesis, where the first hypothesis is that the predicted label corresponding to the adversarial text is c i , and the second hypothesis is that the predicted label corresponding to the adversarial text is not c i , where i = 1, ……, n, and n represents the total number of the multiple labels; Perform hypothesis tests on the first hypothesis and the second hypothesis to obtain test results, where the test results are used to indicate whether the predicted label of the adversarial text is c i ; determine the predicted label of the adversarial text based on the verification result.

10. The apparatus according to claim 9, wherein, the processing unit is further configured to: generate the one or more perturbed texts by scrambling the adversarial text to be repaired based on at least one of random perturbation processing, text error processing, and semantic equivalence adversarial SEAs.

11. A computing device, wherein, comprising at least one processor and a memory, the at least one processor is coupled to the memory and configured to read and execute instructions in the memory to perform the method according to any one of claims 1-5.

12. A computer-readable medium, wherein, the computer-readable medium stores program code, and when the program code runs on a computer, the computer is caused to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Text classification method and device

    CN109582792A

  • Text attack method and device based on similar dictionary and storage medium

    CN111507093A