Training for question-and-answer dialogue systems to prevent hostile attacks
By training question-answering dialogue systems with adversarial statements and bootstrapped policies, the systems become more robust against adversarial attacks, providing accurate answers in multiple languages.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- アンソロピックピービーシー
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-01
AI Technical Summary
Question-answering dialogue systems are vulnerable to adversarial attacks that provide incorrect answers, and existing technologies have not adequately addressed this issue, particularly in multilingual contexts.
A method is employed to enhance the training of question-answering dialogue systems by using adversarial statements to identify and counteract attacks, involving the generation and translation of adversarial statements in multiple languages, and reinforcing the model with bootstrapped policies to improve its robustness.
The enhanced training process makes the dialogue systems more resilient to adversarial attacks, ensuring accurate responses across different languages and contexts.
Smart Images

Figure 2026074008000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of question - answering dialogue systems used to answer questions. More specifically, the present invention relates to the field of protecting question - answering dialogue systems from adversarial attacks that break such question - answering dialogue systems.
Summary of the Invention
[0002] In one or more embodiments of the present invention, a method protects a question - answering dialogue system from being attacked by adversarial statements that answer questions incorrectly. A computing device accesses a plurality of adversarial statements that can perform an adversarial attack on a question - answering dialogue system trained to provide correct answers to specific types of questions. The computing device uses the plurality of adversarial statements to train a machine - learning model for the question - answering dialogue system. Next, the computing device enhances the trained machine - learning model by bootstrapping an adversarial policy that identifies a plurality of types of adversarial statements to the trained machine - learning model. Thereafter, when responding to questions submitted to the question - answering dialogue system, the computing device uses the trained and bootstrapped machine - learning model to prevent adversarial attacks.
[0003] In one or more embodiments of the present invention, a trained and bootstrapped machine learning model is tested by a computing device that performs the following: converting questions to a question-answering dialogue system into statements including placeholders for answers; randomly selecting answer entities from the answers and adding the randomly selected answer entities in place of the placeholders to generate adversarial statements; generating attacks against a trained and bootstrapped machine learning model that include the adversarial statements; measuring the response from the trained and bootstrapped machine learning model to the generated attacks; and modifying the trained and bootstrapped machine learning model to improve the response level of the response to the generated attacks.
[0004] In one or more embodiments of the present invention, a context passage includes a correct answer containing correct answer entities, a specific type of question includes a specific type of question entities, and the method involves a computing device generating / retrieving a Random Answer Random Question (RARQ) adversarial statement, wherein the RARQ adversarial statement includes random answer entities that replace correct answer entities in the correct answer, and the RARQ adversarial statement includes random question entities that replace correct question entities in the correct answer; generating / retrieving a Random Answer Original Question (RAOQ) adversarial statement, wherein the RAOQ adversarial statement includes random answer entities that replace correct answer entities in the correct answer, and the RAOQ adversarial statement includes correct question entities from the correct answer; and generating / retrieving a No Answer Random Question (NARQ) adversarial statement. The process further includes generating / retrieving adversarial statements, wherein the NARQ adversarial statements include random question entities that replace correct answer entities in correct answers with no answer, and the NARQ adversarial statements include correct question entities in correct answers; generating / retrieving No Answer Original Question (NAOQ) adversarial statements, wherein the NAOQ adversarial statements include correct answer entities in correct answers with no answer, and the NAOQ adversarial statements include correct question entities from correct answers; and utilizing the RARQ adversarial statements, RAOQ adversarial statements, NARQ adversarial statements, and NAOQ adversarial statements as input to further train a machine learning model for a question-answering dialogue system to recognize adversarial statements.
[0005] In one or more embodiments of the present invention, the original questions used in the question-answering dialogue system, the original contextual passages used in the question-answering dialogue system, or the adversarial statements generated for the question-answering dialogue system, or a combination thereof, are in one or more different languages, so that the question-answering dialogue system can handle adversarial attacks in multiple languages.
[0006] In one or more embodiments, the methods described herein are performed by the execution of a computer program product or a computer system or both. [Brief explanation of the drawing]
[0007] [Figure 1] This figure shows exemplary systems and networks in which the present invention is implemented in various embodiments. [Figure 2] This figure shows a high-level overview of an exemplary attack pipeline used when executing a question-answering (QA) dialogue / learning system that includes adversarial statements in a contextual passage, according to one or more embodiments of the present invention. [Figure 3] This figure shows various types of adversarial passages used in one or more embodiments of the present invention. [Figure 4] This figure illustrates an exemplary flow of steps used to generate adversarial statements in one or more embodiments of the present invention. [Figure 5] This figure shows an exemplary process for using a trained model in a question-answering dialogue system to defend against adversarial statements / attacks, according to one or more embodiments of the present invention. [Figure 6] This figure shows a high-level overview of recursive training of a transformer model system according to one or more embodiments of the present invention. [Figure 7]Figure 6 shows an exemplary embodiment of a transformer model system that uses a multilanguage bidirectional encoder representation from transformers (e.g., MBERT) according to one or more embodiments of the present invention. [Figure 8] This figure shows an exemplary question-answering dialogue system used in one or more embodiments of the present invention. [Figure 9] This figure shows an exemplary deep neural network used by the QA dialogue system 800 shown in Figure 8 to respond to a new question, according to one or more embodiments of the present invention. [Figure 10] This figure shows a high-level flowchart of one or more steps performed by the method according to one or more embodiments of the present invention. [Figure 11] This figure shows a cloud computing environment according to one or more embodiments of the present invention. [Figure 12] This figure shows an abstraction model layer of a cloud computing environment according to one or more embodiments of the present invention. [Modes for carrying out the invention]
[0008] In one or more embodiments, the present invention is a system, method, or computer program product, or a combination thereof, at any possible level of technical detail of integration. In one or more embodiments, the computer program product includes a computer-readable storage medium containing computer-readable program instructions for causing a processor to perform an aspect of the present invention.
[0009] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exclusive list of further specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooved structures on which instructions are recorded, and any appropriate combination thereof. When used herein, computer-readable storage media should not be construed as themselves being radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmitting media (e.g., light pulses passing through fiber optic cables), or transient signals such as electrical signals transmitted through wires.
[0010] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing device / processing device, or to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof). This network may include copper transmission cables, optical transmission fibers, wireless transmitters, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing device / processing device receives computer-readable program instructions from the network and transfers those computer-readable program instructions for storage on a computer-readable storage medium within the respective computing device / processing device.
[0011] In one or more embodiments, computer-readable program instructions for performing the operations of the present invention include assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java®, Smalltalk®, and C++, and conventional procedural programming languages such as the C programming language or similar programming languages. In one or more embodiments, the computer-readable program instructions are executed entirely on the user's computer, partially as a standalone software package on the user's computer, partially on the user's computer and on a remote computer, respectively, or entirely on a remote computer or on a server. In the latter scenario, in one or more embodiments, the remote computer is connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection is made to an external computer (for example, via the Internet using an Internet Service Provider). In some embodiments, to carry out aspects of the present invention, an electronic circuit including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) executes a computer-readable program instruction by customizing the electronic circuit using state information of the computer-readable program instruction.
[0012] Aspects of the present invention will be described herein by reference to flowcharts or block diagrams, or both, of methods, apparatuses (systems), and computer program products, according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and any combination of blocks contained in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.
[0013] In one or more embodiments, these computer-readable program instructions are provided to a processor of a general-purpose computer, a dedicated computer, or another programmable data processing device to create a machine, so as to create means for instructions executed via the processor of a computer or other programmable data processing device to perform functions / operations specified in one or more blocks of a flowchart or block diagram or both. In one or more embodiments, these computer-readable program instructions are stored on a computer-readable storage medium such that the computer-readable storage medium on which the instructions are stored contains a product containing instructions that perform modes of functions / operations specified in one or more blocks of a flowchart or block diagram or both, and in one or more embodiments, they instruct a computer, a programmable data processing device, or other device, or a combination thereof, to function in a particular manner.
[0014] In one or more embodiments, computer-readable program instructions are also read into a computer, another programmable data processing device, or another device to produce a computer implementation process in which instructions executed on the computer, another programmable device, or another device perform functions / operations specified in one or more blocks of a flowchart or block diagram, or both, causing a series of operable steps to be executed on the computer, another programmable device, or other device that generates the computer implementation process.
[0015] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram represents a module, segment, or portion of an instruction comprising one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions shown in the blocks occur in a different order than shown in the figure. For example, two consecutively shown blocks may actually be executed substantially simultaneously or, in some cases, in reverse order, depending on the functions they contain. It should also be noted that in one or more embodiments of the present invention, each block in the block diagram or flowchart diagram, or both, and any combination of blocks contained in the block diagram or flowchart diagram, or both, are implemented by a dedicated hardware-based system that performs a specified function or operation, or a combination of dedicated hardware and computer instructions.
[0016] Referring here to the figures, and in particular to Figure 1, an exemplary block diagram of a system and network that may be used by, in, or both in, an implementation of the present invention is shown. Note that with respect to computer 101, and including both the depicted hardware and software shown within computer 101, some or all of the exemplary architecture may be used by the artificial intelligence 124 shown in Figure 1, or the software deployment server 150, or the text document server 152, or the audio file server 154, or the question-answering dialogue system 156, or the question transmission system 158, or the video file server 160, or a combination thereof, or the controller 601 shown in Figure 6, or the multilingual transformer bidirectional encoder representation (e.g., MBERT) system 724 shown in Figure 7, or one or more neurons / nodes shown in the deep neural network 924 depicted in Figure 9, or a combination thereof.
[0017] An exemplary computer 101 includes a processor 104 coupled to a system bus 106. The processor 104 may be one or more processors, each of which includes one or more processor cores. A video adapter 108 that drives / supports a display 110 is also coupled to the system bus 106. The system bus 106 is coupled to an input / output (I / O) bus 114 via a bus bridge 112. An I / O interface 116 is coupled to the I / O bus 114. The I / O interface 116 provides communication with various I / O devices, including a keyboard 118, a mouse 120, a media tray 122 (which may include storage devices such as a CD-ROM drive or multimedia interface), artificial intelligence 124, and an external USB port 126. The form of the ports connected to the I / O interface 116 may be any form known to those skilled in the art of computer architecture, but in one embodiment, some or all of those ports are universal serial bus (USB) ports.
[0018] As shown in the figure, computer 101 can also communicate with artificial intelligence 124, or software deployment server 150, or text document server 152, or audio file server 154, or question and answer dialogue system 156, or question sending system 158, or video file server 160, or a combination thereof, using network interface 130 with network 128. Network interface 130 is a hardware network interface such as a network interface card (NIC). Network 128 can be an external network such as the Internet, or an internal Internet such as Ethernet (registered trademark) or virtual private network (VPN). Hereinafter, one or more examples of physical device 154 are presented.
[0019] Hard drive interface 132 is also coupled to system bus 106. Hard drive interface 132 interfaces with hard drive 134. In one embodiment, hard drive 134 inputs data into system memory 136 which is also coupled to system bus 106. System memory is defined as the lowest level of volatile memory within computer 101. This volatile memory includes, but is not limited to, cache memory, registers, and buffers, and may also include additional higher level volatile memory (not shown in the figure). Data input into system memory 136 includes operating system (OS) 138 and application program 144 of computer 101.
[0020] OS138 includes a shell 140 to provide transparent user access to resources such as application program 144. Generally, the shell 140 is a program that provides an interpreter and interface between the user and the operating system. More specifically, the shell 140 executes commands input to the command-line user interface or commands from a file. Thus, the shell 140 (also called a command processor) is typically the highest-level software layer of the operating system and functions as a command interpreter. The shell provides a system prompt, interprets commands input by the keyboard, mouse, or other user input media, and sends the interpreted commands to a more appropriate lower-level operating system (such as kernel 142) for processing. Although the shell 140 is a text-based line-oriented user interface, it should be noted that the present invention equally and appropriately supports other user interface modes such as graphics, voice, and gesture.
[0021] As shown in the figure, OS138 also includes a kernel 142 that includes lower-level functions of OS138, including providing essential services (including memory management, process and task management, disk management, and mouse and keyboard management) required by other parts of OS138 and application program 144.
[0022] Application program 144 includes a renderer, exemplified as browser 146. Browser 146 includes program modules and instructions that enable a World Wide Web (WWW) client (i.e., computer 101) to send and receive network messages to and from the Internet using hypertext transfer protocol (HTTP) messaging, thus enabling communication with software deployment server 150 and other computer systems.
[0023] The application program 144 in the system memory of computer 101 (and the system memory of the software deployment server 150) also includes question answering dialog system protection logic (QADSPL) 148. QADSPL 148 contains code for implementing the processes described below, including the processes described in Figures 2 to 10. In one embodiment, computer 101 can download QADSPL 148 from the software deployment server 150, and this download includes on-demand downloads in which the code for QADSPL 148 is not downloaded until it is needed for execution. Furthermore, note that in one embodiment of the present invention, the software deployment server 150 performs all the functions related to the present invention (including the execution of QADSPL 148), so computer 101 does not need to use its own internal computing resources to execute QADSPL 148.
[0024] The text document server 152 is a server that transmits context (i.e., text passages such as the text passage shown in Figure 3) to the computer 101, AI 124, or QA question-answering dialogue system 156, or a combination thereof, by matching a specific type of question (received by the computer 101, AI 124, or QA question-answering dialogue system 156, or a combination thereof) with a specific set of candidate answer texts.
[0025] The audio file server 154 is a server that transmits context (i.e., audio files) to the computer 101, AI 124, or QA question-answering dialogue system 156, or a combination thereof, by matching a specific type of question (received by the computer 101, AI 124, or QA question-answering dialogue system 156, or a combination thereof) with a specific set of candidate answer audio files. In other words, the audio file server 154 interprets the type of question received and returns the relevant audio files (identified, for example, by metadata describing each audio file) that have a subject that matches that type of question. For example, if the question is about a specific type of music, the audio file server 154 returns audio files that contain metatags describing that specific type of music.
[0026] The QA dialogue system 156 is a system that responds to questions (for example, from the question transmission system 158) with answers using the processes / systems described herein.
[0027] The video file server 160 is a server that transmits context (i.e., video files) to the computer 101, AI 124, or QA question-answering dialogue system 156, or a combination thereof, by matching a specific type of question (received by computer 101, AI 124, or QA question-answering dialogue system 156, or a combination thereof) with a specific set of candidate answer video files. That is, the video file server 160 interprets the type of question received and returns relevant video files that have a subject that matches that type of question (identified, for example, by metadata describing each video file). For example, if the question is about a particular type of visual art, the video file server 160 returns video files that contain metatags describing that particular type of visual art.
[0028] It should be noted that the hardware elements shown in computer 101 are not intended to be exhaustive, but rather representative examples to highlight essential components required by the present invention. For example, computer 101 may include alternative memory storage devices such as magnetic cassettes, digital versatile disks (DVDs), and Bernoulli cartridges. These and other variations are intended to fall within the scope of the present invention.
[0029] Question-answer (QA) systems, also known as question-answer dialogue systems, are essential tools used by people seeking answers. An exemplary QA system receives a question (e.g., "What is the oldest cafe in Paris?"), searches a corpus of resources such as text, video, and audio, and returns the correct answer (e.g., "Cafe X").
[0030] Therefore, such a QA system is preferably robust to ensure that it can provide users with the correct answers. That is, a QA system is weak if it fails to function against a malicious attack (as described in detail below), and robust if it can successfully defend against a malicious attack (as described and asserted in one or more embodiments of the present invention).
[0031] Therefore, one or more embodiments of the present invention provide a robust QA system that not only defends the QA system itself against malicious attacks but also handles multilingual malicious attacks.
[0032] As described herein, one or more embodiments of the present invention utilize one or more novel types of adversarial statements to expose weaknesses in multilingual question answer (MLQA) systems.
[0033] These new types of adversarial statements are used to train QA models, thus making the trained QA models more robust in combating malicious attacks.
[0034] In one or more embodiments of the present invention, a trained QA model is enhanced by bootstrapping adversarial policies (e.g., policies describing which of a new type of adversarial type should be monitored) to create a more effective QA model for training an MLQA system.
[0035] Accordingly, in one or more embodiments of the present invention, the method / apparatus generates attack statements in any language of the MLQA system by: converting an original question into a general statement by using placeholders for answers; randomly selecting various entities to replace question entities or answer entities, or both, found in the original question, in order to create an adversarial statement; randomly adding the adversarial statement to a context for attacking the MLQA system; training the MLQA system with data that includes the adversarial statement in addition to the original data; enhancing the trained MLQA model by bootstrapping an adversarial policy into the trained MLQA (i.e., adding a policy on how to handle adversarial statements); and then using the trained MLQA with the bootstrapped adversarial policy to answer questions that are semantically similar to the enhanced trained MLQA model.
[0036] Recent advances in open domain question answering (QA) systems have primarily focused on machine reading comprehension (MRC), the challenge of which is to read and understand specific text and then answer questions based on that understanding. Much of the confidence in conventional techniques for obtaining state-of-the-art (SOTA) English MRC datasets stems from the invention of large-scale pre-trained language models (LMs). Conventional techniques have paid little attention to multilingual question answering.
[0037] Therefore, one or more embodiments of the present invention focus on multilingual question and answer (MLQA) systems. More specifically, one or more embodiments of the present invention address the problem of adversarial attacks against MLQA datasets (i.e., contexts / passages used by the MLQA system to answer questions) by using novel multilingual adversarial statements to train the MLQA system on how to recognize multilingual attacks using a robust MLQA model.
[0038] In one or more embodiments of the present invention, a multilingual QA model is trained using a multilingual transformer bidirectional encoder representation (e.g., MBERT) that uses transformers, as described in detail below in the example of Figure 7. A transformer is a logical mechanism that reads an entire set of words from a passage without being constrained to reading from left to right or right to left. That is, a transformer is defined as logic that identifies how different words relate to one another, as described below in steps 1 (element 402) and 2 (element 404) of the flowchart 400 shown in Figure 4.
[0039] As described below in Figure 4, a question is transformed into a corresponding statement containing placeholders for the answer, which is then used to create adversarial statements that "look" like the correct answer (due to similar terminology, passages found in the correct answer) but are not actually correct. These adversarial statements, in one or more embodiments of the present invention, include translations of adversarial statements into one or more different languages and are used to attack existing multilingual QA models and to train new multilingual QA models.
[0040] After a trained multilingual QA model is constructed, it is used by an artificial intelligence system to recognize adversarial attacks (including adversarial statements) and prevent adversarial attacks from being returned to the questioner using the QA system.
[0041] Referring here to Figure 2, a high-level overview of an exemplary attack pipeline used when training a question-answer learning system to recognize adversarial statements within a contextual passage, according to one or more embodiments of the present invention, is shown.
[0042] As shown in Figure 2, the original question and the original context (e.g., text passage, video file, etc.) that answers the original question are entered into the holding section of the question / answer (QA) system (e.g., the QA dialogue system 156 shown in Figure 1), as shown in block 202. If both the question and context are text, in one or more embodiments of the present invention, those questions and contexts are in any language.
[0043] As shown in block 204, one or more adversarial statements, which are new statements that contradict information found in the original context / passage / answer, are added to the original context / passage / answer.
[0044] In one or more embodiments of the present invention, these adversarial statements, which are patterned on the original question but contradict the information in the original context / passage / answer, are in a language different from the language used in the original question or the original context / passage / answer, or both.
[0045] In one or more embodiments of the present invention, these adversarial statements are in the same language as the original question or the original context / passage / answer, or both.
[0046] In one or more embodiments of the present invention, as described in detail below, these adversarial statements may take the form of a random answer random question (RARQ) adversarial statement, a random answer original question (RAOQ) adversarial statement, a no answer random question (NARQ) adversarial statement, or a no answer original question (NAOQ) adversarial statement, or a combination thereof, as described in detail below in Figures 3 and 4.
[0047] As shown in block 206 of Figure 2, the original context, including the added adversarial statement, is then executed against a question / answer (QA) model on an artificial intelligence (AI) system. That is, the original context, including the added adversarial statement, is used as input to an AI system that is being trained by a question / answer (QA) model to match specific types of questions (matching the parameters, terms, context, etc. of the original question) with specific types of contexts / passages / answers (matching the parameters, terms, etc. of the original context / passage / answer).
[0048] However, at this point, the system has not been trained to recognize the adversarial statement added in block 204, and therefore the output response shown in block 208 may contain erroneous information caused by the adversarial statement added in block 204.
[0049] Referring now to Figure 3, various types of adversarial passages used in one or more embodiments of the present invention are shown.
[0050] As shown in block 301, assume that the topic of the question is about the item "Paris cafes". Further assume that the original question 304 presented in the QA system is "What is the oldest cafe in Paris?". The correct / original answer to this original question is "Cafe X", which is derived from the original / correct passage / context shown in block 303 and is located at position 302 in block 303. For example, in this example, position 302 is the position of the 25th word in the original / correct passage / context shown in block 303.
[0051] However, the original / correct passage / context shown in block 303 can be modified using adversarial statements such as the adversarial statements shown in adversarial passage A (block 305), adversarial passage B (block 309), adversarial passage C (block 313), and adversarial passage D (block 317).
[0052] Adversarial statements added to an adversarial passage are created by transforming a question into a statement containing placeholders for the answer. These statements can be modified using one of the attack techniques described below, as shown in Figure 3.
[0053] Therefore, with respect to Figure 3, adversarial passage A shown in block 305 contains a random answer random question (RARQ) adversarial statement 307, which contains a random answer entity ("Corporation A"), within which a random question entity ("Arctic Ocean") replaces the correct question entity ("Paris") shown in blocks 301 / 303.
[0054] Adversarial passage B, shown in block 309, contains a random answer original question (RAOQ) adversarial statement 311, which contains a random answer entity ("Alaskan Statehood"), while the specific type of question entity ("Paris") remains the same as that from the correct answer shown in blocks 301 / 303.
[0055] Adversarial passage C shown in block 313 contains an adversarial statement 315 of a no-answer random question (NARQ) in which no answer entity is added (referred to as "_" to indicate that the word does not exist), and a random question entity ("Brooklyn") replaces the correct question entity ("Paris") found in the correct answer shown in blocks 301 / 303.
[0056] Adversarial passage D shown in block 317 contains an unanswered original question (NAOQ) adversarial statement 319, in which no answer entity is added (referred to as "_" to indicate the absence of a word), and the correct question entity ("Paris") remains the same as that from the correct answer shown in blocks 301 / 303.
[0057] As previously described in the description of block 204 in Figure 2, in one or more embodiments of the present invention, the adversarial statement is in a language other than the language of the original question or the original context / passage / answer or both. Such adversarial statements in another language are either a result of the QA system retrieving a foreign language passage or a result of the QA system translating one of the aforementioned adversarial statements. In either embodiment, block 321 shows adversarial passage A', in which the RARQ adversarial statement 307 ("Corporation A is the oldest cafe in the Arctic Ocean.") is translated into the German adversarial statement 323 ("Corporation A ist das alteste Cafe in Arktischen Ozean.") and inserted into the original / correct passage shown in block 303.
[0058] Referring now to Figure 4, an exemplary flowchart 400 of the steps used to generate the exemplary adversarial statement shown in Figure 3, according to one or more embodiments of the present invention, is shown.
[0059] As shown in block 402, in one or more embodiments of the present invention, step 1 performs a language preprocessing step on question 412 ("What is the oldest cafe in Paris?") also shown in block 301 of Figure 3. These exemplary language preprocessing steps include (1) universal dependency parsing (UDP) and (2) named entity recognition (NER). Step 1 uses markup rules and parsing to identify a tagged question entity 426 (location, e.g., Paris) in addition to a root term (e.g., element 446) that broadly identifies the type of question being asked ("what"). That is, this analysis identifies that within question 412, focus words (e.g., which, what, etc.) are generated by the parser using corresponding part of speech (POS) tags (e.g., wrb for adverbs such as "where," vb for verbs such as "is"). This analysis leads to a depth-first search in parsing, marking all POS tokens that are at the same level as the focus word or are children of the focus word as part of the question rules. This technique generates thousands of patterns within the question-answer dataset used as the training set, some of which occur only once. Some exemplary patterns include "what nn", "what vb", "who vb", "how many", and "what vb vb".
[0060] In addition, in one or more embodiments of the present invention, the system marks up all entities (e.g., words) within the question 412.
[0061] In one or more embodiments of the present invention, entities tagged by NER that are not part of the question pattern are given priority. However, if no such entities are found, the system prefers to examine nouns and then verbs to ensure better coverage.
[0062] Therefore, in the example shown in Figure 4, "what vb" is the pattern found in "What is the oldest cafe in Paris?".
[0063] As shown in block 404, in one or more embodiments of the present invention, step 2 translates question 412 into statement 414.
[0064] In one or more embodiments of the present invention, the patterns found in step 1 are used to select from multiple rules based on a catch-all for common question words {"who", "what", "when", "why", "which", "where", "how"} and any patterns that do not contain question words (these are usually due to grammatically incorrect questions or misspellings, such as "Mr. Smith's grandmother's name was?"). This rule includes question 412 ("What is the oldest cafe in Paris?") with a tagged question entity 426 ("Paris") and placeholder 424 instead of the root term "what is" (element 446). <answer> ( <answer>)) Add statement 414 (" <answer>is the oldest cafe in Paris( <answer>It is the oldest café in Paris.)
[0065] If the first question word found within the pattern is "what", then as shown in statement 414, <answer>is the oldest cafe in Paris( <answer>The rule "what vb" is used when referring to "what" (what is) as in "It is the oldest café in Paris." <answer> ( <answer>Replace with ). The answer may be added to the end of the statement. The "when vb vb" pattern triggers a rule about "when", which changes "When did Rock Band ABC release their second album?" to "Rock Band ABC released their second album in <answer>(The rock band ABC released their second album <answer>This translates to "(Released on)".
[0066] As shown in block 406, in one or more embodiments of the present invention, step 3 generates one or more adversarial statements based on different strategies. In the exemplary embodiment shown in Figure 4, given question 412 and statement 414, exemplary attack statements RARQ407 (similar to adversarial statement 307 shown in Figure 3), attack statement RAOQ411 (similar to adversarial statement 311 shown in Figure 3), attack statement NARQ415 (similar to adversarial statement 315 shown in Figure 3), and attack statement NAOQ419 (similar to adversarial statement 319 shown in Figure 3) are generated.
[0067] As shown in Figure 4, based on the attack <answer>Alternatively, RARQ407, RAOQ411, NARQ415, and NAOQ419 are generated to replace the question entity or both. In one or more embodiments of the present invention, candidate entities are randomly selected from entities found in the training data of the question answer dataset based on the entity type. The type of answer entity is selected based on the entities that the system predicts for development / test questions in a non-adversarial setting.
[0068] In one or more embodiments of the present invention, date and numerical entities are not selected from the training data of the question response dataset, but are simply generated randomly.
[0069] Candidate entities are applied to create adversarial statements using the following transformations, from the most complex to the simplest.
[0070] RARQ407 is a random answer random question adversarial / attack statement, placeholder 424( <answer>It includes a random answer entity 428 ("Corporation A") which replaces ), and question entity 430 ("Arctic Ocean") which is randomly changed from the tagged question entity 426 ("Paris") found in statement 414. Note that "Corporation A" is an incorrect answer to question 412, and this is intentional, as RARQ407 is used to train the QA system on how to recognize RARQ attack / adversarial statements.
[0071] RAOQ411, a random answer original question attack / adversarial statement, contains a random answer entity 432 ("Alaskan Statehood"), while question entity 426 ("Paris") is the same question entity 426 found in statement 414. Note that RAOQ411 is also an incorrect statement used to train a QA system on how to recognize RARQ attack / adversarial statements.
[0072] NARQ415, a no-answer random question attack / adversarial statement, does not contain an answer entity in section 436 and contains a randomly generated question entity 438 ("Brooklyn"). Note that NARQ415 is also an incorrect statement used to train a QA system on how to recognize RARQ attack / adversarial statements.
[0073] NAOQ419, an original question attack / adversarial statement with no answer, does not contain an answer entity in section 440, but does contain question entity 426 ("Paris") found in statement 414. Note that NAOQ419 is also an incorrect statement that can be used to train a QA system on how to recognize NAOQ attack / adversarial statements.
[0074] As shown in block 408, step 4 translates one or more of the attack / adversarial statements created in step 3 into another language. That is, in one or more embodiments of the present invention, the attack / adversarial statements created in step 3 are initially generated using the same language used by the question (e.g., English). Since the QA system evaluates multilingual datasets and models, these attack / adversarial statements are then translated by the QA system into several other languages if they are not already in another language when sent from the text document server 152 shown in Figure 1.
[0075] For example, to create RARQ423 (similar to the German adversarial statement 323 shown in Figure 3), the RARQ407 attack / adversarial statement is translated into German.
[0076] As shown in block 410, step 5 then randomly inserts the attack / adversarial statements created in step 3 or step 4 or both into the context (e.g., the original / correct passage / context shown in block 303 of Figure 3) to create adversarial passages A, B, C, D, and A' shown in Figure 3. That is, the generated adversarial statements (e.g., RARQ407, RAOQ411, NARQ415, NAOQ419, RARQ423, etc.) are inserted at random positions within the context, such as the original / correct passage shown in block 303 of Figure 3, which is shown as adversarial passage 425 in Figure 4. This generates new instances (Qx, Cy, Ay, Sz) where x, y, z ∈ L are the language of the question, context, and statement, respectively, and do not need to be the same, as shown in blocks 305, 309, 313, 317, and 321 of Figure 3.
[0077] The aforementioned attack / adversarial statements allow the QA system to investigate vulnerabilities in the MLQA dataset and MBERT by forcing the QA system to predict incorrect answers in multiple languages, not just one, so that the questions, context, and adversarial statements can all be in the same or different languages.
[0078] Therefore, referring here to Figure 5, an exemplary process for using a trained model in a question-answering dialogue system to defend against adversarial attacks / statements is shown according to one or more embodiments of the present invention.
[0079] As shown in block 501, the process begins by retrieving a question / answer (QA) dataset of known questions and their known correct answers (e.g., question and answer pairs such as the question and answer pairs shown in block 301 of Figure 3).
[0080] As shown in block 503 of Figure 5, attack / adversarial statements of multiple types (e.g., RARQ, RAOQ, NARQ, or NAOQ, or a combination thereof) or in multiple languages or both are added to the context / passage of the entire training dataset (question and answer pairs), as described in Figure 4.
[0081] As mentioned earlier when describing Figure 4, a QA model is created by converting question 412 into statement 414 and then associating question 412 with statement 414. This QA model is then modified to create a multilingual QA (MLQA) model, as described in block 505 of Figure 5. In one or more embodiments of the present invention, the MLQA model is created in two steps.
[0082] The first step is to intentionally contaminate / input data into passage 303 using one or more attack / adversarial statements in multiple languages, as illustrated in Figure 4, in order to create additional training data for the MLQA model.
[0083] Furthermore, in order to create new passages as additional training data for the MLQA model, one or more adversarial statements may be input into the passage multiple times using the same or different attacks.
[0084] As illustrated in Figure 5, the original question / answer / passage and the new question / answer / passage created in Figure 4 are used to retrain the original MLQA model.
[0085] The second step is to improve the retrained MLQA model by bootstrapping the adversarial policies into a retrained version of the MLQA model using attacks (i.e., adding policies on how to handle various attack / adversarial statements in different languages). This retrained MLQA model is then recursively trained by an artificial intelligence (AI) system using reinforcement learning, as shown in arrow block 506.
[0086] As shown in block 507, during each iteration, adversarial attacks are performed by the retrained MLQA, which includes multiple languages, in order to evaluate whether the newly retrained MLQA model is robust (i.e., unaffected by attacks).
[0087] In one or more embodiments of the present invention, the process shown in block 503 or block 505 or both shown in Figure 5 uses artificial intelligence such as the artificial intelligence 124 shown in Figure 1. Such artificial intelligence 124 can take various forms according to one or more embodiments of the present invention. Such forms include, but are not limited to, transformer-based reinforcement learning systems utilizing multilanguage bidirectional encoder representation from transformers (MBERT), deep neural networks (DNNs), recursive neural networks (RNNs), convolutional neural networks (CNNs), and the like.
[0088] Accordingly, in one or more embodiments of the present invention, the MBERT system described below in Figure 7 is a transformer-based system used in conjunction with enhanced learning as shown in Figure 5. That is, the combination of transformers and enhanced learning allows the system to determine which bootstrapped adversarial policies to use when deciding whether to (1) create RAOQ adversarial statements, NAOQ adversarial statements, etc., as described in Figures 3 and 4, from a context such as the exemplary passage shown in block 303 of Figure 3; (2) translate questions such as the exemplary question 412 shown in Figure 4 into another language; or (3) translate answers such as the exemplary statement 414 shown in Figure 4 into another language, or a combination thereof.
[0089] In other words, in the reinforcement learning setup of one or more embodiments of the present invention, the system (e.g., the QA dialogue system 156 shown in Figure 1) finds the best combination of one or more adversarial policies by a policy gradient algorithm such as the REINFORCE algorithm (described below), and then applies those policies to a large pool of adversarial statements, translations, etc., which may be newly created during each iteration and used to train the system's defenses.
[0090] As described herein, in one or more embodiments of the present invention, candidate contexts (e.g., one or more of the contexts / passages shown in blocks 303, 305, 309, 313, 317, 321 of Figure 3) are evaluated to determine the location of the correct answer within such contexts / passages, even if they are corrupted in some cases with adversarial conditions (e.g., elements 307, 311, 315, 319, 323 shown in Figure 3).
[0091] Referring now to Figure 6, a high-level overview of one or more embodiments of the present invention is shown.
[0092] A transformer model system 624, similar to the AI 124 shown in Figure 1 (i.e., a system that models context by using transformers as described herein), receives a question 604 (similar to question 304 shown in Figure 3) and a candidate context 600 (similar to some or all of the contexts shown in blocks 303, 305, 309, 313, 317, and 321 in Figure 3) as input. The candidate context 600 also includes a candidate answer position 602, which indicates a position within the candidate context 600 where the candidate context 600 is expected to hold the correct answer to question 604. The transformer model system 624 is trained on how to accurately identify the correct answer position from the candidate answer positions 602 using these different answer positions 602. In one or more embodiments of the present invention, as shown in block 604, the question 604, the candidate context 600, and the candidate answer positions 602 are combined into a single group. Regardless of whether the question 604, candidate context 600, and candidate answer position 602 are combined into a single group, in one or more embodiments of the present invention, a controller 601 (e.g., computer 101 shown in Figure 1) transmits various questions, candidate contexts, or candidate answer positions, or combinations thereof, to the transformer model system 624 for training the transformer model system 624, or for evaluating various questions, candidate contexts, or candidate answer positions, or combinations thereof, or both.
[0093] Referring now to Figure 7, an exemplary multilanguage bidirectional encoder representation from transformers (MBERT) system 724, as used in one or more embodiments of the present invention, is shown.
[0094] The MBERT system 724 (i.e., a training system that uses artificial intelligence to identify the position of correct answer terms within a context / passage, including contexts / passages corrupted by adversarial statements as shown in Figures 3 and 4) takes candidate context 600, candidate answer position 602, and question 604 as inputs, as described in Figure 6. These inputs are converted into embeddings (vectors). The embedding Eap (element 702) for candidate answer position 602 represents the candidate position of the correct answer within candidate context 600. Embeddings Eq1~Eqn (elements 703~705) are different vectors representing terms within question 604. Embeddings Ecc1~Eccm (elements 707~709) are different vectors representing terms within candidate context 600.
[0095] Next, node 711 (i.e., the artificial intelligence computing node) uses weights, algorithms, biases, etc. (similar to the weights, algorithms, biases, etc. described in block 911 with respect to the deep neural network 924 shown below in Figure 9) to evaluate candidate answer positions 602 as correct positions within the candidate context 600 for providing the correct answer to question 604.
[0096] Node 711 outputs a confidence level 713 that the position within the candidate context 600, starting at position 715 and ending at position 717, is accurate. This confidence level 713 is output as an answerability prediction 719 (i.e., the confidence that a particular start / end position contains an answer to question 604), as shown in the start / end position prediction 721. The answerability prediction 719 and the start / end position prediction 721 are then sent to the controller 701.
[0097] Line 723, as indicated by line 723 going to block 604, then indicates that the controller 701 uses different candidate contexts / questions / answer positions from the candidate contexts / questions / answer positions trained by the MBERT system 724. As shown in Figure 6, according to one or more embodiments of the present invention, these different candidate answer positions, questions, or candidate contexts, or combinations thereof, can be input to the MBERT system 724 collectively, individually, or both.
[0098] Referring now to Figure 8, a QA dialogue system 800 is shown that uses a transformer-based system to answer a question using the correct answer 816 from a candidate context 801 (for example, one or more of the passages shown in Figure 3).
[0099] In one or more embodiments of the present invention, a transformer (such as the transformer used by MBERT as described herein) combines tokens (e.g., words in a sentence) with a position identifier of the token's position in the sentence and a sentence identifier of the sentence to create an embedding. These embeddings are used to answer a question within a particular context, which may or may not include an adversarial statement.
[0100] Next, reinforcement systems (e.g., REINFORCE, which uses gradients such as Monte Carlo policy gradients) allow the system to learn which policies are effective in order for the MLQA model to understand that a statement is an adversarial attack.
[0101] Assume there are multiple bootstrapped adversarial policies 804 that can be used by a transformer-based reinforcement learning system 802 to understand adversarial statements (e.g., the exemplary adversarial statements shown above in Figure 4). The transformer-based reinforcement system (e.g., QA dialogue system 800) then uses a gradient-based algorithm such as the REINFORCE algorithm, which applies various adversarial policies and response positions 302 from the bootstrapped adversarial policies 804 until a suitable adversarial statement (e.g., RARQ adversarial statement 806 or its corresponding translated adversarial statement 814 or both), determined by comparison with real-world types of adversarial statements that attack the QA dialogue system 156 shown in Figure 1, is no longer considered the optimal training statement. For example, if the question statement "Cafe X is the oldest cafe in Paris" is translated into an adversarial statement (e.g., the RARQ adversarial statement "Corporation A is the oldest cafe in the Arctic Ocean") or its translated adversarial statement ("Corporation A ist das alteste Cafe im Arktischen Ozean") or both, it is indicated that one or both of these adversarial statements match the type of adversarial statement that would actually attack (or be expected to attack) the QA dialogue system 156, and these adversarial statements or their translated adversarial statements or both are then sent to the controller (e.g., computer 101 shown in Figure 1) to retrain the MLQA model (block 505 in Figure 5) and execute the attack pipeline (block 507 in Figure 5).
[0102] In one or more embodiments of the present invention, a transformer-based learning system (e.g., the transformer model system 624 shown in Figure 6) translates the correct statements into another language (the translated correct statements or the original questions or both) for use in performing the steps described in Figure 4, thus enabling the QA dialogue system 156 to process questions / statements in multiple languages.
[0103] In one or more embodiments of the present invention, the artificial intelligence 124 utilizes an electronic neural network architecture other than a transformer-based system (e.g., a transformer model system 624), such as an electronic neural network architecture found in deep neural networks (DNNs), convolutional neural networks (CNNs), or recurrent neural networks (RNNs), together with an enhanced learning system.
[0104] In a preferred embodiment, a deep neural network (DNN) is used to evaluate text / numerical data in documents from a text corpus received from the text document server 152 shown in Figure 1, while a CNN is used to evaluate images from an audio or image corpus (for example, from the audio file server 154 or video file server 160 shown in Figure 1, respectively).
[0105] CNNs are similar to DNNs in that they both utilize interconnected electronic neurons. However, CNNs differ from DNNs in that (1) CNNs include neural layers with sizes based on filter size, stride value, padding value, etc., and (2) CNNs analyze image data using a convolutional approach. CNNs get their name "convolutional" from the fact that they are based on the convolution (i.e., the mathematical operation on two functions to obtain the result) of filtering and pooling (the mathematical operation on two functions) of pixel data to produce a predicted output (to obtain the result).
[0106] RNNs are similar to DNNs in that they both utilize interconnected electronic neurons. However, RNNs have a very simple architecture in which child nodes feed to parent nodes using a weight matrix and nonlinearity (such as trigonometric functions) that is adjusted until the parent node produces the desired vector.
[0107] The logical units within an electronic neural network (DNN, CNN, or RNN) are called "neurons" or "nodes." When an electronic neural network is implemented entirely in software, each neuron / node is a separate piece of code (i.e., an instruction that performs a specific action). When an electronic neural network is implemented entirely in hardware, each neuron / node is a separate piece of hardware logic (e.g., a processor, gate array, etc.). When an electronic neural network is implemented as a combination of hardware and software, each neuron / node is a set of instructions, a piece of hardware logic, or both.
[0108] Neural networks, as the name suggests, are broadly modeled after biological neural networks (e.g., the human brain). Biological neural networks consist of a series of interconnected neurons that influence each other. For example, a synapse allows a first neuron to be electrically connected to a second neuron through the release of neurotransmitters received by the second neuron (from the first neuron). These neurotransmitters can cause the second neuron to be excited or inhibited. The patterns of excited / inhibited interconnected neurons ultimately lead to biological outcomes, including thought, muscle movement, and memory retrieval. While this explanation of biological neural networks is highly simplified, the high-level overview is that one or more biological neurons influence the behavior of one or more other bioelectrically connected biological neurons.
[0109] Electronic neural networks are similarly composed of electronic neurons. However, unlike biological neurons, electronic neurons cannot technically be "inhibitory," and are often only "excitatory" to varying degrees.
[0110] In an electronic neural network, neurons are arranged in layers known as the input layer, hidden layers, and output layer. The input layer contains neurons / nodes that receive input data and send it to a series of hidden layers of neurons, with all neurons from one layer of hidden layers interconnected with all neurons in the next layer of hidden layers. The final layer of hidden layers then outputs the computation result to the output layer, which is often one or more nodes that hold vector information.
[0111] In one or more embodiments of the present invention, a deep neural network is used to create an MLQA model for a question-answering dialogue system.
[0112] Next, referring to Figure 7, an exemplary form of a deep neural network (DNN) is shown, which is a transformer (i.e., part of the MBERT system 724) used to create and utilize an MLQA model when answering questions according to one or more embodiments of the present invention.
[0113] For illustrative purposes, we assume that the input to the transformer / DNN includes the original question 412 (e.g., "What is the oldest cafe in Paris?") and the correct answer location (e.g., the location of "(Cafe X)" within one or more candidate contexts). Such a DNN can use these inputs to create an initial QA model by aligning the answer entities (e.g., elements 446 and 424 shown in Figure 4) and the question entities (e.g., element 426 shown in Figure 4).
[0114] As shown in Figure 8, this DNN (shown as QA Dialogue System 800) includes bootstrapped adversarial policies (e.g., policies that determine how to recognize different types of attack / adversarial statements within a passage), RARQ adversarial statements 806 (examples of which are shown in Figures 3 and 4), RAOQ adversarial statements 808 (examples of which are shown in Figures 3 and 4), NARQ adversarial statements 810 (examples of which are shown in Figures 3 and 4), NAOQ adversarial statements 812 (examples of which are shown in Figures 3 and 4), as well as algorithms, rules, etc. that use translations of these adversarial statements (e.g., 423 shown in Figure 4), shown as translated adversarial statements 814 that are input into context 801. In other words, it should be understood that RARQ adversarial statement 806, RAOQ adversarial statement 808, NARQ adversarial statement 810, NAOQ adversarial statement 812, or the translated adversarial statement 814, or any combination thereof, are part of (integrated into) context 801, which are shown in different boxes in Figure 8 simply for clarity.
[0115] The algorithms and rules used in the DNN / QA dialogue system 800 can be recursively defined and improved upon in a pre-trained MLQA model.
[0116] Figure 9 shows a high-level overview of an exemplary trained deep neural network (DNN) 924 that could be used to provide the correct answer location 915 within a proposed answer context / passage 902 when responding to a new question 901.
[0117] When automatically adjusted, "backpropagation" is used to adjust the mathematical functions, output values, weights, or biases, or combinations thereof, and in backpropagation, the "gradient descent" method determines how each mathematical function, output value, weight, or bias, or combination thereof should be adjusted to provide the exact output 917. That is, the mathematical functions, output values, weights, or biases, or combinations thereof shown in block 911 of the exemplary node 909 are recursively adjusted until the expected vector values of the trained MLQA model 915 are reached.
[0118] The new question 901 (for example, "What is the oldest cafe in Madrid?") is also input to the input layer 903 along with the proposed answer context / passage 902 (provided by a question / answer database, such as the aforementioned question / answer database), and the input layer 903 processes such information before passing it to the intermediate layer 905. That is, one or more answer entities and one or more question entities within the new question 901 are used to retrieve answers (similar to the answer described by statement 414) from the QA dataset, which is used to retrieve answers of a similar type from context / passage using a similar process described above in Figures 3 to 5. In order to correctly answer the new question 1001, the DNN 924 determines one or more of these contexts / passages, while adversarial statements are ignored.
[0119] Therefore, the mathematical functions, output values, weights, and bias values of the elements shown in block 911 and found in one, more, or all of the neurons in DNN924 cause the output layer 907 to produce output 917, which includes the correct answer position 915 of the correct answer to the new question 901, including the answer found in the passage containing adversarial statements about the new question 901.
[0120] In one or more embodiments of the present invention, the correct answer position 915 is then returned to the questioner.
[0121] Therefore, in one or more embodiments of the present invention, the present invention not only searches for a specific known correct answer ("Cafe X") to a specific type of question ("What is the oldest cafe in Paris?") within a context / passage, but also searches for the correct answer location of the correct answer to a specific type of question, thus providing a system that is far more robust than a simple word lookup program.
[0122] Referring now to Figure 10, a high-level flowchart of one or more steps performed according to one or more embodiments of the present invention is shown.
[0123] Following the introductory block 1002, as shown in block 1004, a computing device (e.g., the computer 101 shown in Figure 1, or artificial intelligence 124, or QA question-answering dialogue system 156, or a combination thereof, implemented as the MBERT system 724 or DNN or both shown in Figure 7) accesses several adversarial statements (e.g., elements 307, 311, 315, 319 shown in Figure 3) that can launch an adversarial attack against the question-answering dialogue system. The question-answering dialogue system shown in Figure 1 (e.g., artificial intelligence 124 or QA question-answering dialogue system 156 or both) is a QA system designed / trained to provide correct answers to a specific type of question, such as "What is the oldest cafe in a certain city?".
[0124] As shown in block 1006, multiple adversarial statements are used in training a machine learning model (e.g., the trained MLQA model 915 shown in Figure 9).
[0125] As shown in block 1008, the computing device enhances the trained machine learning model by bootstrapping it with adversarial policies that identify multiple types of adversarial statements (e.g., bootstrapped adversarial policies 804 shown in Figure 8).
[0126] As shown in block 1010, when a computing device responds to a question submitted to the question-answering dialogue system 800 shown in Figure 8 (e.g., the MBERT system 724 shown in Figure 7), it utilizes a trained, bootstrapped machine learning model (e.g., an updated, bootstrapped, trained MLQA model) to prevent adversarial attacks.
[0127] As shown by line 1014, the process operates in a recursive manner by returning to block 1004 until it is determined that the QA dialogue system has been properly trained (for example, by exceeding a default level of the correct percentage for identifying and overcoming adversarial statements).
[0128] The flowchart ends at termination block 1012.
[0129] In one or more embodiments of the present invention, a trained bootstrapped machine learning model is tested by a computing device that performs the following: converting questions to a question-answering dialogue system into statements including placeholders for answers; randomly selecting answer entities from the answers and adding the randomly selected answer entities in place of the placeholders to generate adversarial statements; generating attacks against a trained bootstrapped machine learning model using questions and contexts / passages containing the adversarial statements; measuring the response from the trained bootstrapped machine learning model to the generated attacks; and modifying the trained bootstrapped machine learning model to improve the response level of the response to the generated attacks.
[0130] Specifically, as shown in Figures 3 to 10, the computing device translates a question to a question-answering dialogue system into a statement containing placeholders for the answer (see, for example, steps 1 and 2 in Figure 4). The computing device then randomly selects answer entities from the answer and adds the randomly selected answer entities in place of the placeholders to generate an adversarial statement (see, for example, step 3 in Figure 4). As described herein, the process randomly inputs the adversarial statement into a passage (e.g., context / passage) to create an adversarial passage. The computing device then uses the question and context / passage containing the adversarial passage to generate an attack against a trained bootstrapped machine learning model (see, for example, block 206 in Figure 2 or block 507 in Figure 5 or both) and measures the response from the trained bootstrapped machine learning model to the generated attack (e.g., by neurons in the trained DNN924 shown in Figure 9). The computing device then modifies the trained, bootstrapped machine learning model (for example, by backpropagation within the DNN924 shown in Figure 9) to improve the response level of the generated attack (i.e., to more clearly indicate the presence of the attack).
[0131] In one or more embodiments of the present invention, the adversarial statements include a first adversarial statement in a first language and a second adversarial statement in a different second language, but both the first and second adversarial statements provide the same incorrect answer to the question. For example, the first adversarial statement (e.g., RARQ307 shown in Figure 3 - "Corporation A is the oldest cafe in the Arctic Ocean") is in a first language (English), and the second adversarial statement (e.g., RARQ323 shown in Figure 3 - "Corporation A ist das alteste Cafe in Arktischen Ozean") is in a different second language (German), but both of these adversarial statements provide the same incorrect answer to the question "What is the oldest cafe in Paris?". Therefore, as described herein, the QA training system (e.g., DNN924) can handle adversarial statements in different languages.
[0132] In one or more embodiments of the present invention, a computing device generates RARQ adversarial statements, RAOQ adversarial statements, NARQ adversarial statements, or NAOQ adversarial statements, or a combination thereof (for example, by actually generating one or more of these adversarial statements).
[0133] In one or more embodiments of the present invention, a computing device retrieves RARQ adversarial statements, RAOQ adversarial statements, NARQ adversarial statements, or NAOQ adversarial statements, or a combination thereof, (for example, from a dataset that has already been created).
[0134] In one or more embodiments of the present invention, a computing device uses a generated or retrieved RARQ adversarial statement, RAOQ adversarial statement, NARQ adversarial statement, or NAOQ adversarial statement, or a combination thereof, as input to further train a machine learning model for a question-answering dialogue system to recognize adversarial statements (see Figure 6 of this patent application).
[0135] In one or more embodiments of the present invention, multiple adversarial statements are randomly placed in a single context / passage at once.
[0136] In one or more embodiments of the present invention, multiple adversarial statements are randomly and individually placed in a single context / passage, and each original context / passage containing a new adversarial statement becomes a new context / passage.
[0137] Accordingly, this specification describes a novel multilingual QA system in which the questions, context, and adversarial statements may be in the same language or different languages. Adversarial / attack statements may be generated in one language and then translated into other languages, or adversarial / attack statements may be received in different languages. In any case, the QA systems described herein utilize a single trained MLQA model capable of handling multiple languages so that the QA system's defense against attacks is effective, regardless of whether the model is zero-shot (trained on data in which the questions, context, and adversarial statements are in a different language from the test data), multilingual (trained on data in which the questions, context, and adversarial statements are in two or more different languages), or both.
[0138] In one or more embodiments, the present invention is implemented using cloud computing. Nevertheless, although this disclosure includes a detailed description of cloud computing, it should be understood in advance that implementations of the contents described herein are not limited to cloud computing environments. Embodiments of the present invention can be implemented in combination with any other type of computing environment that is currently known or may be developed in the future.
[0139] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), allowing these resources to be provisioned and released quickly with minimal administrative effort or interaction with service providers. This cloud model includes at least five features, at least three service models, and at least four deployment models.
[0140] The features are as follows:
[0141] On-demand self-service: Cloud users can unilaterally and automatically provision computing power, such as server time and network storage, as needed, without requiring human interaction with service providers.
[0142] Broad network access: The capabilities of the cloud are available over the network and accessed using standard mechanisms that facilitate use by heterogeneous thin-client or thick-client platforms (e.g., mobile phones, laptops, and PDAs).
[0143] Resource Pooling: A provider's computing resources are pooled and, using a multi-tenant model, various physical and virtual resources are dynamically allocated and reallocated as needed to serve multiple users. Users typically have a sense of location independence in that they have neither control nor know the exact location of the resources they are served, although at a higher level of abstraction they can still specify a location (e.g., country, state, or data center).
[0144] Rapid Flexibility: Cloud capabilities can be provisioned quickly and flexibly, sometimes automatically, scale out rapidly, and be released quickly to scale in rapidly. These capabilities available for provisioning often appear unlimited to the user, and any amount can be purchased at any time.
[0145] Services measured: Cloud systems automatically control and optimize resource usage by leveraging metric capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both the service providers and users of the services being utilized.
[0146] SaaS (Software as a Service): The capability provided to the user is the use of the provider's applications running on cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., web-based email). Users do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or individual application functions, except for the possibility of making limited user-specific application configuration settings.
[0147] PaaS (Platform as a Service): The ability provided to the user is to deploy applications created or acquired by the user, using programming languages and tools supported by the provider, onto a cloud infrastructure. The user does not manage or control the underlying cloud infrastructure, including the network, servers, operating system, or storage, but can control the configuration of the deployed application and, in some cases, the application hosting environment.
[0148] IaaS (Infrastructure as a Service): The capabilities provided to the user are the provisioning of processing, storage, networking, and other basic computing resources, allowing the user to deploy and run any software, including operating systems and applications. The user does not manage or control the underlying cloud infrastructure, but can control the operating system, storage, and deployed applications, and in some cases, has limited control over selected network components (e.g., host firewalls).
[0149] The deployment model is as follows:
[0150] Private Cloud: This cloud infrastructure operates solely for the benefit of an organization. In one or more embodiments, this cloud infrastructure is managed by these organizations or third parties, or resides on-premises, off-premises, or both.
[0151] Community Cloud: This cloud infrastructure is shared by multiple organizations and supports a specific community that shares common interests (e.g., missions, security requirements, policies, and compliance considerations). In one or more embodiments, this cloud infrastructure is managed by these organizations or third parties, or resides on-premises, off-premises, or both.
[0152] Public Cloud: This cloud infrastructure is made available for use by the general public or large industry groups and is owned by the organization that sells the cloud services.
[0153] Hybrid Cloud: This cloud infrastructure is a combination of two or more clouds (private, community, or public) that are linked together while maintaining their own distinct entities through standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable the portability of data and applications.
[0154] Cloud computing environments are service-oriented environments that emphasize statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure consisting of a network of interconnected nodes.
[0155] Referring now to Figure 11, an exemplary cloud computing environment 50 is shown. As illustrated, the cloud computing environment 50 includes one or more cloud computing nodes 10 that communicate with each other with local computing devices used by cloud users (e.g., a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, or an automotive computer system 54N, or a combination thereof). Furthermore, the nodes 10 communicate with each other. In one embodiment, these nodes are grouped physically or virtually (not shown) within one or more networks, such as the private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof. This allows the cloud computing environment 50 to provide an infrastructure, platform, or SaaS, or a combination thereof, that does not require cloud users to maintain resources on their local computing devices. The types of computing devices 54A to 54N shown in Figure 11 are intended for illustrative purposes only, and it is understood that the computing node 10 and the cloud computing environment 50 can communicate with any type of computer-controlled device through any type of network or network-addressable connection (e.g., a connection using a web browser) or both.
[0156] Referring now to Figure 12, a set of functional abstraction layers provided by the cloud computing environment 50 (Figure 11) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 12 are intended to be illustrative only and that embodiments of the present invention are not limited thereto. The following layers and corresponding functions are provided as illustrated:
[0157] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include a mainframe 61, servers 62, 63, and blade servers 64 based on a RISC (Reduced Instruction Set Computer) architecture, storage devices 65, and networks and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0158] The virtualization layer 70 includes an abstraction layer, in one or more embodiments, which provides examples of virtual entities such as a virtual server 71, virtual storage 72, a virtual network 73 including a virtual private network, a virtual application and operating system 74, and a virtual client 75.
[0159] For example, the management layer 80 provides the following functions: Resource provisioning 81 dynamically procures computing and other resources used to perform tasks within the cloud computing environment. Measurement and pricing 82 tracks the costs of using resources within the cloud computing environment and charges or sends invoices for the use of those resources. For example, those resources include application software licenses. Security verifies the identities of cloud users and tasks and protects data and other resources. The user portal 83 provides users and system administrators with access to the cloud computing environment. Service level management 84 allocates and manages cloud computing resources to meet the required service levels. Service Level Agreement (SLA) planning and execution 85 prepares and procures cloud computing resources in advance of anticipated future demands in accordance with the SLA.
[0160] The workload layer 90 illustrates an example of functionality in which a cloud computing environment is utilized in one or more embodiments. Examples of workloads and functionality provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and QA interaction system protection processing 96 which performs one or more of the functions of the present invention described herein.
[0161] The terms used herein are for the sole purpose of describing specific embodiments and are not intended to limit the invention. Where used herein, the singular forms “a,” “an,” and “the” are intended to include the plural form unless otherwise explicitly indicated. It will be further understood that the terms “equipped with” or “possessing” or both, as used herein, indicate the presence of a described function, integer, step, operation, element, or component, or a combination thereof, but do not exclude the presence or addition of one or more other functions, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.
[0162] All means or steps and functional elements within the following claims, corresponding structures, materials, actions, and equivalents are intended to include any structures, materials, or actions for performing a function in combination with other specifically claimed elements. The descriptions of various embodiments of the present invention are presented for illustrative and explanatory purposes, but are not intended to be exhaustive or to limit the present invention to the disclosed forms. Many changes and modifications will be apparent to those skilled in the art without departing from the scope and spirit of the present invention. Embodiments have been selected and described in order to best illustrate the principles and practical applications of the present invention and to enable those skilled in the art to understand the present invention with respect to various embodiments with various modifications suitable for a particular intended use.
[0163] In one or more embodiments of the present invention, any method described herein is implemented by using a VHDL (VHSIC Hardware Description Language) program and a VHDL chip. VHDL is an example of a design input language for field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and other similar electronic devices. Accordingly, in one or more embodiments of the present invention, any method implemented by software described herein is emulated by a hardware-based VHDL program and then applied to a VHDL chip such as an FPGA.
[0164] Therefore, by describing in detail the embodiments of the present invention in this application and by referring to the examples thereof, it will be clear that modifications and variations are possible without departing from the scope of the invention as defined in the appended claims.< / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer>
Claims
1. Accessing multiple adversarial statements that can be used by a computing device to launch an adversarial attack against a question-answering dialogue system, wherein the question-answering dialogue system is trained to provide correct answers to specific types of questions. The computing device is used to train a machine learning model for the question-answering dialogue system using the multiple adversarial statements, The computing device enhances the trained machine learning model by bootstrapping it with adversarial policies that identify multiple types of adversarial statements. The computing device, when responding to questions submitted to the question-answering dialogue system, utilizes the trained and bootstrapped machine learning model to prevent adversarial attacks. Methods that include...
2. The computing device converts the questions to the question-answering dialogue system into statements that include placeholders for the answers, The computing device randomly selects answer entities from the answers, adds the randomly selected answer entities in place of the placeholders, and generates adversarial statements. The computing device generates an attack against the trained and bootstrapped machine learning model, including the adversarial statement. The computing device measures the response of the trained and bootstrapped machine learning model to the generated attack, The computing device modifies the trained and bootstrapped machine learning model to improve the response level of the response to the generated attack. The method according to claim 1, further comprising testing the trained and bootstrapped machine learning model by means of...
3. The method according to claim 1, wherein the plurality of adversarial statements include a first adversarial statement in a first language and a second adversarial statement in a different second language, and both the first adversarial statement and the second adversarial statement provide the same incorrect answer to the question.
4. The aforementioned correct answer includes a correct answer entity and is associated with a correct question entity, and the method is The computing device generates a random answer random question (RARQ) adversarial statement, wherein the RARQ adversarial statement is a first type of attack statement, the RARQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer. The computing device generates a random answer original question (RAOQ) adversarial statement, wherein the RAOQ adversarial statement is a second type of attack statement, the RAOQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RAOQ adversarial statement includes the correct question entity from the correct answer. The computing device generates a random question (NARQ) adversarial statement with no answer, wherein the NARQ adversarial statement is a third type of attack statement, and the NARQ adversarial statement includes a random question entity that replaces the correct answer entity in the correct answer with no answer, and the NARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer. The computing device generates an unanswered original question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement is a fourth type of attack statement, the NAOQ adversarial statement replaces the correct answer entity in the correct answer with an unanswered one, and the NAOQ adversarial statement includes the correct question entity from the correct answer. The method according to claim 1, further comprising using the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as input for further training the machine learning model for the question-answering dialogue system to recognize adversarial statements within a contextual passage by the computing device.
5. The aforementioned correct answer includes a correct answer entity and is associated with a correct question entity, and the method is The retrieval of a Random Answer Random Question (RARQ) adversarial statement, wherein the RARQ adversarial statement is a first type of attack statement, the RARQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer, The extraction involves extracting a Random Answer Original Question (RAOQ) adversarial statement, wherein the RAOQ adversarial statement is a second type of attack statement, the RAOQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RAOQ adversarial statement includes the correct question entity from the correct answer. The retrieval of a No-Answer Random Question (NARQ) adversarial statement, wherein the NARQ adversarial statement is a third type of attack statement, and the NARQ adversarial statement includes a random question entity that replaces the correct answer entity in the correct answer with no answer, and the NARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer. Extracting an unanswered original question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement is a fourth type of attack statement, the NAOQ adversarial statement replaces the correct answer entity in the correct answer with no answer, and the NAOQ adversarial statement includes the correct question entity from the correct answer. The method according to claim 1, further comprising using the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as input for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements within a contextual passage by the computing device.
6. The aforementioned correct answer includes a correct answer entity and a correct question entity, and the method is Extracting a Random Answer Original Question (RAOQ) adversarial statement, wherein the RAOQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RAOQ adversarial statement includes the correct question entity from the correct answer. The method according to claim 1, further comprising using the computing device to utilize the RAOQ adversarial statements as input for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements.
7. The aforementioned correct answer includes a correct answer entity and a correct question entity, and the method is The retrieval of a No-Answer Random Question (NARQ) adversarial statement, wherein the NARQ adversarial statement includes a random question entity that replaces the correct answer entity in the correct answer with no answer, and the NARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer. The method according to claim 1, further comprising using the computing device to utilize the NARQ adversarial statements as input for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements.
8. The aforementioned correct answer includes a correct answer entity and a correct question entity, and the method is Extracting an unanswered original question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement replaces the correct answer entity in the correct answer with an unanswered statement, and the NAOQ adversarial statement includes the correct question entity from the correct answer. The method according to claim 1, further comprising using the computing device to utilize the NAOQ adversarial statements as input for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements.
9. A computer program product comprising a computer-readable storage medium on which program instructions are embodied, wherein the computer-readable storage medium itself is not a transient signal, the program code is readable and executable by the processor, and a method is performed to prevent adversarial attacks against a question-answering dialogue system, the method is Accessing multiple adversarial statements that can be used to launch an adversarial attack against a question-answering dialogue system, wherein the question-answering dialogue system is trained to provide correct answers to specific types of questions. Using the aforementioned multiple adversarial statements, a machine learning model for the question-answering dialogue system is trained. The trained machine learning model is enhanced by bootstrapping an adversarial policy that identifies multiple types of adversarial statements into the trained machine learning model, When responding to questions submitted to the aforementioned question-answering dialogue system, the trained and bootstrapped machine learning model is used to prevent adversarial attacks. Computer program products, including [the following].
10. The method described above is Converting a question to the aforementioned question-answering dialogue system into a statement containing placeholders for the answer, The process involves randomly selecting answer entities from the aforementioned answers, adding the randomly selected answer entities in place of the placeholders, and generating adversarial statements. The process involves randomly inputting the aforementioned adversarial statements into a passage to create an adversarial passage, To generate an attack against the trained and bootstrapped machine learning model, including the adversarial passage, Measuring the response of the trained and bootstrapped machine learning model to the generated attack, Testing the trained and bootstrapped machine learning model by modifying the trained and bootstrapped machine learning model using a computer device to improve the response level of the response to the generated attack. The computer program product according to claim 9, further comprising:
11. The computer program product according to claim 9, wherein the plurality of adversarial statements include a first adversarial statement in a first language and a second adversarial statement in a different second language, and both the first adversarial statement and the second adversarial statement provide the same incorrect answer to the question.
12. The aforementioned correct answer includes a correct answer entity and is associated with a correct question entity, and the method is Generating a Random Answer Random Question (RARQ) adversarial statement, wherein the RARQ adversarial statement is a first type of attack statement, the RARQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer, Generating a Random Answer Original Question (RAOQ) adversarial statement, wherein the RAOQ adversarial statement is a second type of attack statement, the RAOQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RAOQ adversarial statement includes the correct question entity from the correct answer. Generating a No-Answer Random Question (NARQ) adversarial statement, wherein the NARQ adversarial statement is a third type of attack statement, and the NARQ adversarial statement includes a random question entity that replaces the correct answer entity in the correct answer with no answer, and the NARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer. To generate an unanswered original question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement is a fourth type of attack statement, the NAOQ adversarial statement replaces the correct answer entity in the correct answer with no answer, and the NAOQ adversarial statement includes the correct question entity from the correct answer, The RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements are used by the computing device as input to further train the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements within a contextual passage. The computer program product according to claim 9, further comprising:
13. The aforementioned correct answer includes a correct answer entity and is associated with a correct question entity, and the method is The retrieval of a Random Answer Random Question (RARQ) adversarial statement, wherein the RARQ adversarial statement is a first type of attack statement, the RARQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer, The extraction involves extracting a Random Answer Original Question (RAOQ) adversarial statement, wherein the RAOQ adversarial statement is a second type of attack statement, the RAOQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RAOQ adversarial statement includes the correct question entity from the correct answer. The retrieval of a No-Answer Random Question (NARQ) adversarial statement, wherein the NARQ adversarial statement is a third type of attack statement, and the NARQ adversarial statement includes a random question entity that replaces the correct answer entity in the correct answer with no answer, and the NARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer. Extracting an unanswered original question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement is a fourth type of attack statement, the NAOQ adversarial statement replaces the correct answer entity in the correct answer with no answer, and the NAOQ adversarial statement includes the correct question entity from the correct answer. The RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements are used by the computing device as input to further train the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements within a contextual passage. The computer program product according to claim 9, further comprising:
14. The aforementioned correct answer includes a correct answer entity and a correct question entity, and the method is Extracting a Random Answer Original Question (RAOQ) adversarial statement, wherein the RAOQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RAOQ adversarial statement includes the correct question entity from the correct answer. The aforementioned RAOQ adversarial statements are used as input to further train the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements. The computer program product according to claim 9, further comprising:
15. The aforementioned correct answer includes a correct answer entity and a correct question entity, and the method is The retrieval of a No-Answer Random Question (NARQ) adversarial statement, wherein the NARQ adversarial statement includes a random question entity that replaces the correct answer entity in the correct answer with no answer, and the NARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer. The aforementioned NARQ adversarial statements are used as input to further train the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements. The computer program product according to claim 9, further comprising:
16. The aforementioned correct answer includes a correct answer entity and a correct question entity, and the method is Extracting an unanswered original question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement replaces the correct answer entity in the correct answer with an unanswered statement, and the NAOQ adversarial statement includes the correct question entity from the correct answer. The aforementioned NAOQ adversarial statements are used as input for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements. The computer program product according to claim 9, further comprising:
17. The computer program product according to claim 9, wherein the program code is provided as a service in a cloud environment.
18. A computer system comprising one or more processors, one or more computer-readable memories, and one or more computer-readable non-transient storage media, wherein program instructions are stored in at least one of the one or more computer-readable non-transient storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, and the stored program instructions Accessing multiple adversarial statements that can be used to launch an adversarial attack against a question-answering dialogue system, wherein the question-answering dialogue system is trained to provide correct answers to specific types of questions. Using the aforementioned multiple adversarial statements, a machine learning model for the question-answering dialogue system is trained. The trained machine learning model is enhanced by bootstrapping an adversarial policy that identifies multiple types of adversarial statements into the trained machine learning model, When responding to questions submitted to the aforementioned question-answering dialogue system, the trained and bootstrapped machine learning model is used to prevent adversarial attacks. A computer system that is used to perform a method that includes [a specific action].
19. The aforementioned correct answer includes a correct answer entity and is associated with a correct question entity, and the method is The retrieval of a Random Answer Random Question (RARQ) adversarial statement, wherein the RARQ adversarial statement is a first type of attack statement, the RARQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer, The extraction involves extracting a Random Answer Original Question (RAOQ) adversarial statement, wherein the RAOQ adversarial statement is a second type of attack statement, the RAOQ adversarial statement includes a random answer entity that replaces the correct answer entity in the correct answer, and the RAOQ adversarial statement includes the correct question entity from the correct answer. The retrieval of a No-Answer Random Question (NARQ) adversarial statement, wherein the NARQ adversarial statement is a third type of attack statement, and the NARQ adversarial statement includes a random question entity that replaces the correct answer entity in the correct answer with no answer, and the NARQ adversarial statement includes a random question entity that replaces the correct question entity in the correct answer. Extracting an unanswered original question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement is a fourth type of attack statement, the NAOQ adversarial statement replaces the correct answer entity in the correct answer with no answer, and the NAOQ adversarial statement includes the correct question entity from the correct answer. The computer system according to claim 18, further comprising using the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as input for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements within a contextual passage by the computing device.
20. The computer system according to claim 18, wherein the stored program instructions are provided as a service in a cloud environment.