Training a Question-Answering Dialogue System to Prevent Adversarial Attacks
By training question-answering dialogue systems with adversarial statements and applying a bootstrapped policy, the model enhances its ability to recognize and defend against attacks, ensuring accurate responses in multilingual contexts.
Patent Information
- Application Number
- JP2023521038
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-21
- Filing Date
- 2021-08-30
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-08-30
AI Technical Summary
Question-answering dialogue systems are vulnerable to adversarial attacks that cause them to provide incorrect answers, and existing technologies have not adequately addressed this issue, particularly in multilingual contexts.
A method is employed to train a machine learning model by using adversarial statements to enhance its ability to recognize and defend against attacks, involving the generation of Random Answer Random Question (RARQ), Random Answer Original Question (RAOQ), No Answer Random Question (NARQ), and No Answer Original Question (NAOQ) statements in multiple languages, and applying a bootstrapped adversarial policy to improve the model's robustness.
The trained model effectively prevents adversarial attacks by recognizing and mitigating incorrect responses, enhancing the system's ability to provide accurate answers across different languages, thereby improving the resilience of multilingual question-answering dialogue systems.
Smart Images

Figure 0007807441000001 
Figure 0007807441000002 
Figure 0007807441000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of question-answering dialogue systems used to answer questions, and more particularly to the field of protecting such dialogue systems from adversarial attacks that corrupt them. Summary of the Invention
[0002] In one or more embodiments of the present invention, a method protects a question-answering dialogue system from being attacked by adversarial statements that incorrectly answer questions. A computing device accesses a plurality of adversarial statements capable of launching an adversarial attack against a question-answering dialogue system trained to provide correct answers to a particular type of question. The computing device utilizes the plurality of adversarial statements to train a machine learning model for the question-answering dialogue system. The computing device then enhances the trained machine learning model by bootstrapping an adversarial policy that identifies the plurality of types of adversarial statements onto the trained machine learning model. The computing device then utilizes the trained, bootstrapped machine learning model to prevent the adversarial attack when responding to questions submitted to the question-answering dialogue system.
[0003] In one or more embodiments of the present invention, the trained and bootstrapped machine learning model is tested by a computing device that performs the following operations: converting questions for a question-answering dialogue system into statements that include placeholders for answers; randomly selecting answer entities from the answers and adding the randomly selected answer entities in place of the placeholders to generate adversarial statements; generating attacks against the trained and bootstrapped machine learning model that include the adversarial statements; measuring responses from the trained and bootstrapped machine learning model to the generated attacks; and modifying the trained and bootstrapped machine learning model to improve its level of response to the generated attacks.
[0004] In one or more embodiments of the invention, the context passage includes a correct answer that includes a correct answer entity, and the particular type of question includes a particular type of question entity, and the method includes, in a computing device, generating / retrieving Random Answer Random Question (RARQ) adversarial statements, the RARQ adversarial statements including a random answer entity that replaces a correct answer entity in the correct answer, the RARQ adversarial statements including a random question entity that replaces a correct question entity in the correct answer; generating / retrieving Random Answer Original Question (RAOQ) adversarial statements, the RAOQ adversarial statements including a random answer entity that replaces a correct answer entity in the correct answer, the RAOQ adversarial statements including a correct question entity from the correct answer; and generating / retrieving No Answer Random Question (NARQ) adversarial statements. generating / retrieving No Answer Original Question (NAOQ) adversarial statements, where the NAOQ adversarial statements replace the correct answer entity in the correct answer with no answer and the NAOQ adversarial statements include a random question entity that replaces the correct question entity in the correct answer; generating / retrieving No Answer Original Question (NAOQ) adversarial statements, where the NAOQ adversarial statements replace the correct answer entity in the correct answer with no answer and the NAOQ adversarial statements include the correct question entity from the correct answer; and utilizing the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as inputs for further training a machine learning model for the question-answering dialogue system to recognize the adversarial statements.
[0005] In one or more embodiments of the present invention, the original questions used in the question-answering dialogue system, the original context passages used in the question-answering dialogue system, and / or the adversarial statements generated for the question-answering dialogue system are in one or more different languages so that the question-answering dialogue system can handle adversarial attacks in multiple languages.
[0006] In one or more embodiments, the methods described herein are performed by execution of a computer program product and / or a computer system. [Brief explanation of the drawings]
[0007] [Figure 1] 1 illustrates an exemplary system and network in which the present invention may be implemented in various embodiments. [Figure 2] FIG. 1 illustrates a high-level overview of an exemplary attack pipeline used when running a question answering (QA) dialogue / learning system that includes adversarial statements in context passages, in accordance with one or more embodiments of the present invention. [Figure 3] FIG. 1 illustrates various types of adversarial passages used in one or more embodiments of the present invention. [Figure 4] FIG. 1 illustrates an exemplary flow of steps used to generate adversarial statements in one or more embodiments of the present invention. [Figure 5] FIG. 1 illustrates an exemplary process for using a trained model in a question-answering dialogue system to defend against adversarial statements / attacks, in accordance with one or more embodiments of the present invention. [Figure 6] FIG. 1 illustrates a high-level overview of recursive training of a Transformer model system in accordance with one or more embodiments of the present invention. [Figure 7]FIG. 7 illustrates an example embodiment of the Transformer model system shown in FIG. 6 using a multilanguage bidirectional encoder representation from transformers (e.g., MBERT) in accordance with one or more embodiments of the present invention. [Figure 8] FIG. 1 illustrates an exemplary question-and-answer dialogue system utilized in one or more embodiments of the present invention. [Figure 9] FIG. 9 illustrates an exemplary deep neural network used by the QA dialogue system 800 shown in FIG. 8 to respond to new questions, in accordance with one or more embodiments of the present invention. [Figure 10] FIG. 1 depicts a high-level flowchart of one or more steps performed by a method in accordance with one or more embodiments of the present invention. [Figure 11] FIG. 1 illustrates a cloud computing environment in accordance with one or more embodiments of the present invention. [Figure 12] FIG. 1 illustrates abstraction model layers of a cloud computing environment in accordance with one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0008] In one or more embodiments, the present invention is a system, method, and / or computer program product, at any possible level of technical detail of integration. In one or more embodiments, the computer program product includes a computer-readable storage medium containing computer-readable program instructions for causing a processor to perform aspects of the present invention.
[0009] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device, such as, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or ridge-in-groove structures on which instructions are recorded, and any suitable combination thereof. As used herein, computer-readable storage media should not itself be construed as being ephemeral signals such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through fiber optic cable), or electrical signals transmitted over wires.
[0010] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or storage device over a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). This network may comprise copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface within each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage on a computer-readable storage medium within the respective computing / processing device.
[0011] In one or more embodiments, the computer-readable program instructions for carrying out the operations of the present invention include assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java®, Smalltalk®, C++, and traditional procedural programming languages such as the “C” programming language or similar programming languages. In one or more embodiments, the computer-readable program instructions execute entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and on a remote computer, or entirely on the remote computer or server. In the latter scenario, in one or more embodiments, the remote computer is connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection is to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, to carry out aspects of the present invention, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), executes computer-readable program instructions by utilizing state information of the computer-readable program instructions to customize the electronic circuitry.
[0012] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0013] In one or more embodiments, these computer-readable program instructions are provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, the instructions of which execute via the processor of the computer or other programmable data processing apparatus to create means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, to create a machine. In one or more embodiments, these computer-readable program instructions are stored on a computer-readable storage medium such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture containing instructions that implement aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams, and in one or more embodiments, direct a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner.
[0014] In one or more embodiments, the computer-readable program instructions are also loaded into a computer, other programmable data processing apparatus, or other device to create a computer-implemented process, causing a series of operable steps to be performed on the computer, other programmable apparatus, or other device that produces a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus, or other device, perform the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0015] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram represents a module, segment, or portion of instructions, comprising one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or in the reverse order, depending on the functionality involved. It should also be noted that, in one or more embodiments of the present invention, each block in the block diagrams and / or flowchart diagrams, and combinations of blocks included in the block diagrams and / or flowchart diagrams, are implemented by a special-purpose hardware-based system that performs the specified function or operation or executes a combination of special-purpose hardware and computer instructions.
[0016] Referring now to the figures, and in particular to Figure 1, there is shown a block diagram of an exemplary system and network that may be utilized by and / or in an implementation of the present invention. Note that some or all of the exemplary architecture, including both the depicted hardware and software, shown with respect to and within computer 101 may be utilized by one or more of the neurons / nodes shown in artificial intelligence 124, software deployment server 150, text document server 152, audio file server 154, question answering dialogue system 156, question submission system 158, or video file server 160, or combinations thereof, shown in Figure 1, controller 601 shown in Figure 6, Multilingual Transformer Bidirectional Encoder Representation (e.g., MBERT) system 724 shown in Figure 7, or deep neural network 924 depicted in Figure 9, or combinations thereof.
[0017] The exemplary computer 101 includes a processor 104 coupled to a system bus 106. The processor 104 may utilize one or more processors, each including one or more processor cores. A video adapter 108, which drives / supports a display 110, is also coupled to the system bus 106. The system bus 106 is coupled to an input / output (I / O) bus 114 via a bus bridge 112. An I / O interface 116 is coupled to the I / O bus 114. The I / O interface 116 provides communication with various I / O devices, including a keyboard 118, a mouse 120, a media tray 122 (which may include storage devices such as a CD-ROM drive, a multimedia interface, etc.), an artificial intelligence 124, and an external USB port 126. While the type of ports connected to the I / O interface 116 can be any type known to those skilled in the art of computer architecture, in one embodiment, some or all of the ports are universal serial bus (USB) ports.
[0018] As shown, computer 101 may also communicate with artificial intelligence 124, software deployment server 150, text document server 152, audio file server 154, question-and-answer dialogue system 156, question submission system 158, video file server 160, or any combination thereof, using network interface 130 with network 128. Network interface 130 is a hardware network interface, such as a network interface card (NIC). Network 128 can be an external network, such as the Internet, or an internal network, such as an Ethernet or virtual private network (VPN). One or more examples of physical device 154 are presented below.
[0019] A hard drive interface 132 is also coupled to the system bus 106. The hard drive interface 132 interfaces with a hard drive 134. In one embodiment, the hard drive 134 inputs data to a system memory 136, which is also coupled to the system bus 106. The system memory is defined as the lowest level of volatile memory within the computer 101. This volatile memory includes additional higher levels of volatile memory (not shown), including, but not limited to, cache memory, registers, and buffers. Data input to the system memory 136 includes the operating system (OS) 138 and application programs 144 of the computer 101.
[0020] The OS 138 includes a shell 140 for providing transparent user access to resources, such as application programs 144. The shell 140 is generally a program that provides an interpreter and interface between a user and the operating system. More specifically, the shell 140 executes commands entered into a command-line user interface or from a file. Thus, the shell 140 (also called a command processor) is typically the highest software layer of an operating system and functions as a command interpreter. The shell provides a system prompt, interprets commands entered via a keyboard, mouse, or other user input medium, and sends the interpreted commands to the appropriate lower-level operating system (e.g., kernel 142) for processing. While the shell 140 is a text-based, line-oriented user interface, the present invention adequately supports other user interface modes, such as graphics, speech, and gestures.
[0021] As shown in the figure, OS 138 also includes a kernel 142 that contains the lower level functions of OS 138, including providing essential services required by other parts of OS 138 and application programs 144 (including memory management, process and task management, disk management, and mouse and keyboard management).
[0022] Application programs 144 include a renderer illustratively shown as browser 146. Browser 146 includes program modules and instructions that enable a World Wide Web (WWW) client (i.e., computer 101) to send and receive network messages to the Internet using hypertext transfer protocol (HTTP) messaging, thus enabling communication with software deployment server 150 and other computer systems.
[0023] The application program 144 in the system memory of the computer 101 (and the system memory of the software deployment server 150) also includes question answering dialog system protection logic (QADSPL) 148. QADSPL 148 includes code for implementing the processes described below, including those described in Figures 2-10. In one embodiment, the computer 101 can download QADSPL 148 from the software deployment server 150, which includes an on-demand download, in which the code for QADSPL 148 is not downloaded until needed for execution. Furthermore, in one embodiment of the present invention, the software deployment server 150 performs all functionality related to the present invention (including executing QADSPL 148), and therefore the computer 101 does not need to use its own internal computing resources to execute QADSPL 148.
[0024] The text document server 152 is a server that transmits context (i.e., a text passage, such as the text passage shown in FIG. 3) to the computer 101, the AI 124, and / or the QA question-answering dialogue system 156 by matching a particular type of question (received by the computer 101, the AI 124, and / or the QA question-answering dialogue system 156) with a particular set of candidate answer texts.
[0025] Audio file server 154 is a server that transmits context (i.e., audio files) to computer 101, AI 124, and / or QA question-answering dialogue system 156 by matching a particular type of question (received by computer 101, AI 124, and / or QA question-answering dialogue system 156) with a particular set of candidate answer audio files. That is, audio file server 154 interprets the type of question received and returns relevant audio files (e.g., identified by metadata describing each audio file) having subject matter that matches that type of question. For example, if the question is about a particular type of music, audio file server 154 returns audio files that contain metatags that describe that particular type of music.
[0026] QA dialogue system 156 is a system that utilizes the processes / systems described herein to respond to questions (eg, from question submission system 158) with answers.
[0027] Video file server 160 is a server that transmits context (i.e., video files) to computer 101, AI 124, and / or QA question-answering dialogue system 156 by matching a particular type of question (received by computer 101, AI 124, and / or QA question-answering dialogue system 156) with a particular set of candidate answer video files. That is, video file server 160 interprets the type of question received and returns relevant video files (e.g., identified by metadata describing each video file) that have subject matter that matches that type of question. For example, if the question is about a particular type of visual art, video file server 160 returns video files that contain metatags that describe that particular type of visual art.
[0028] It should be noted that the hardware elements illustrated in computer 101 are not intended to be exhaustive, but rather are representative examples to highlight the essential components required by the present invention. For example, computer 101 may include alternative memory storage devices such as magnetic cassettes, digital versatile disks (DVDs), Bernoulli cartridges, etc. These and other variations are intended to be within the scope of the present invention.
[0029] Question-answering (QA) systems, also known as question-answering dialogue systems, are important tools used by people seeking answers. An exemplary QA system receives a question (e.g., "What is the oldest cafe in Paris?"), searches a corpus of resources such as text, video, and audio, and returns the correct answer (e.g., "Cafe X").
[0030] Therefore, such a QA system is preferably robust to ensure that it can provide the correct answer to the user: a QA system is weak if it fails against a malicious attack (described in detail below), and robust if it can successfully defend against malicious attacks (as described and claimed in one or more embodiments of the present invention).
[0031] Thus, one or more embodiments of the present invention provide a robust QA system that not only defends itself against malicious attacks, but is also able to handle multilingual malicious attacks.
[0032] As described herein, one or more embodiments of the present invention utilize one or more new types of adversarial statements to expose weaknesses in multilingual question answer (MLQA) systems.
[0033] These new kinds of adversarial statements are used to train QA models, thus making the trained QA models more robust in fighting malicious attacks.
[0034] In one or more embodiments of the present invention, the trained QA model is enhanced by bootstrapping an adversarial policy (e.g., a policy that describes which new types of adversarial classes should be monitored), thereby creating an even more effective QA model for training an MLQA system.
[0035] Thus, in one or more embodiments of the present invention, a method / apparatus generates attack statements in any language of an MLQA system by converting an original question into a general statement by using placeholders for answers; randomly selecting various entities to replace question entities and / or answer entities found in the original question to create adversarial statements; randomly adding the adversarial statements to a context for attacking the MLQA system; training the MLQA system using data that includes the adversarial statements in addition to the original data; augmenting the trained MLQA model by bootstrapping an adversarial policy to the trained MLQA (i.e., appending a policy on how to handle adversarial statements); and then using the trained MLQA with the bootstrapped adversarial policy to answer questions that are semantically similar to the augmented trained MLQA model.
[0036] Recent advances in open domain question answering (QA) systems have primarily revolved around machine reading comprehension (MRC), where the MRC challenge is to read and understand a given text and then answer questions based on it. Much of the reliance in the prior art for obtaining state-of-the-art English MRC datasets is due to the invention of large-scale pre-trained language models (LMs). Little attention has been paid to multilingual question answering in the prior art.
[0037] As such, one or more embodiments of the present invention focus on multilingual QA (MLQA) systems. More particularly, one or more embodiments of the present invention address the problem of adversarial attacks on MLQA datasets (i.e., the contexts / passages used by MLQA systems to answer questions) by using novel multilingual adversarial statements to train MLQA systems on how to recognize the multilingual attacks using robust MLQA models.
[0038] In one or more embodiments of the present invention, a multilingual QA model is trained using a multilingual transformer bidirectional encoder representation (e.g., MBERT) that uses a transformer, as described in detail below in the example of Figure 7. A transformer is a logical mechanism that reads an entire sequence of words from a passage without being constrained to reading from left to right or right to left. That is, a transformer is defined as logic that identifies how various words are related to each other, as described below in step 1 (element 402) and step 2 (element 404) of flowchart 400 shown in Figure 4.
[0039] As described below in Figure 4, questions are converted into corresponding statements containing placeholders for answers, which are then used to create adversarial statements that "look" like the correct answer (due to similar terms, passages found in the correct answer) but are not. These adversarial statements, in one or more embodiments of the invention, include translations of the adversarial statements translated into one or more different languages, and are used to attack existing multilingual QA models and train new multilingual QA models.
[0040] After the trained multilingual QA model is constructed, it is used by an artificial intelligence system to recognize adversarial attacks (containing adversarial statements) and prevent them from being returned to questioners using the QA system.
[0041] Referring now to FIG. 2, there is shown a high-level overview of an exemplary attack pipeline used when training a question-answering learning system to recognize adversarial statements within a context passage, in accordance with one or more embodiments of the present invention.
[0042] 2, an original question and an original context (e.g., a text passage, a video file, etc.) that answers the original question are entered into a holding section of a question / answer (QA) system (e.g., QA dialogue system 156 shown in FIG. 1), as shown in block 202. While the question and context are both textual, in one or more embodiments of the present invention, the question and context may be in any language.
[0043] As shown in block 204, one or more adversarial statements, which are new statements that contradict information discovered in the original context / passage / answer, are added to the original context / passage / answer.
[0044] In one or more embodiments of the present invention, these adversarial statements that are patterned on the original question but contradict information in the original context / passage / answer are in a language that is different from the language used in the original question and / or the original context / passage / answer.
[0045] In one or more embodiments of the present invention, these adversarial statements are in the same language as the original question and / or the original context / passage / answer.
[0046] In one or more embodiments of the present invention, as described in detail below, these adversarial statements are in the form of random answer random question (RARQ) adversarial statements, random answer original question (RAOQ) adversarial statements, no answer random question (NARQ) adversarial statements, or no answer original question (NAOQ) adversarial statements, or combinations thereof, as described in detail below in Figures 3 and 4.
[0047] As shown in block 206 of Figure 2, the original context with the added adversarial statements is then run against a question / answer (QA) model on an artificial intelligence (AI) system. That is, the original context with the added adversarial statements is used as input to an AI system that is trained by a question / answer (QA) model to match specific types of questions (matching the original question's parameters, terminology, context, etc.) with specific types of contexts / passages / answers (matching the original context / passages / answer's parameters, terminology, etc.).
[0048] However, at this point, the system has not been trained to recognize the adversarial statements added in block 204, and therefore the output answer shown in block 208 may contain erroneous information caused by the adversarial statements added in block 204.
[0049] Referring now to FIG. 3, various types of adversarial passages are shown that may be used in one or more embodiments of the present invention.
[0050] Assume that the topic of the question is about the item "Paris cafes," as shown in block 301. Assume further that the original question 304 being posed to the QA system is "What is the oldest cafe in Paris?" The correct / original answer to this original question is "Cafe X," derived from the original / correct passage / context shown in block 303, and is located at position 302 in block 303. For instance, in this example, position 302 is the 25th word position in the original / correct passage / context shown in block 303.
[0051] However, the original / correct passage / context shown in block 303 may be modified using adversarial statements such as those shown in adversarial passage A (block 305), adversarial passage B (block 309), adversarial passage C (block 313), and adversarial passage D (block 317).
[0052] The adversarial statements added to the adversarial passage are created by transforming the question into a statement containing a placeholder for the answer. The statement can be modified using one of the attack techniques described below, as shown in Figure 3.
[0053] Thus, with reference to FIG. 3, adversarial passage A shown in block 305 includes a random answer random question (RARQ) adversarial statement 307 that includes a random answer entity ("Corporation A") in which a random question entity ("Arctic Ocean") replaces the correct question entity ("Paris") shown in blocks 301 / 303.
[0054] Adversarial passage B, shown in block 309, includes a random answer original question (RAOQ) adversarial statement 311, which includes a random answer entity ("Alaskan Statehood") and keeps the specific type of question entity ("Paris") from the correct answer shown in blocks 301 / 303 the same.
[0055] Adversarial passage C, shown in block 313, includes a no-answer random question (NARQ) adversarial statement 315 in which no answer entity is added (referred to as "_" to indicate the absence of a word) and a random question entity ("Brooklyn") replaces the correct question entity ("Paris") found in the correct answer shown in blocks 301 / 303.
[0056] Adversarial passage D, shown in block 317, contains a No Answer Original Question (NAOQ) adversarial statement 319 in which no answer entity is added (referred to as "_" to indicate the absence of the word) and the correct question entity ("Paris") from the correct answer shown in blocks 301 / 303 remains the same.
[0057] As previously discussed in the description of block 204 of FIG. 2 , in one or more embodiments of the present invention, the adversarial statements are in a language other than the language of the original question and / or the original context / passage / answer. Such adversarial statements in other languages may result from the QA system retrieving a foreign language passage or from the QA system translating one of the aforementioned adversarial statements. In either embodiment, block 321 shows adversarial passage A′, in which RARQ adversarial statement 307 (“Corporation A is the oldest cafe in the Arctic Ocean.”) is translated into German adversarial statement 323 (“Corporation A ist das alteste Cafe in Arktischen Ozean.”) and inserted into the original / correct passage shown in block 303.
[0058] Referring now to FIG. 4, there is shown an example flowchart 400 of steps used to generate the example adversarial statements shown in FIG. 3 in accordance with one or more embodiments of the present invention.
[0059] As shown in block 402, in one or more embodiments of the present invention, step 1 performs linguistic preprocessing steps on question 412 (“What is the oldest cafe in Paris?”), also shown in block 301 of FIG. 3 . These exemplary linguistic preprocessing steps include (1) universal dependency parsing (UDP) and (2) named entity recognition (NER). Step 1 uses markup rules and syntactic analysis to identify tagged question entities 426 (locations, e.g., Paris) in addition to root terms (e.g., element 446) that broadly identify the type of question being asked (“what”). That is, this analysis identifies focus words (e.g., which, what, etc.) within question 412 that use corresponding part-of-speech (POS) tags (e.g., wrb for adverbs such as “where” and vb for verbs such as “is”) to be generated by the parser. This analysis leads to a depth-first search in the syntactic analysis, marking all POS tokens that are at the same level as the focus word or that are children of the focus word as part of the question rule. This technique creates thousands of patterns in the question-answering dataset used as a training set, some of which occur only once. Some example patterns include "what nn," "what vb," "who vb," "how many," and "what vb vb."
[0060] Additionally, in one or more embodiments of the present invention, the system marks up all entities (eg, words) within the question 412 .
[0061] In one or more embodiments of the present invention, priority is given to entities tagged by NER that are not part of the query pattern, but if no such entities are found, the system preferably looks at nouns and then verbs to ensure better coverage.
[0062] So in the example shown in Figure 4, "what vb" is a pattern discovered in "What is the oldest cafe in Paris?"
[0063] As shown in block 404 , in one or more embodiments of the present invention, step 2 converts questions 412 into statements 414 .
[0064] In one or more embodiments of the invention, the pattern discovered in step 1 is used to select from multiple rules based on a catch-all for common question words (“who,” “what,” “when,” “why,” “which,” “where,” “how”) and any patterns that do not contain question words (these usually result from grammatically incorrect questions or misspellings, such as “Mr. Smith’s grandmother’s name was?”). This rule selects question 412 (“What is the oldest cafe in Paris?”) containing tagged question entity 426 (“Paris”) with placeholder 424 ( ) instead of the root term “what is” (element 446). <answer> ( <answer>)) and add statement 414 (" <answer>is the oldest cafe in Paris( <answer>is the oldest cafe in Paris)
[0065] If the first question word found in the pattern is "what", then the " <answer>is the oldest cafe in Paris( <answer>The rule "what vb" can be used to translate "what" into ", such as "is the oldest cafe in Paris." <answer> ( <answer>)". The answer may be added to the end of the statement. The "when vb vb" pattern triggers the "when" rule, which changes "When did Rock Band ABC release their second album?" to "Rock Band ABC released their second album in <answer>(Rock band ABC released their second album <answer>(released on ).
[0066] As shown in block 406, in one or more embodiments of the present invention, step 3 generates one or more adversarial statements based on different strategies. In the exemplary embodiment shown in Figure 4, given question 412 and statement 414, exemplary attack statement RARQ407 (similar to adversarial statement 307 shown in Figure 3), attack statement RAOQ411 (similar to adversarial statement 311 shown in Figure 3), attack statement NARQ415 (similar to adversarial statement 315 shown in Figure 3), and attack statement NAOQ419 (similar to adversarial statement 319 shown in Figure 3) are generated.
[0067] Based on the attack, as shown in Figure 4 <answer>RARQ407, RAOQ411, NARQ415, and NAOQ419 are generated to replace the question entity or both. In one or more embodiments of the present invention, candidate entities are randomly selected from entities discovered in the training data of the question-answering dataset based on entity type. The answer entity type is selected based on entities the system predicts for development / test questions in a non-adversarial setting.
[0068] In one or more embodiments of the present invention, the date and numeric entities are not selected from the training data of the question-answering dataset, but are simply randomly generated.
[0069] Candidate entities are applied to create adversarial statements using the following transformations from most complex to simplest:
[0070] Random Answer Random Question The hostile / attack statement RARQ407 is placed inside the placeholder 424 ( <answer>), and question entity 430 ("Arctic Ocean") is randomly changed from tagged question entity 426 ("Paris") found in statement 414. Note that "Corporation A" is an incorrect answer to question 412, and this is intentional, as RARQ 407 is used to train QA systems on how to recognize RARQ attacks / adversarial statements.
[0071] Random Answer The original question attack / adversarial statement, RAOQ411, contains a random answer entity 432 ("Alaskan Statehood"), but question entity 426 ("Paris") is the same question entity 426 found in statement 414. Note that RAOQ411 is also an incorrect statement used to train a QA system on how to recognize RARQ attacks / adversarial statements.
[0072] No-answer random question attack / adversarial statement NARQ415 does not contain an answer entity in section 436, but rather a randomly generated question entity 438 ("Brooklyn"). Note that NARQ415 is also an incorrect statement used to train QA systems on how to recognize RARQ attacks / adversarial statements.
[0073] The original unanswered question attack / adversarial statement, NAOQ419, does not contain an answer entity in section 440, but does contain the question entity 426 ("Paris") found in statement 414. Note that NAOQ419 is also an incorrect statement used to train a QA system on how to recognize NAOQ attacks / adversarial statements.
[0074] As shown in block 408, step 4 translates one or more of the attack / adversarial statements created in step 3 into another language. That is, in one or more embodiments of the present invention, the attack / adversarial statements created in step 3 are initially generated using the same language (e.g., English) as the language used by the questions. These attack / adversarial statements are then translated by the QA system into multiple other languages, if they were not already in another language when sent from the text document server 152 shown in FIG. 1, in order for the QA system to evaluate multilingual datasets and models.
[0075] For example, the RARQ407 attack / adversarial statement is translated into German to create RARQ423 (similar to the German adversarial statement 323 shown in FIG. 3).
[0076] As shown in block 410, step 5 then randomly inserts the attack / adversarial statements created in step 3 and / or step 4 into a context (e.g., the original / correct passage / context shown in block 303 of FIG. 3) to create adversarial passages A, B, C, D, and A' shown in FIG. 3. That is, the generated adversarial statements (e.g., RARQ407, RAOQ411, NARQ415, NAOQ419, RARQ423, etc.) are inserted into random positions in a context, such as the original / correct passage shown in block 303 of FIG. 3, which is shown in FIG. 4 as adversarial passage 425. This generates a new instance (Qx, Cy, Ay, Sz), where x, y, z∈L are the languages of the question, context, and statement, respectively, which do not have to be the same, as shown in blocks 305, 309, 313, 317, and 321 in Figure 3.
[0077] The aforementioned attacks / adversarial statements allow QA systems to explore vulnerabilities in MLQA datasets and MBERT by forcing them to predict incorrect answers not just in one language but in multiple languages, such that the question, context, and adversarial statements can all be in the same language or different languages.
[0078] Therefore, referring now to FIG. 5, an exemplary process for using a trained model in a question-answering dialogue system to defend against adversarial attacks / statements is shown in accordance with one or more embodiments of the present invention.
[0079] As shown in block 501, the process begins by retrieving a question / answer (QA) dataset of known questions and their known correct answers (e.g., question and answer pairs such as the question and answer pairs shown in block 301 of FIG. 3).
[0080] As shown in block 503 of FIG. 5, attack / adversarial statements of multiple types (e.g., RARQ, RAOQ, NARQ, and / or NAOQ) and / or multiple languages are added to the context / passages of the entire training dataset (of question and answer pairs) as described in FIG. 4.
[0081] As previously discussed in describing Figure 4, a QA model is created by converting questions 412 into statements 414 and then associating questions 412 with statements 414. This QA model is then modified to create a multilingual QA (MLQA) model, as described in block 505 of Figure 5. In one or more embodiments of the present invention, the MLQA model is created in two steps.
[0082] The first step is to intentionally pollute / populate passage 303 with one or more of the attack / adversarial statements in multiple languages, as described in Figure 4, to create additional training data for the MLQA model.
[0083] Additionally, a passage may be fed with one or more adversarial statements multiple times, using the same or different attacks, to create new passages as additional training data for the MLQA model.
[0084] As illustrated in Figure 5, the original questions / answers / passages and the new questions / answers / passages created in Figure 4 are used to retrain the original MLQA model.
[0085] The second step is to improve the retrained MLQA model by bootstrapping the adversarial policy (i.e., adding a policy on how to handle various attacks / adversarial statements in different languages) onto the retrained version of the MLQA model using attacks. This retrained MLQA model is then recursively trained using reinforcement learning by an artificial intelligence (AI) system, as indicated by arrow block 506.
[0086] As shown in block 507, during each iteration, questions / answers / passages containing adversarial attacks are run by the retrained MLQA model containing multiple languages to evaluate whether the newly retrained MLQA model is robust (i.e., immune to attacks).
[0087] In one or more embodiments of the present invention, the process illustrated in block 503 and / or block 505 shown in Figure 5 uses artificial intelligence, such as artificial intelligence 124 shown in Figure 1. Such artificial intelligence 124 may take various forms, according to one or more embodiments of the present invention, including, but not limited to, a transformer-based reinforcement learning system utilizing multilanguage bidirectional encoder representation from transformers (MBERT), a deep neural network (DNN), a recursive neural network (RNN), a convolutional neural network (CNN), and the like.
[0088] Thus, in one or more embodiments of the invention, the MBERT system described below in Figure 7 is a Transformer-based system used with reinforced learning as shown in Figure 5. That is, the combination of Transformers and reinforced learning allows the system to determine which bootstrapped adversarial policy to use in determining whether to (1) create RAOQ adversarial statements, NAOQ adversarial statements, etc., described in Figures 3 and 4 from a context such as the example passage shown in block 303 of Figure 3, (2) translate a question, such as example question 412 shown in Figure 4, into another language, or (3) translate an answer, such as example statement 414 shown in Figure 4, into another language, or a combination thereof.
[0089] That is, in a reinforcement learning setting in one or more embodiments of the present invention, a system (e.g., the QA dialogue system 156 shown in FIG. 1) finds the best combination of one or more adversarial policies via a policy gradient algorithm, such as the REINFORCE algorithm (described below), and then applies those policies to a large pool of adversarial statements, translations, etc., which may be newly created during each iteration and are used to train the system's defenses.
[0090] As described herein, in one or more embodiments of the present invention, candidate contexts (e.g., one or more of the contexts / passages shown in blocks 303, 305, 309, 313, 317, 321 of FIG. 3 ) are evaluated to determine the location of a correct answer within such contexts / passages, even if they are corrupted, possibly with adversarial conditions (e.g., elements 307, 311, 315, 319, 323 shown in FIG. 3 ).
[0091] Referring now to FIG. 6, a high level overview of one or more embodiments of the present invention is shown.
[0092] A Transformer model system 624 (i.e., a system that models a context by using a Transformer, as described herein), similar to AI 124 shown in FIG. 1, receives as input a question 604 (similar to question 304 shown in FIG. 3) and a candidate context 600 (similar to some or all of the contexts shown in blocks 303, 305, 309, 313, 317, and 321 of FIG. 3). Candidate context 600 also includes candidate answer locations 602, which indicate locations within candidate context 600 where candidate context 600 is predicted to hold the correct answer to question 604. Transformer model system 624 uses these different answer locations 602 to train Transformer model system 624 on how to accurately identify the correct answer location from the candidate answer locations 602. As shown in block 604, in one or more embodiments of the present invention, question 604, candidate context 600, and candidate answer locations 602 are combined into a single group. Regardless of whether the questions 604, candidate contexts 600, and candidate answer locations 602 are combined into a single group, in one or more embodiments of the present invention, a controller 601 (e.g., computer 101 shown in FIG. 1) sends various questions, candidate contexts, and / or candidate answer locations to the transformer model system 624 for training the transformer model system 624 and / or for evaluating various questions, candidate contexts, and / or candidate answer locations.
[0093] Referring now to FIG. 7, an exemplary multilanguage bidirectional encoder representation from transformers (MBERT) system 724 as used in one or more embodiments of the present invention is shown.
[0094] The MBERT system 724 (i.e., a training system that uses artificial intelligence to identify the location of correct answer terms within contexts / passages, including contexts / passages corrupted by adversarial statements such as those shown in Figures 3 and 4) uses the candidate context 600, candidate answer positions 602, and question 604 described in Figure 6 as input. These inputs are converted into embeddings (vectors). The embedding Ep (element 702) for candidate answer position 602 represents the candidate position of the correct answer within candidate context 600. The embeddings Eq1 through Eqn (elements 703 through 705) are different vectors representing the terms within question 604. The embeddings Ecc1 through Eccm (elements 707 through 709) are different vectors representing the terms within candidate context 600.
[0095] Next, node 711 (i.e., an artificial intelligence computation node) evaluates candidate answer location 602 as being the correct location within candidate context 600 for providing the correct answer to question 604 using weights, algorithms, biases, etc. (similar to the weights, algorithms, biases, etc. described in block 911 with respect to deep neural network 924 shown below in FIG. 9).
[0096] Node 711 outputs a confidence 713 that the location within candidate context 600, beginning at start location 715 and ending at end location 717, is accurate. This confidence 713 is output as an answerability prediction 719 (i.e., the confidence that a particular start / end location contains the answer to question 604), as shown in start / end location prediction 721. Answerability prediction 719 and start / end location prediction 721 are then sent to controller 701.
[0097] Line 723 indicates that controller 701 then uses different candidate contexts / question / answer locations from the candidate contexts / question / answer locations trained by MBERT system 724, as indicated by line 723 going to block 604. As in Figure 6, these different candidate answer locations, questions, or candidate contexts, or combinations thereof, may be input collectively, individually, or both, to MBERT system 724 in accordance with one or more embodiments of the present invention.
[0098] Referring now to FIG. 8, there is shown a QA dialogue system 800 that utilizes a transformer-based system to answer questions using correct answers 816 from candidate contexts 801 (e.g., one or more of the passages shown in FIG. 3).
[0099] In one or more embodiments of the present invention, a transformer (such as the transformer used by MBERT described herein) combines tokens (e.g., words within a sentence) with a positional identifier for the token's position within the sentence and a sentence identifier for the sentence to create embeddings. These embeddings are used to answer questions within a particular context, where a given context may or may not include adversarial statements.
[0100] Next, a reinforcement system (e.g., REINFORCE, which uses gradients such as Monte Carlo policy gradients) allows the system to learn which policies are effective in helping the MLQA model understand when a statement is an adversarial attack.
[0101] Assume there are multiple bootstrapped adversarial policies 804 that can be used by a Transformer-based reinforcement learning system 802 to understand an adversarial statement (e.g., the example adversarial statement shown above in FIG. 4). The Transformer-based reinforcement system (e.g., QA dialogue system 800) then uses a gradient-based algorithm, such as the REINFORCE algorithm, to apply various adversarial policies and answer positions 302 from the bootstrapped adversarial policies 804 until a suitable adversarial statement (e.g., RARQ adversarial statement 806 and / or its corresponding translated adversarial statement 814), as determined by comparison with real-world types of adversarial statements that attack the QA dialogue system 156 shown in FIG. 1, is no longer considered the optimal training statement. For example, if the query statement “Cafe X is the oldest cafe in Paris” is transformed into an adversarial statement (e.g., the RARQ adversarial statement “Corporation A is the oldest cafe in the Arctic Ocean”) or its translated adversarial statement (“Corporation A ist das alteste Cafe im Arktischen Ozean”), or both, it is shown that one or both of these adversarial statements matches the type of adversarial statements that actually attack (or are predicted to attack) the QA dialogue system 156, and then these adversarial statements and / or translated adversarial statements are sent to a controller (e.g., the computer 101 shown in FIG. 1 ) to retrain the MLQA model (block 505 of FIG. 5 ) and run the attack pipeline (block 507 of FIG. 5 ).
[0102] In one or more embodiments of the present invention, a transformer-based learning system (e.g., the transformer model system 624 shown in FIG. 6) also translates the correct statement into another language (the translated correct statement and / or the original question) for use in performing the steps described in FIG. 4, thus enabling the QA dialogue system 156 to process questions / statements in multiple languages.
[0103] In one or more embodiments of the present invention, the artificial intelligence 124 utilizes electronic neural network architectures other than Transformer-based systems (e.g., the Transformer model system 624), such as those found in deep neural networks (DNNs), convolutional neural networks (CNNs), or recurrent neural networks (RNNs), in conjunction with reinforced learning systems.
[0104] In a preferred embodiment, a deep neural network (DNN) is used to evaluate text / numerical data in documents from a text corpus received from the text document server 152 shown in FIG. 1, while a CNN is used to evaluate images from an audio or image corpus (e.g., from the audio file server 154 or video file server 160 shown in FIG. 1, respectively).
[0105] CNNs are similar to DNNs in that they both utilize interconnected electronic neurons. However, CNNs differ from DNNs in that (1) CNNs include neural layers whose size is based on filter size, stride value, padding value, etc., and (2) CNNs utilize a convolutional approach to analyze image data. CNNs get their name "convolutional" because they rely on the convolution (i.e., a mathematical operation on two functions to obtain a result) of filtering and pooling pixel data to generate a predicted output (i.e., obtain a result).
[0106] RNNs are similar to DNNs in that they both utilize interconnected electronic neurons, but they are a much simpler architecture in which child nodes feed into parent nodes using weight matrices and nonlinearities (such as trigonometric functions) that are adjusted until the parent node produces the desired vector.
[0107] The logical units within an electronic neural network (DNN or CNN or RNN) are called "neurons" or "nodes." When an electronic neural network is implemented entirely in software, each neuron / node is an individual piece of code (i.e., instructions that perform a specific operation). When an electronic neural network is implemented entirely in hardware, each neuron / node is an individual piece of hardware logic (e.g., processor, gate array, etc.). When an electronic neural network is implemented as a combination of hardware and software, each neuron / node is a set of instructions and / or a piece of hardware logic.
[0108] As the name implies, neural networks are loosely modeled after biological neural networks (e.g., the human brain). Biological neural networks consist of a series of interconnected neurons that influence each other. For example, a first neuron can be electrically connected to a second neuron by a synapse through the release of neurotransmitters that are received by the second neuron. These neurotransmitters can cause the second neuron to be excited or inhibited. Patterns of excitation / inhibition of interconnected neurons ultimately lead to biological outcomes, including thoughts, muscle movements, memory retrieval, and more. While this description of a biological neural network is highly simplified, the high-level overview is that one or more biological neurons influence the behavior of one or more other bioelectrically connected biological neurons.
[0109] Electronic neural networks are similarly made up of electronic neurons, but unlike biological neurons, electronic neurons are never technically "inhibitory" and are often only "excitatory" to varying degrees.
[0110] In an electronic neural network, neurons are arranged in layers known as the input layer, hidden layer, and output layer. The input layer contains neurons / nodes that receive input data and send it to a series of hidden layers of neurons, and every neuron from one hidden layer is interconnected with every neuron in the next hidden layer. The final hidden layer then outputs the results of its calculations to the output layer, which is often one or more nodes to hold vector information.
[0111] In one or more embodiments of the present invention, deep neural networks are used to create MLQA models for question-answering dialogue systems.
[0112] Referring now to FIG. 7, there is shown an exemplary deep neural network (DNN) form of a Transformer (i.e., part of an MBERT system 724) that is used to create and utilize MLQA models when answering questions in accordance with one or more embodiments of the present invention.
[0113] For illustrative purposes, assume that the input to the Transformer / DNN includes the original question 412 (e.g., "What is the oldest cafe in Paris?") and the correct answer location (e.g., the location of "(Cafe X) Cafe X" within one or more of the candidate contexts). Such a DNN can use these inputs to create an initial QA model by aligning answer entities (e.g., element 446 and element 424 shown in FIG. 4) and question entities (e.g., element 426 shown in FIG. 4).
[0114] As shown in FIG. 8, this DNN (denoted as QA dialogue system 800) also includes a bootstrapped adversarial policy (e.g., a policy that determines how to recognize different types of attacks / adversarial statements within a passage), algorithms, rules, etc. that use RARQ adversarial statements 806 (examples of which are described in FIGS. 3 and 4), RAOQ adversarial statements 808 (examples of which are described in FIGS. 3 and 4), NARQ adversarial statements 810 (examples of which are described in FIGS. 3 and 4), and NAOQ adversarial statements 812 (examples of which are described in FIGS. 3 and 4), as well as translations of these adversarial statements (e.g., 423 shown in FIG. 4), which are input into context 801 and shown as translated adversarial statements 814. That is, it should be understood that RARQ adversarial statement 806, RAOQ adversarial statement 808, NARQ adversarial statement 810, NAOQ adversarial statement 812, or translated adversarial statement 814, or a combination thereof, are part of (incorporated into) context 801, but are shown in different boxes in FIG. 8 solely for purposes of clarity.
[0115] The algorithms, rules, etc. used in the DNN / QA dialogue system 800 can recursively define and improve the trained MLQA model.
[0116] FIG. 9 shows a high-level overview of an exemplary trained deep neural network (DNN) 924 that can be used to provide correct answer locations 915 within suggested answer contexts / passages 902 when responding to a new question 901.
[0117] When automatically adjusted, the mathematical functions, output values, weights, and / or biases are adjusted using "backpropagation," where a "gradient descent" method determines how each mathematical function, output value, weight, and / or bias should be adjusted to provide an accurate output 917. That is, the mathematical functions, output values, weights, and / or biases shown in block 911 of example node 909 are recursively adjusted until the expected vector value of the trained MLQA model 915 is reached.
[0118] A new question 901 (e.g., "What is the oldest cafe in Madrid?") along with suggested answer contexts / passages 902 (e.g., provided by a question / answer database such as the question / answer database described above) are also input to the input layer 903, which processes such information before passing it to the intermediate layer 905. That is, using a similar process described above in FIGS. 3-5, one or more answer entities and one or more question entities in the new question 901 are used to retrieve answers (similar to the answer described by statement 414) from the QA dataset, which is used to retrieve similar types of answers from the contexts / passages. One or more of these contexts / passages are determined by the DNN 924 to correctly answer the new question 1001, while ignoring the adversarial statements.
[0119] Thus, the mathematical functions, output values, weights, and bias values of the elements shown in block 911 and found in one or more or all of the neurons in the DNN 924 cause the output layer 907 to produce an output 917, which includes a correct answer location 915 of the correct answer to the new question 901, including the answer found in the passage containing the adversarial statement regarding the new question 901.
[0120] In one or more embodiments of the present invention, the correct answer location 915 is then returned to the questioner.
[0121] Thus, in one or more embodiments of the present invention, the present invention not only looks for a specific known correct answer ("Cafe X") to a specific type of question ("What is the oldest cafe in Paris?") within a context / passage, but also looks for the correct answer location of the correct answer to the specific type of question, thus providing a system that is much more robust than just a word search program.
[0122] Referring now to FIG. 10, there is shown a high-level flowchart of one or more steps performed in accordance with one or more embodiments of the present invention.
[0123] After start block 1002, as shown in block 1004, a computing device (e.g., computer 101 shown in FIG. 1 , implemented as MBERT system 724 and / or DNN shown in FIG. 7 , or artificial intelligence 124, or QA question-answering dialogue system 156, or a combination thereof) accesses a plurality of adversarial statements (e.g., elements 307, 311, 315, and 319 shown in FIG. 3 ) that can be used to launch adversarial attacks against the question-answering dialogue system. The question-answering dialogue system (e.g., artificial intelligence 124 and / or QA question-answering dialogue system 156) shown in FIG. 1 is a QA system designed / trained to provide correct answers to a particular type of question, such as "What is the oldest cafe in a certain city?"
[0124] As shown in block 1006, multiple adversarial statements are utilized in training a machine learning model (e.g., trained MLQA model 915 shown in FIG. 9).
[0125] As shown in block 1008, the computing device enhances the trained machine learning model by bootstrapping an adversarial policy that identifies multiple types of adversarial statements onto the trained machine learning model (e.g., bootstrapped adversarial policy 804 shown in FIG. 8).
[0126] As shown in block 1010, when a computing device responds to a question posed to the question-answering dialogue system 800 shown in FIG. 8 (e.g., the MBERT system 724 shown in FIG. 7), the computing device utilizes the trained, bootstrapped machine learning model (e.g., the updated, bootstrapped, trained MLQA model) to prevent adversarial attacks.
[0127] As indicated by line 1014, the process operates in a recursive manner by returning to block 1004 until it is determined that the QA dialogue system is adequately trained (e.g., by exceeding a predetermined level of correct percentage for identifying and overcoming adversarial statements).
[0128] The flowchart ends at end block 1012 .
[0129] In one or more embodiments of the present invention, the trained bootstrapped machine learning model is tested by a computing device that converts questions for a question-answering dialogue system into statements that include placeholders for answers; randomly selecting answer entities from the answers and adding the randomly selected answer entities in place of the placeholders to generate adversarial statements; generating attacks against the trained bootstrapped machine learning model using the questions and contexts / passages that include the adversarial statements; measuring responses from the trained bootstrapped machine learning model to the generated attacks; and modifying the trained bootstrapped machine learning model to improve its level of response to the generated attacks.
[0130] That is, as shown in FIGS. 3-10, a computing device converts a question for a question-answering dialogue system into a statement that includes a placeholder for the answer (e.g., see steps 1 and 2 in FIG. 4). Next, the computing device randomly selects an answer entity from the answer and adds the randomly selected answer entity in place of the placeholder to generate an adversarial statement (e.g., see step 3 in FIG. 4). As described herein, the process randomly inputs the adversarial statement into a passage (e.g., a context / passage) to create the adversarial passage. The computing device then generates an attack against the trained, bootstrapped machine learning model using the question and context / passage that include the adversarial passage (e.g., see block 206 in FIG. 2 and / or block 507 in FIG. 5) and measures the response from the trained, bootstrapped machine learning model to the generated attack (e.g., by neurons in the trained DNN 924 shown in FIG. 9). Finally, the computing device modifies the trained, bootstrapped machine learning model (e.g., by backpropagation within the DNN 924 shown in Figure 9) to improve the response level of the response to the generated attack (i.e., to more clearly indicate the presence of an attack).
[0131] In one or more embodiments of the present invention, the plurality of adversarial statements includes a first adversarial statement in a first language and a second adversarial statement in a different second language, but the first adversarial statement and the second adversarial statement both provide the same incorrect answer to a question. For example, the first adversarial statement (e.g., RARQ307 shown in FIG. 3 - "Corporation A is the oldest cafe in the Arctic Ocean") is in a first language (English), and the second adversarial statement (e.g., RARQ323 shown in FIG. 3 - "Corporation A ist das alteste Cafe in Arktischen Ozean") is in a different second language (German), but both of these adversarial statements provide the same incorrect answer to the question "What is the oldest cafe in Paris?" Therefore, as described herein, QA training systems (e.g., DNN924) can respond to adversarial statements in different languages.
[0132] In one or more embodiments of the present invention, a computing device generates RARQ adversarial statements, RAOQ adversarial statements, NARQ adversarial statements, or NAOQ adversarial statements, or a combination thereof (e.g., by actually generating one or more of these adversarial statements).
[0133] In one or more embodiments of the present invention, a computing device retrieves RARQ adversarial statements, RAOQ adversarial statements, NARQ adversarial statements, or NAOQ adversarial statements, or a combination thereof (e.g., from an already created dataset).
[0134] In one or more embodiments of the present invention, the computing device further trains a machine learning model for the question-answering dialogue system to recognize adversarial statements using the generated or retrieved RARQ adversarial statements, RAOQ adversarial statements, NARQ adversarial statements, or NAOQ adversarial statements, or a combination thereof, as input (see Figure 6 of this patent application).
[0135] In one or more embodiments of the present invention, multiple adversarial statements are randomly placed into a single context / passage at a time.
[0136] In one or more embodiments of the present invention, multiple adversarial statements are randomly placed individually into a single context / passage, with each original context / passage containing a new adversarial statement becoming a new context / passage.
[0137] Thus, a novel multilingual QA system is described herein in which the questions, context, and adversarial statements can be in the same language or in different languages. Adversarial / attack statements can be generated in one language and then translated into another language, or the adversarial / attack statements can be received in a different language. In either case, the QA system described herein utilizes a single trained MLQA model that can handle multiple languages, such that the QA system's defenses against attacks are effective, whether the model is zero-shot (trained on data in which the questions, context, and adversarial statements are in a different language than the test data), multilingual (trained on data in which the questions, context, and adversarial statements are in two or more different languages), or both.
[0138] In one or more embodiments, the present invention is implemented using cloud computing. Nevertheless, although this disclosure includes detailed descriptions of cloud computing, it is understood that implementation of the subject matter presented herein is not limited to cloud computing environments. Embodiments of the present invention may be implemented in conjunction with any other type of computing environment now known or later developed.
[0139] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computational resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) and for rapidly provisioning and releasing these resources with minimal administrative effort or interaction with the service provider. This cloud model includes at least five characteristics, at least three service models, and at least four deployment models.
[0140] The features are as follows:
[0141] On-demand self-service: Cloud customers can automatically and unilaterally provision computing power, such as server time and network storage, as needed, without the need for human interaction with the service provider.
[0142] Wide network access: Cloud capabilities are available over the network and accessed using standard mechanisms that facilitate usage by heterogeneous thin-client or thick-client platforms (e.g., mobile phones, laptops, and PDAs).
[0143] Resource Pool: The provider's computing resources are pooled and offered to multiple consumers using a multi-tenant model, with various physical and virtual resources dynamically allocated and reallocated according to demand. There is a sense of location independence in that consumers typically have no control or knowledge regarding the exact location of the resources offered, although at a higher level of abstraction they can still specify location (e.g., country, state, or data center).
[0144] Rapid Elasticity: Cloud capacity can be provisioned quickly and elastically, in some cases automatically, to scale out quickly, and released quickly to scale in quickly. This capacity available for provisioning often appears to consumers as unlimited, and any amount can be purchased at any time.
[0145] Metered Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of utilized services.
[0146] SaaS (Software as a Service): The consumer is provided with the ability to use the provider's applications running on a cloud infrastructure. Those applications can be accessed from a variety of client devices through thin-client interfaces such as web browsers (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or individual application features, except for the possibility of setting limited user-specific application configuration settings.
[0147] PaaS (Platform as a Service): The ability offered to a consumer is to deploy applications they create or acquire, written using programming languages and tools supported by the provider, onto a cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of the application hosting environment.
[0148] Infrastructure as a Service (IaaS): The capability provided to a consumer is the provisioning of processing, storage, network, and other basic computing resources, upon which the consumer can deploy and run any software, which may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating system, storage, deployed applications, and in some cases, limited control over selected network components (e.g., host firewalls).
[0149] The deployment model is as follows:
[0150] Private Cloud: The cloud infrastructure is operated solely for an organization. In one or more embodiments, the cloud infrastructure is managed by the organization or a third party and / or resides on-premises or off-premises.
[0151] Community Cloud: The cloud infrastructure is shared by multiple organizations and supports a specific community of shared interests (e.g., mission, security requirements, policy, and compliance considerations). In one or more embodiments, the cloud infrastructure is managed by these organizations or a third party, and / or resides on-premises or off-premises.
[0152] Public Cloud: This cloud infrastructure is available for use by the general public or large industry organizations and is owned by an organization that sells cloud services.
[0153] Hybrid cloud: This cloud infrastructure is a combination of two or more clouds (private, community, or public) that maintain their unique identity and are bound together by standardized or proprietary technologies that allow for the portability of data and applications (e.g., cloud bursting for load balancing between clouds).
[0154] A cloud computing environment is a service-oriented environment that emphasizes statelessness, loose coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that consists of a network of interconnected nodes.
[0155] Referring now to FIG. 11 , an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 that communicate with each other and with local computing devices used by cloud consumers (e.g., a personal digital assistant (PDA) or mobile phone 54A, a desktop computer 54B, a laptop computer 54C, and / or an automobile computer system 54N). The nodes 10 also communicate with each other. In one embodiment, these nodes are physically or virtually grouped together in one or more networks (not shown), such as the aforementioned private cloud, community cloud, public cloud, and / or hybrid cloud. This enables the cloud computing environment 50 to provide infrastructure, platform, and / or SaaS services that do not require cloud consumers to maintain resources on their local computing devices. The types of computing devices 54A-54N shown in FIG. 11 are intended to be illustrative only, and it is understood that computing node 10 and cloud computing environment 50 can communicate with any type of computer-controlled device through any type of network and / or network-addressable connection (e.g., a connection using a web browser).
[0156] Referring now to Figure 12, a set of functional abstraction layers provided by cloud computing environment 50 (Figure 11) is shown. It should be understood in advance that the components, layers, and functions shown in Figure 12 are intended to be illustrative only, and that embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0157] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include mainframe 61, servers based on RISC (Reduced Instruction Set Computer) architecture 62, servers 63, blade servers 64, storage devices 65, and networks and network components 66. In some embodiments, software components include network application server software 67 and database software 68.
[0158] The virtualization layer 70, in one or more embodiments, comprises an abstraction layer in which examples of virtual entities such as virtual servers 71, virtual storage 72, virtual networks including virtual private networks 73, virtual applications and operating systems 74, and virtual clients 75 are provided.
[0159] By way of example, the management layer 80 provides the following functions: Resource provisioning 81 dynamically procures computing and other resources used to execute tasks within the cloud computing environment; Metering and pricing 82 tracks costs as resources are utilized within the cloud computing environment and charges or bills for the use of those resources; these resources include, for example, application software licenses; Security verifies the identities of cloud users and tasks and protects data and other resources; User portal 83 provides users and system administrators with access to the cloud computing environment; Service level management 84 allocates and manages cloud computing resources to meet required service levels; and Service Level Agreement (SLA) planning and execution 85 proactively prepares and procures cloud computing resources in accordance with SLAs in anticipation of future demand.
[0160] Workload tier 90 illustrates example functions for which a cloud computing environment may be utilized in one or more embodiments. Examples of workloads and functions provided from this tier include mapping and navigation 91, software development and lifecycle management 92, virtual classroom instruction delivery 93, data analytics processing 94, transaction processing 95, and QA interactive system protection processing 96, which may perform one or more of the inventive functions described herein.
[0161] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless expressly stated otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used herein, indicate the presence of stated features, integers, steps, operations, elements, or components, or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, or groups thereof, or combinations thereof.
[0162] Corresponding structures, materials, acts, and equivalents of all means or steps and functional elements within the scope of the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of various embodiments of the present invention has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the invention to the form disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The embodiments were chosen and described to best explain the principles and practical application of the invention and to enable those skilled in the art to understand the invention in terms of various embodiments with various modifications as suited to the particular uses contemplated.
[0163] In one or more embodiments of the present invention, any of the methods described in this disclosure are implemented using VHDL (VHSIC Hardware Description Language) programs and VHDL chips. VHDL is an example of a design entry language for field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and other similar electronic devices. Thus, in one or more embodiments of the present invention, any of the software-implemented methods described herein are emulated by a hardware-based VHDL program and then applied to a VHDL chip, such as an FPGA.
[0164] Thus, having described in detail the embodiments of the invention in this application, and with reference to the example embodiments thereof, it will be apparent that changes and modifications are possible without departing from the scope of the invention as defined in the appended claims.< / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer> < / answer>
Claims
1. accessing, by a computing device, a plurality of adversarial statements capable of conducting adversarial attacks against a question-answering dialogue system, the question-answering dialogue system being trained to provide correct answers to specific types of questions; training, by the computing device, a machine learning model for the question-answering dialogue system using the plurality of adversarial statements; enhancing, by the computing device, the trained machine learning model by appending to the trained machine learning model an adversarial policy for processing multiple types of adversarial statements; utilizing the trained and enhanced machine learning model to fend off adversarial attacks when responding to questions submitted by the computing device to the question-answering dialogue system; and A method comprising:
2. converting, by the computing device, a question for the question-answering dialogue system into a statement including a placeholder for an answer; randomly selecting, by the computing device, an answer entity from the answer and adding the randomly selected answer entity in place of the placeholder to generate an adversarial statement; generating, by the computing device, an attack against the trained and enhanced machine learning model that includes the adversarial statement; measuring, by the computing device, a response by the trained and enhanced machine learning model to the generated attack; modifying the trained and enhanced machine learning model to improve the response level of the response to the generated attack by the computing device; 10. The method of claim 1, further comprising testing the trained and enhanced machine learning model by:
3. 3. The method of claim 1, wherein the plurality of adversarial statements comprises a first adversarial statement in a first language and a second adversarial statement in a different second language, and the first adversarial statement and the second adversarial statement both provide the same incorrect answer to the question.
4. the correct answer comprises a correct answer entity and is associated with a correct question entity, and the method further comprises: generating, by the computing device, a random answer random question (RARQ) adversarial statement, the RARQ adversarial statement being a first type of attack statement, the RARQ adversarial statement including a random answer entity replacing the correct answer entity in the correct answer, and the RARQ adversarial statement including a random question entity replacing the correct question entity in the correct answer; generating, by the computing device, random answer original question (RAOQ) adversarial statements, the RAOQ adversarial statements being attack statements of a second type, the RAOQ adversarial statements including random answer entities replacing the correct answer entities in the correct answer, and the RAOQ adversarial statements including the correct question entities from the correct answer; generating, by the computing device, a no-answer random question (NARQ) adversarial statement, the NARQ adversarial statement being a third type of attack statement, the NARQ adversarial statement replacing the correct answer entity in the correct answer with no answer, and the NARQ adversarial statement including a random question entity replacing the correct question entity in the correct answer; generating, by the computing device, a No Answer Original Question (NAOQ) adversarial statement, the NAOQ adversarial statement being a fourth type of attack statement, the NAOQ adversarial statement replacing the correct answer entity in the correct answer with no answer, and the NAOQ adversarial statement including the correct question entity from the correct answer; utilizing, by the computing device, the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize adversarial statements within context passages; The method of any one of claims 1 to 3, further comprising:
5. the correct answer comprises a correct answer entity and is associated with a correct question entity, and the method further comprises: retrieving a random answer random question (RARQ) adversarial statement, the RARQ adversarial statement being a first type of attack statement, the RARQ adversarial statement including a random answer entity replacing the correct answer entity in the correct answer, and the RARQ adversarial statement including a random question entity replacing the correct question entity in the correct answer; retrieving a random answer original question (RAOQ) adversarial statement, the RAOQ adversarial statement being a second type of attack statement, the RAOQ adversarial statement including a random answer entity replacing the correct answer entity in the correct answer, and the RAOQ adversarial statement including the correct question entity from the correct answer; retrieving No Answer Random Question (NARQ) adversarial statements, the NARQ adversarial statements being a third type of attack statement, the NARQ adversarial statements replacing the correct answer entity in the correct answer with no answer, and the NARQ adversarial statements including a random question entity replacing the correct question entity in the correct answer; retrieving a No Answer Original Question (NAOQ) hostile statement, wherein the NAOQ hostile statement is a fourth type of attack statement, the NAOQ hostile statement replaces the correct answer entity in the correct answer with No Answer, and the NAOQ hostile statement includes the correct question entity from the correct answer; utilizing, by the computing device, the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements within context passages; The method of any one of claims 1 to 3, further comprising:
6. the correct answer comprises a correct answer entity and a correct question entity, and the method comprises: retrieving random answer original question (RAOQ) adversarial statements, the RAOQ adversarial statements including random answer entities that replace the correct answer entities in the correct answer, and the RAOQ adversarial statements including the correct question entities from the correct answer; utilizing, by the computing device, the RAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements; The method of any one of claims 1 to 3, further comprising:
7. the correct answer comprises a correct answer entity and a correct question entity, and the method comprises: retrieving No Answer Random Question (NARQ) adversarial statements, wherein the NARQ adversarial statements replace the correct answer entity in the correct answer with no answer, and the NARQ adversarial statements include random question entities replacing the correct question entity in the correct answer; utilizing, by the computing device, the NARQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements; The method of any one of claims 1 to 3, further comprising:
8. the correct answer comprises a correct answer entity and a correct question entity, and the method comprises: retrieving a No Answer Original Question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement replaces the correct answer entity in the correct answer with No Answer, and the NAOQ adversarial statement includes the correct question entity from the correct answer; utilizing, by the computing device, the NAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements; The method of any one of claims 1 to 3, further comprising:
9. A computer program comprising: accessing a plurality of adversarial statements capable of conducting adversarial attacks against a question-answering dialogue system, the question-answering dialogue system being trained to provide correct answers to specific types of questions; utilizing the plurality of adversarial statements to train a machine learning model for the question-answering dialogue system; and Enhancing the trained machine learning model by appending to the trained machine learning model an adversarial policy on how to handle multiple types of adversarial statements; Utilizing the trained and enhanced machine learning model to prevent adversarial attacks when responding to questions posed to the question-answering dialogue system. A computer program for executing the above.
10. the processor, converting a question to the question-answering dialogue system into a statement containing a placeholder for an answer; randomly selecting answer entities from the answers and adding the randomly selected answer entities in place of the placeholders to generate adversarial statements; randomly inputting the adversarial statements into a passage to create an adversarial passage; generating an attack against the trained and enhanced machine learning model that includes the adversarial passage; and measuring a response by the trained and enhanced machine learning model to the generated attacks; modifying, by a computing device, the trained and enhanced machine learning model to improve a response level of the response to the generated attack; 10. The computer program product of claim 9, further comprising:
11. 11. The computer program product of claim 9 or 10, wherein the plurality of adversarial statements comprises a first adversarial statement in a first language and a second adversarial statement in a different second language, and the first adversarial statement and the second adversarial statement both provide the same incorrect answer to the question.
12. the correct answer comprises a correct answer entity and is associated with a correct question entity, and the processor: generating a random answer random question (RARQ) adversarial statement, the RARQ adversarial statement being a first type of attack statement, the RARQ adversarial statement including a random answer entity replacing the correct answer entity in the correct answer, and the RARQ adversarial statement including a random question entity replacing the correct question entity in the correct answer; generating random answer original question (RAOQ) adversarial statements, the RAOQ adversarial statements being attack statements of a second type, the RAOQ adversarial statements including random answer entities replacing the correct answer entities in the correct answer, and the RAOQ adversarial statements including the correct question entities from the correct answer; generating No Answer Random Question (NARQ) adversarial statements, the NARQ adversarial statements being a third type of attack statement, the NARQ adversarial statements replacing the correct answer entity in the correct answer with no answer, and the NARQ adversarial statements including a random question entity replacing the correct question entity in the correct answer; generating a No Answer Original Question (NAOQ) adversarial statement, the NAOQ adversarial statement being a fourth type of attack statement, the NAOQ adversarial statement replacing the correct answer entity in the correct answer with no answer, and the NAOQ adversarial statement including the correct question entity from the correct answer; utilizing the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements within context passages; The computer program according to any one of claims 9 to 11, further comprising:
13. the correct answer comprises a correct answer entity and is associated with a correct question entity, and the processor: retrieving a random answer random question (RARQ) adversarial statement, the RARQ adversarial statement being a first type of attack statement, the RARQ adversarial statement including a random answer entity replacing the correct answer entity in the correct answer, and the RARQ adversarial statement including a random question entity replacing the correct question entity in the correct answer; retrieving a random answer original question (RAOQ) adversarial statement, the RAOQ adversarial statement being a second type of attack statement, the RAOQ adversarial statement including a random answer entity replacing the correct answer entity in the correct answer, and the RAOQ adversarial statement including the correct question entity from the correct answer; retrieving No Answer Random Question (NARQ) adversarial statements, the NARQ adversarial statements being a third type of attack statement, the NARQ adversarial statements replacing the correct answer entity in the correct answer with no answer, and the NARQ adversarial statements including a random question entity replacing the correct question entity in the correct answer; retrieving a No Answer Original Question (NAOQ) hostile statement, wherein the NAOQ hostile statement is a fourth type of attack statement, the NAOQ hostile statement replaces the correct answer entity in the correct answer with No Answer, and the NAOQ hostile statement includes the correct question entity from the correct answer; utilizing the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements within context passages; The computer program according to any one of claims 9 to 11, further comprising:
14. the correct answer includes a correct answer entity and a correct question entity, and retrieving random answer original question (RAOQ) adversarial statements, the RAOQ adversarial statements including random answer entities that replace the correct answer entities in the correct answer, and the RAOQ adversarial statements including the correct question entities from the correct answer; utilizing the RAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements; The computer program according to any one of claims 9 to 11, further comprising:
15. the correct answer includes a correct answer entity and a correct question entity, and retrieving No Answer Random Question (NARQ) adversarial statements, wherein the NARQ adversarial statements replace the correct answer entity in the correct answer with no answer, and the NARQ adversarial statements include random question entities replacing the correct question entity in the correct answer; utilizing the NARQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements; The computer program according to any one of claims 9 to 11, further comprising:
16. the correct answer includes a correct answer entity and a correct question entity, and retrieving a No Answer Original Question (NAOQ) adversarial statement, wherein the NAOQ adversarial statement replaces the correct answer entity in the correct answer with No Answer, and the NAOQ adversarial statement includes the correct question entity from the correct answer; Utilizing the NAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements; The computer program according to any one of claims 9 to 11, further comprising:
17. The computer program according to any one of claims 9 to 16, wherein the program is provided as a service in a cloud environment.
18. 1. A computer system comprising one or more processors, one or more computer-readable memories, and one or more computer-readable non-transitory storage media, wherein program instructions are stored in at least one of the one or more computer-readable non-transitory storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, the stored program instructions comprising: accessing a plurality of adversarial statements capable of conducting adversarial attacks against a question-answering dialogue system, the question-answering dialogue system being trained to provide correct answers to specific types of questions; utilizing the plurality of adversarial statements to train a machine learning model for the question-answering dialogue system; and Enhancing the trained machine learning model by appending to the trained machine learning model an adversarial policy on how to handle multiple types of adversarial statements; Utilizing the trained and enhanced machine learning model to prevent adversarial attacks when responding to questions posed to the question-answering dialogue system. A computer system executed to perform a method including:
19. the correct answer comprises a correct answer entity and is associated with a correct question entity, and the method further comprises: retrieving a random answer random question (RARQ) adversarial statement, the RARQ adversarial statement being a first type of attack statement, the RARQ adversarial statement including a random answer entity replacing the correct answer entity in the correct answer, and the RARQ adversarial statement including a random question entity replacing the correct question entity in the correct answer; retrieving a random answer original question (RAOQ) adversarial statement, the RAOQ adversarial statement being a second type of attack statement, the RAOQ adversarial statement including a random answer entity replacing the correct answer entity in the correct answer, and the RAOQ adversarial statement including the correct question entity from the correct answer; retrieving No Answer Random Question (NARQ) adversarial statements, the NARQ adversarial statements being a third type of attack statement, the NARQ adversarial statements replacing the correct answer entity in the correct answer with no answer, and the NARQ adversarial statements including a random question entity replacing the correct question entity in the correct answer; retrieving a No Answer Original Question (NAOQ) hostile statement, wherein the NAOQ hostile statement is a fourth type of attack statement, the NAOQ hostile statement replaces the correct answer entity in the correct answer with No Answer, and the NAOQ hostile statement includes the correct question entity from the correct answer; utilizing the RARQ adversarial statements, the RAOQ adversarial statements, the NARQ adversarial statements, and the NAOQ adversarial statements as inputs for further training the machine learning model for the question-answering dialogue system to recognize and ignore adversarial statements within context passages; 20. The computer system of claim 18, further comprising:
20. 20. The computer system of claim 18 or 19, wherein the stored program instructions are provided as a service in a cloud environment.
Citation Information
Patent Citations
Trend analysis device, method and program
JP2014081882A
Method for processing information, information processor, and program
JP2019082987A
Providing suggestions for interactions with an automated assistant in a multi-user message exchange thread
JP2019522266A
Question answering system, question reception answering system, primary answering system, and question answering method using them
JP2020123371A
Question Answering Method and Apparatus
US20190371299A1