Constructing imaginary discourse tree to improve ability to answer convergent questions
By constructing a fictitious discourse tree to supplement the missing entities in the autonomous agent system, the problem that the existing system is unable to answer complex questions is solved, and the generation of complete answers and intelligent answers are achieved.
Patent Information
- Application Number
- CN202510727739.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-10
- Filing Date
- 2019-05-09
- Publication Date
- 2025-09-16
AI Technical Summary
Existing autonomous agent systems are unable to effectively answer complex, multi-sentence, or convergent questions and cannot form complete and accurate answers to the questions, especially when the answers are provided in multiple resources.
By constructing a fictitious discourse tree, using computing devices to represent the rhetorical interrelationships between questions and initial answers, accessing a text corpus to supplement missing entities, and generating a fictitious discourse tree to form a complete answer.
It improves the recall of question-answering for complex, multi-sentence, convergent questions, provides complete and accurate answers, eliminates the need for ontology, and enables intelligent answering.
Smart Images

Figure CN120654693A_ABST
Abstract
Description
[0001] This application is a divisional application of invention patent application 201980030899.2, filed on May 9, 2019, and entitled “Constructing a fictional discourse tree to improve the ability to answer convergent questions”.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 729,335, filed September 10, 2018, and U.S. Provisional Application No. 62 / 668,963, filed May 9, 2018, both of which are incorporated herein by reference in their entireties. Background Art
[0004] Linguistics is the scientific study of language. One aspect of linguistics is the application of computer science to natural human languages, such as English. Due to the significant increase in processor speed and memory capacity, computer applications of linguistics are increasing. For example, computer-enabled analysis of language discourse has facilitated many applications, such as automated agents, that can answer questions from users. The use of autonomous agents or "chatbots" to answer questions, facilitate discussions, manage conversations, and provide social promotion is becoming increasingly popular.
[0005] Autonomous agents can serve queries received from user devices by generating answers based on information found using specific resources such as databases or by querying search engines. However, sometimes a single resource or search engine result cannot fully solve complex user queries or convergent problems. Convergent problems are problems that require answers with a high degree of accuracy.
[0006] Therefore, existing systems are unable to form a complete and accurate answer to a question when the answer is provided in multiple resources. Therefore, new solutions are needed. Summary of the Invention
[0007] Various aspects described herein use a fictitious discourse tree to improve question-answering recall for complex, multi-sentence, convergent questions. More specifically, an improved autonomous agent accesses an initial answer to a question received from a user device. The initial answer partially, but not completely, addresses the question. The agent represents the question and the initial answer as a discourse tree and identifies entities in the question that are not addressed in the answer. The agent accesses additional resources, such as a text corpus, and determines an answer that rhetorically connects the missing entity to another entity in the answer. The agent selects this additional resource, thereby forming a fictitious discourse tree that, when combined with the discourse tree of the answer, can be used to generate an answer that is improved over existing solutions.
[0008] In one aspect, a method includes using a computing device to construct a question discourse tree including a question entity based on a question. The question discourse tree represents the rhetorical relationship between the basic discourse units of the question. The method also includes using a computing device to access an initial answer from a text corpus. The method also includes using the computing device to construct an answer discourse tree including an answer entity from the initial answer. The answer discourse tree represents the rhetorical relationship between the basic discourse units of the initial answer. The method also includes using the computing device to determine that a score indicating the relevance of the answer entity to the question entity is below a threshold. The method also includes generating a fictitious discourse tree. Generating the fictitious discourse tree includes creating an additional discourse tree from the text corpus. Generating the fictitious discourse tree includes determining that the additional discourse tree includes a rhetorical relationship connecting the question entity to the answer entity. Generating the fictitious discourse tree includes extracting a subtree of the additional discourse tree including the question entity, the answer entity, and the rhetorical relationship, thereby generating the fictitious discourse tree. Generating the fictitious discourse tree includes outputting an answer represented by a combination of the answer discourse tree and the fictitious discourse tree.
[0009] In an example, accessing the initial answer includes determining an answer relevance score for a portion of the text, and in response to determining that the answer relevance score is greater than a threshold, selecting the portion of the text as the initial answer.
[0010] In an example, the fictitious discourse tree includes nodes representing rhetorical relations. The method also includes integrating the fictitious discourse tree into the answer discourse tree by connecting the nodes to the answer entities.
[0011] In an example, creating the additional utterance trees includes calculating, for each additional utterance tree, a score indicating a number of question entities including mappings to one or more answer entities in the corresponding additional utterance tree. Creating the additional utterance trees includes selecting, from the additional utterance trees, an additional utterance tree having a highest score.
[0012] In an example, creating the additional utterance trees includes, for each additional utterance tree, calculating a score by applying a trained classification model to one or more of (a) the question utterance tree and (b) the corresponding additional answer utterance tree; and selecting the additional utterance tree with the highest score from the additional utterance trees.
[0013] In an example, the question includes keywords, and accessing the initial answers includes obtaining answers based on a search query including the keywords by performing a search of an electronic document. Accessing the initial answers includes determining, for each answer, an answer score indicating a degree of match between the question and the corresponding answer. Accessing the initial answers includes selecting an answer with a highest score from the answers as the initial answer.
[0014] In an example, calculating the score includes applying a trained classification model to one or more of (a) the question utterance tree and (b) the answer utterance tree; and receiving the score from the classification model.
[0015] In an example, constructing a discourse tree includes accessing sentences including segments. At least one segment includes a verb and a word, each word including a role of the word in the segment. Each segment is a basic discourse unit. Constructing the discourse tree also includes generating a discourse tree representing rhetorical relationships between the segments. The discourse tree includes nodes, each non-terminal node represents a rhetorical relationship between two segments, and each terminal node in the nodes of the discourse tree is associated with one of the segments.
[0016] In an example, the method further includes determining a question communication discourse tree including a question root node from the question discourse tree. The communication discourse tree is a discourse tree including communication actions. The generating further includes determining an answer communication discourse tree from the fictitious discourse tree. The answer communication discourse tree includes an answer root node. The generating includes merging the communication discourse trees by identifying that the question root node and the answer root node are the same. The generating includes calculating a degree of complementarity between the question communication discourse tree and the answer communication discourse tree by applying a prediction model to the merged communication discourse trees. The generating includes outputting a final answer corresponding to the fictitious discourse tree in response to determining that the degree of complementarity is above a threshold.
[0017] In an example, a discourse tree represents rhetorical interrelationships between text segments. The discourse tree includes nodes. Each non-terminal node representing a rhetorical interrelationship between two segments in the segment and each terminal node in the nodes of the discourse tree is associated with a segment in the segment. Constructing the communication discourse tree includes matching each segment having a verb with a verb signature. The matching includes accessing the verb signature. The verb signature includes the verb and thematic role sequence of the segment. The thematic role describes the relationship between the verb and the related words. Matching also includes, for each verb signature in the verb signature, determining the thematic role of the corresponding signature that matches the role of the word in the segment. The matching also includes selecting a particular verb signature from the verb signatures based on the particular verb signature including the maximum number of matches. The matching also includes associating the particular verb signature with the segment.
[0018] In an example, a method includes constructing a question discourse tree including a question entity for a question. The method also includes constructing an answer discourse tree including an answer entity from an initial answer. The method also includes establishing a mapping between a first question entity of each question entity and an answer entity in each answer entity, the mapping establishing a correlation between the answer entity and the first question entity. The method also includes, in response to determining that a second question entity in the question entity is not resolved by any answer entity, generating a fictitious discourse tree by combining an additional discourse tree corresponding to an additional answer with the answer discourse tree. The method also includes determining a question communication discourse tree from the question discourse tree. The method also includes determining an answer communication discourse tree from the fictitious discourse tree. The method also includes calculating a degree of complementarity between the question communication discourse tree and the answer communication discourse tree by applying a prediction model, the question communication discourse tree, and the answer communication discourse tree. The method also includes outputting a final answer corresponding to the fictitious discourse tree in response to determining that the degree of complementarity is higher than a threshold.
[0019] The above method may be implemented as a system including one or more processing devices and / or a non-transitory computer-readable medium, on which program instructions may be stored, the program instructions causing one or more processors to perform the operations described with respect to the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Illustrative aspects of the invention are described in detail below with reference to the following drawings.
[0021] Figure 1 An exemplary rhetorical classification environment according to an aspect is shown.
[0022] Figure 2 Depicted is an example of a discourse tree according to an aspect.
[0023] Figure 3 A further example of a discourse tree according to an aspect is depicted.
[0024] Figure 4 An illustrative mode according to one aspect is depicted.
[0025] Figure 5 Depicted is a node-link representation of a hierarchical binary tree according to an aspect.
[0026] Figure 6 Described according to one aspect Figure 5 Example indented text encoding for representations in .
[0027] Figure 7 Depicted is an exemplary DT regarding an example request for property taxes according to an aspect.
[0028] Figure 8 Describes the Figure 7 Example responses to the questions represented in .
[0029] Figure 9 Illustrated is a discourse tree for an official answer according to an aspect.
[0030] Figure 10 Illustrated is an utterance tree for an original answer according to an aspect.
[0031] Figure 11 Illustrated is a communication utterance tree for a first agent's claim according to an aspect.
[0032] Figure 12 Illustrated is a communication utterance tree for a second agent's statement according to an aspect.
[0033] Figure 13 Illustrated is a communication utterance tree for a third agent's statement according to an aspect.
[0034] Figure 14 Illustrated is a diagram of parsing thickets according to an aspect.
[0035] Figure 15 An exemplary process for constructing an exchange utterance tree according to one aspect is illustrated.
[0036] Figure 16 An exemplary process for constructing a fictional utterance tree according to an aspect is illustrated.
[0037] Figure 17 Depicted are example discourse trees of questions, answers, and two fictitious discourse trees according to one aspect.
[0038] Figure 18 A simplified diagram of a distributed system for implementing one of these aspects is depicted.
[0039] Figure 19 is a simplified block diagram of components of a system environment according to one aspect, through which services provided by the components of the system according to one aspect can be offered as cloud services.
[0040] Figure 20 An exemplary computer system is illustrated in which various aspects of the invention may be implemented. DETAILED DESCRIPTION
[0041] As discussed above, existing systems for autonomous agents have shortcomings. For example, such systems cannot answer complex, multi-sentence, or convergent queries. These systems may also rely on ontologies, or interrelationships between different concepts in a domain, which can be difficult and expensive to build. Furthermore, some existing solutions employ knowledge graph-based approaches, which can limit expressiveness and coverage.
[0042] In contrast, aspects described herein can answer complex user queries by employing domain-independent discourse analysis. Aspects described herein use a fictitious discourse tree to verify and, in some cases, complete the rhetorical links between questions and answers, thereby improving question-answering recall for complex, multi-sentence, convergent questions. The fictitious discourse tree is a discourse tree that represents a combination of an initial answer to a question and supplementary additional answers. In this way, the fictitious discourse tree represents a complete answer to a user query.
[0043] For a given answer to be relevant to a given question, the answer's entities should overlap with the question's entities. An "entity" has an independent and unique existence. Examples include objects, places, and people. An entity can also be a subject or topic, such as "electric car," "brake," or "France."
[0044] In some cases, however, one or more entities in the question are not addressed in the initial answer. For example, a user may ask about the "engine" of their car, but the initial answer does not discuss the "engine" at all. In other cases, more specific entities appear in the answer instead, but their connection to the question entities is not immediately obvious. Continuing with the previous example, the initial answer may contain the entities "spark plug" or "transmission" without connecting these entities to "engine." In other cases, some answer entities may not be explicitly mentioned in the question, but are instead assumed in the question. In order to provide a complete answer, missing entities should be considered and, where appropriate, explained.
[0045] To fill this gap, certain aspects of the present disclosure use a fictitious discourse tree. For example, an autonomous agent application can identify answers that provide links between entities that exist in the question but are missing from the answer. The application combines relevant fragments of the fictitious answer discourse tree into a fictitious discourse tree that represents the complete answer to the question. Thus, the fictitious discourse tree provides rhetorical links that can be used to verify the answer and, in some cases, expand the answer. The application can then present the complete answer to the user device.
[0046] The disclosed solution eliminates the need for ontologies by creating a fictitious discourse tree augmented with tree fragments obtained from documents mined on-demand from various sources. For example, considering the automotive domain, an ontology might represent the interrelationships between brakes and wheels, or between transmissions and engines. In contrast, each aspect obtains a canonical discourse representation of the answer that is independent of the thought structure of a given author.
[0047] In another example, the disclosed solution can use a communication discourse tree to verify rhetorical consistency between two parts of text. For example, a rhetorical consistency application working with an autonomous agent application can verify rhetorical consistency, e.g., style, between a question and an answer, a question and a portion of an answer, or an initial answer and an answer represented by a fictitious discourse tree.
[0048] A "communication discourse tree" or "CDT" comprises a discourse tree supplemented with communication actions. Communication actions are collaborative actions taken by individuals based on mutual negotiation and argumentation, or otherwise as is known in the art. By incorporating labels that identify communication actions, learning of the communication discourse tree can be performed on a feature set that is richer than just rhetorical relations and the syntax of basic discourse units (EDUs). With this feature set, additional techniques such as classification can be used to determine the level of rhetorical consistency between questions and answers or request-response pairs, thereby enabling improved automated agents. In doing so, the computing system implements an autonomous agent that is capable of intelligently answering questions and other messages.
[0049] Certain definitions
[0050] As used in this paper, "rhetorical structure theory" is a field of study and research that provides a theoretical foundation within which the coherence of discourse can be analyzed.
[0051] As used herein, a "discourse tree" or "DT" refers to a structure that represents the rhetorical relationships of a sentence for a portion of a sentence.
[0052] As used herein, a "rhetorical relationship," "rhetorical interrelationship," or "coherence relationship" or "discourse relationship" refers to how two segments of a discourse are logically connected to each other. Examples of rhetorical relationships include elaboration, contrast, and attribution.
[0053] As used herein, a "sentence fragment" or "fragment" is a portion of a sentence that can be separated from the rest of the sentence. A fragment is a basic unit of speech. For example, for the sentence "Dutch accident investigators say that evidence points to pro-Russian rebels as being responsible for shooting down the plane," the two fragments are "Dutch accident investigators say that evidence points to pro-Russian rebels" and "as being responsible for shooting down the plane." A fragment can, but need not, contain a verb.
[0054] As used herein, a "signature" or "frame" refers to the properties of a verb in a segment. Each signature can include one or more topic roles. For example, for the segment "Dutch accident investigators say that evidence points to pro-Russian rebels," the verb is "say," and the signature for this particular use of the verb "say" can be "agent verb topic," where "investigators" is the agent and "evidence" is the topic.
[0055] As used herein, a "topic role" refers to a component of a signature that describes the role of one or more words. Continuing with the previous example, "agent" and "subject" are topic roles.
[0056] As used herein, "nuclearity" refers to which text segment, fragment, or span is more critical to the author's purpose. The nucleus is the more critical span, while the satellites are the less critical spans.
[0057] As used herein, "coherency" refers to linking two rhetorical relationships together.
[0058] As used herein, a "communicative verb" is a verb that indicates communication. For example, the verb "deny" is a communicative verb.
[0059] As used herein, a "communicative action" describes an action performed by one or more agents and agents' subjects.
[0060] Figure 1 An exemplary rhetorical classification environment according to an aspect is shown. Figure 1 Depicted are an autonomous agent 101, an input question 130, an output answer 150, a data network 104, and a server 160. The autonomous agent 101 may include one or more of the following: an autonomous agent application 102, a database 115, a rhetorical consistency application 112, a rhetorical consistency classifier 120, or training data 125. The server 160 may be a public or private internet server, such as a public database of user questions and answers. Examples of functionality provided by the server 160 include a search engine and a database. The data network 104 may be any public or private network, wired or wireless network, wide area network, local area network, or the internet. The autonomous agent application 102 and the rhetorical consistency application 112 may execute within a distributed system 1800, as described with respect to FIG. Figure 18 Further discussion.
[0061] The autonomous agent application 102 can access an input question 130 and generate an output answer 150. The input question 130 can be a single question or a stream of questions such as a chat. As shown, the input question 130 is "What is an advantage of an electric car?" The autonomous agent application 102 can receive the input question 130 and analyze the input question 130, for example, by creating one or more fictitious utterance trees 110.
[0062] For example, the autonomous agent application 102 creates a question utterance tree based on the input question 130 and creates a question utterance tree based on initial answers obtained from a resource such as a database. In some cases, the autonomous agent application 102 may obtain a set of candidate answers and select the best match as the initial answer. By analyzing and comparing entities in the question utterance tree with entities in the answer utterance tree, the autonomous agent application 102 determines one or more entities in the question that are not addressed in the answer.
[0063] Continuing with this example, the autonomous agent application 102 accesses additional resources, such as a text corpus, and determines one or more additional answer discourse trees from the text. The autonomous agent application 102 determines that one of the additional answer discourse trees establishes a rhetorical link between a missing question entity and another entity in the question or answer, thereby designating the particular additional answer discourse tree as a fictitious discourse tree. The autonomous agent application 102 can then form a complete answer from the text represented by the initial answer discourse tree and the fictitious discourse tree. The autonomous agent application 102 provides the answer as an output answer 150. For example, the autonomous agent application 102 outputs the text "There is no need for gasoline (no need to refuel)".
[0064] In some cases, the rhetorical consistency application 112 can work in conjunction with the autonomous agent application 102 to promote improved rhetorical consistency between the input question 130 and the output answer 150. For example, the rhetorical consistency application 112 can verify the degree of rhetorical consistency between one or more discourse trees or portions thereof by using one or more communication discourse trees 114. By using the communication discourse trees, the rhetorical consistency and communication actions between the question and the answer can be accurately modeled. For example, the rhetorical consistency application 112 can verify that the input question and the output answer maintain rhetorical consistency, thereby ensuring not only a complete and responsive answer, but also an answer that is consistent with the style of the question.
[0065] For example, the rhetorical consistency application 112 can create a question communication discourse tree representing the input question 130 and an additional communication discourse tree for each candidate answer. The rhetorical consistency application 112 determines the most appropriate answer from the candidate answers. Various methods can be used. In one aspect, the rhetorical consistency application 112 can create a candidate answer communication discourse tree for each candidate answer and compare the question communication discourse tree to each candidate discourse tree. The rhetorical consistency application 112 can then identify the best match between the question communication discourse tree and the candidate answer communication discourse tree.
[0066] In another example, the rhetorical consistency application 112 creates a question-answer pair for each candidate answer, including the question 130 and the candidate answer. The rhetorical consistency application 112 provides the question-answer pair to a predictive model, such as the rhetorical consistency classifier 120. The rhetorical consistency application 112 uses the trained rhetorical consistency classifier 120 to determine whether the question-answer pair matches above a threshold level of match, e.g., indicating whether the answer solves the question. If not, the rhetorical consistency application 112 continues analyzing additional pairs including questions and different answers until a suitable answer is found. In some cases, the rhetorical consistency application 112 can train the rhetorical consistency classifier 120 using the training data 125.
[0067] Rhetorical Structure Theory and Discourse Tree
[0068] Linguistics is the scientific study of language. For example, linguistics can include the structure of sentences (syntax), such as subject-verb-object; the meaning of sentences (semantics), such as dog bites man versus man bites dog; and what speakers do in conversation, i.e., language analysis beyond sentences or discourse analysis.
[0069] The theoretical foundation of discourse, Rhetorical Structure Theory (RST), can be attributed to Mann, William and Thompson, Sandra, "Rhetorical structure theory: A Theory of Text organization," Text-Interdisciplinary Journal for the Study of Discourse, 8(3):243–281, 1988. Similar to how the syntax and semantics of programming language theory help implement modern software compilers, RST helps implement discourse analysis. More specifically, RST places structural blocks on at least two levels: a first level of coreness and rhetorical relations, and a second level of structure or pattern. A discourse parser or other computer software can parse a text into a discourse tree.
[0070] Rhetorical structure theory models the logical organization of a text, a structure adopted by the author that relies on the relationships between its components. Rhetorical structure theory (RST) models textual coherence by forming a hierarchical, connected structure of the text via a discourse tree. Rhetorical relations are categorized as coordinate and subordinate; these relations hold across two or more text segments, thus achieving coherence. These text segments are referred to as elementary discourse units (EDUs). Clauses within a sentence and sentences within a text are logically connected by the author. The meaning of a given sentence is related to the meaning of the preceding and following sentences. These logical relationships between clauses are referred to as the coherent structure of the text. RST, one of the most popular discourse theories, is based on a tree-like discourse structure called a discourse tree (DT). The leaves of the DT correspond to EDUs, or consecutive atomic text segments. Adjacent EDUs are connected by coherence relations (e.g., attribution, order), forming higher-level discourse units. These units are then also constrained by these relational links. EDUs linked by relations are then differentiated based on their relative importance: cores are the core of the relationship, while satellites are the peripheral parts. As discussed, to determine accurate request-response pairs, both topic and rhetorical consistency are analyzed. When a speaker answers a question, such as a phrase or sentence, the speaker's answer should address the topic of the question. In cases where the question is implicitly posed via the seed text of the message, an appropriate answer that not only maintains the topic but also matches the broad cognitive state of the seed is desirable.
[0071] Rhetorical Relationship
[0072] As discussed, the aspects described herein use a fictitious discourse tree. Rhetorical relations can be described in different ways. For example, Mann and Thompson describe 23 possible relations. See William C. Mann, William & Thompson, Sandra (1987) (“Mann and Thompson”). Maite “Rhetorical structure theory: A Theory of Text organization,” Text-Interdisciplinary Journal for the Study of Discourse, 8(3): 243-281, 1988. Other numbers of relations are also possible.
[0073]
[0074]
[0075] Some empirical research assumes that most texts are structured using core-satellite relationships. See Mann and Thompson 1988. However, other relationships do not carry an explicit choice of core. Examples of such relationships are shown below.
[0076] Relationship Name Segment Other sections contrast A substitute Another alternative joint (Not restricted) (Not restricted) List project Next Project sequence project Next Project
[0077] Figure 2 Depicted is an example of a discourse tree according to an aspect. Figure 2 A discourse tree 200 is included. The discourse tree includes a text segment 201, a text segment 202, a text segment 203, a relationship 210, and a relationship 228. Figure 2 The numbers in correspond to the three text sections. Figure 3 Corresponding to the following example text with three text sections numbered 1, 2, and 3:
[0078] 1. Honolulu, Hawaii will be the site of the 2017 Conference on Hawaiian History.
[0079] 2. It is expected that 200 historians from the US and Asia will attend (It is expected that 200 historians from the US and Asia will attend)
[0080] 3.The conference will be concerned with how the Polynesians sailed to Hawaii.
[0081] For example, relationship 210, or elaboration, describes the interrelationship between text segments 201 and 202. Relationship 228 depicts the interrelationship, or elaboration, between text segments 203 and 204. As depicted, text segments 202 and 203 further elaborate on text segment 201. In the example above, given the goal of informing readers about a meeting, text segment 1 is the core. Text segments 2 and 3 provide more detailed information about the meeting. Figure 2 In the text, horizontal numbers (e.g., 1-3, 1, 2, 3) cover sections of the text (which may be composed of further sections); vertical lines indicate one or more cores; and curved lines represent rhetorical relationships (elaborations) and the direction of the arrows points from satellites to the core. If a text section only serves as a satellite and not as a core, deleting the satellite will still leave a coherent text. If Figure 2If you remove the core from the text, then text sections 2 and 3 will be difficult to understand.
[0082] Figure 3 Another example of a discourse tree according to an aspect is depicted. Figure 3 It includes components 301 and 302, text sections 305-307, relationship 310, and relationship 328. Relationship 310 depicts the mutual relationship 310 between components 306 and 305, and 307 and 305—enablement. Figure 3 The following text sections are involved:
[0083] 1.The new Tech Report abstracts are now in the journal area of the library near the abridged dictionary.
[0084] 2.Please sign your name by any means that you would be interested inseeing.
[0085] 3. Last day for sign-ups is 31 May. (The last day for signing is May 31).
[0086] As can be seen, relationship 328 depicts a mutual relationship between entities 307 and 306 (the mutual relationship being an enablement). Figure 3 The diagram illustrates that although cores can be nested, there is only one core-most text segment.
[0087] Constructing a discourse tree
[0088] Discourse trees can be generated using different methods. A simple example of a bottom-up approach to constructing a DT is:
[0089] (1) Divide the discourse text into units in the following way:
[0090] (a) The cell size can vary depending on the objective of the analysis.
[0091] (b) Typically, the unit is a clause
[0092] (2) Check each cell and its neighbors. Do they maintain a relationship with each other?
[0093] (3) If yes, mark the relationship.
[0094] (4) If not, the unit may be at the boundary of a higher-level relationship. Look at the relationships that hold between larger units (segments).
[0095] (5) Continue until all units in the text have been considered.
[0096] Mann and Thompson also describe a second level of building block structure, called pattern applications. In RST, rhetorical relations are not mapped directly onto text; they are assembled onto structures called pattern applications, and these structures are in turn assembled onto text. Pattern applications are derived from simpler structures called patterns (e.g. Figure 4 Each pattern indicates how to decompose a specific unit of text into other smaller units of text. A rhetorical structure tree, or DT, is a hierarchical system of pattern applications. Pattern applications link multiple consecutive text segments and create complex text segments, which can in turn be linked by higher-level pattern applications. RST asserts that the structure of every coherent discourse can be described by a single rhetorical structure tree, whose top pattern creates a segment that encompasses the entire discourse.
[0097] Figure 4 An illustrative mode according to one aspect is depicted. Figure 4 The joint pattern is shown to be a list of items consisting of cores and no satellites. Figure 4 Patterns 401-406 are depicted. Pattern 401 depicts a contextual relationship between text segments 410 and 428. Pattern 402 depicts a sequential relationship between text segments 420 and 421, and a sequential relationship between text segments 421 and 422. Pattern 403 depicts a contrasting relationship between text segments 430 and 431. Pattern 404 depicts a joint interrelationship between text segments 440 and 441. Pattern 405 depicts a motivational interrelationship between 450 and 451, and an enabling interrelationship between 452 and 451. Pattern 406 depicts a joint interrelationship between text segments 460 and 462. Figure 4 An example of a joint pattern for the following three text segments is shown in :
[0098] 1. Skies will be partly sunny in the New York metropolitan area today.
[0099] 2. It will be more humid, with temperatures in the middle 80's.
[0100] 3.Tonight will be mostly cloudy, with the low temperature between 65 and 70.
[0101] Although Figure 2-Figure 4 Some graphical representations of discourse trees are depicted, but other representations are possible.
[0102] Figure 5 Depicted is a node-link representation of a hierarchical binary tree according to one aspect. Figure 6 As can be seen in [1], the leaves of a DT correspond to contiguous, non-overlapping segments of text called elementary discourse units (EDUs). Adjacent EDUs are connected by relations (e.g., elaboration, attribution, etc.) to form larger discourse units, which are also connected by relations. "Discourse analysis in RST involves two subtasks: discourse segmentation is the task of identifying EDUs, and discourse parsing is the task of linking discourse units into a labeled tree."
[0103] Figure 5 depicts the text segments as leaves or terminal nodes on the tree, each numbered in the order in which they appear in the full text, e.g. Figure 6 shown. Figure 5 Tree 500 is included. Tree 500 includes, for example, nodes 501-507. Nodes indicate relationships. Nodes are either non-terminal nodes (such as node 501) or terminal nodes (such as nodes 502-507). As can be seen, nodes 503 and 504 are related by a joint relationship. Nodes 502, 505, 506, and 508 are cores. Dashed lines indicate branches or text segments that are satellites. Relationships are nodes in gray boxes. Figure 6 Described according to one aspect Figure 5 Example indented text encoding for representations in . Figure 6 Includes text more suited to computer programming 600. "N" is the core, and "S" is the satellite.
[0104] Example of a discourse parser
[0105] Automatic discourse segmentation can be performed in different ways. For example, given a sentence, a segmentation model identifies the boundaries of compound basic discourse units by predicting whether a boundary should be inserted before each specific token in the sentence. For example, one framework considers each token in a sentence sequentially and independently. In this framework, the segmentation model scans the sentence token by token and uses a binary classifier (such as a support vector machine or logistic regression) to predict whether it is appropriate to insert a boundary before the token being examined. In another example, the task is a sequential tagging problem. Once the text is segmented into basic discourse units, sentence-level discourse parsing can be performed to construct a discourse tree. Machine learning techniques can be used.
[0106] In one aspect of the present invention, two Rhetorical Structure Theory (RST) discourse parsers are used: CoreNLPProcessor, which relies on a composition grammar, and FastNLPProcessor, which uses a dependency grammar. They are described in Mihai Surdeanu, Thomas Hicks, and Marco A. Valenzuela-Escarcega, "Two Practical Rhetorical Structure Theory Parsers", Proceedings of the Conference of the North American Chapterof the Association for Computational Linguistics-Human Language Technologies: Software Demonstrations (NAACL HLT) 2015.
[0107] In addition, the above two discourse parsers, i.e., CoreNLPProcessor and FastNLPProcessor, use natural language processing (NLP) for syntactic parsing. For example, Stanford CoreNLP gives the basic form of words, their parts of speech, whether they are the names of companies, people, etc.; normalizes dates, times and numerical quantities; marks the structure of sentences according to phrases and syntactic dependencies; indicates which noun phrases refer to the same entity. In fact, RST is still a theory that may work in many discourse situations but may not work in some situations. There are many variables, including but not limited to, which EDUs are in the coherent text, i.e., what discourse segmenter is used, which relation lists are used and which relations are selected for the EDUs, the corpus of documents used for training and testing, and even what parser is used. Therefore, for example, in "Two Practical Rhetorical Structure Theory Parsers" by Surdeanu et al. cited above, a specific corpus must be tested using a dedicated metric to determine which parser gives better performance. Thus, unlike computer language parsers, which give predictable results, discourse parsers (and segmenters) may give unpredictable results depending on the training and / or test text corpus. Thus, discourse trees are a mix of predictable domains (e.g., compilers) and unpredictable domains (e.g., chemistry, where you need to experiment to determine which combinations will give you the desired result).
[0108] To objectively determine the quality of discourse analysis, a range of metrics are being used, such as the Precision / Recall / F1 metric from Daniel Marcu, "The Theory and Practice of Discourse Parsing and Summarization", MIT Press, November 2000, ISBN: 9780262123722. Precision or positive predictive value is the proportion of relevant instances among the retrieved instances, while recall (also known as sensitivity) is the proportion of relevant instances that have been retrieved out of the total number of relevant instances. Therefore, both precision and recall are based on an understanding and measurement of relevance. Suppose a computer program for identifying dogs in photographs identifies 8 dogs in a picture containing 12 dogs and some cats. Of the eight dogs identified, five are actually dogs (true positives) and the rest are cats (false positives). The program has a precision of 5 / 8 and a recall of 5 / 12. When a search engine returns 30 pages, of which only 20 are relevant, and fails to return 40 additional relevant pages, its precision is 20 / 30 = 2 / 3, while its recall is 20 / 60 = 1 / 3. Therefore, in this case, precision is "how useful the search results are," and recall is "how complete the results are." The F1 score (also known as the F-score or F-measure) is a measure of the accuracy of a test. It takes into account both the precision and recall of a test to calculate the score: F1 = 2 × ((precision × recall) / (precision + recall)) and is the harmonic mean of precision and recall. The F1 score reaches its optimal value (perfect precision and recall) when it is 1, and its worst value when it is 0.
[0109] Autonomous agents or chatbots
[0110] A conversation between human A and human B is a form of discourse. For example, there are Messenger, For applications such as SMS, in addition to the more traditional email and voice conversations, a conversation between A and B can typically also be conducted via messages. A chatbot (which may also be called an intelligent robot or a virtual assistant, etc.) is an "intelligent" machine that, for example, replaces human B and imitates a conversation between two humans to varying degrees. The ultimate goal of the example is that human A cannot tell whether B is a human or a machine (the Turing test developed by Alan Turing in 1950). Discourse analysis, artificial intelligence including machine learning, and natural language processing have made great progress in the long-term goal of passing the Turing test. Of course, as computers become increasingly capable of searching and processing large data repositories and performing complex analyses on data, including predictive analysis, the long-term goal is to make chatbots like people and combined with computers.
[0111] For example, users can interact with intelligent bot platforms through conversational interactions. Also known as a conversational user interface (UI), this interaction is a dialogue between an end user and a chatbot, much like a conversation between two humans. It could be as simple as an end user saying "Hello" to a chatbot, which then responds with "Hi" and asks the user how it can help. It could also be a transactional interaction in a banking chatbot, such as transferring funds from one account to another, an informational interaction in an HR chatbot, such as checking remaining vacation time, or a FAQ in a retail chatbot, such as how to process a return. Natural language processing (NLP) and machine learning (ML) algorithms, combined with other methods, can be used to classify end-user intent. A high-level intent is the goal the end user wants to accomplish (e.g., getting an account balance, making a purchase). An intent is essentially a mapping of customer input to a unit of work that should be performed by the backend. Therefore, based on the phrases spoken by the user into the chatbot, they are mapped to specific and discrete use cases or units of work. For example, checking balances, transferring funds, and tracking expenses are all "use cases" that the chatbot should support. The chatbot is able to infer which unit of work should be triggered from the free text entries typed in natural language by the end user.
[0112] The fundamental principle behind making AI chatbots respond like humans is that the human brain can formulate and understand requests, and then respond to them much better than machines. Therefore, if we mimic human B, the chatbot's requests and responses should improve significantly. Therefore, the initial question is how does the human brain formulate and understand requests? To mimic this, we use models. RST and DT allow this to be done in a formal and repeatable manner.
[0113] At a high level, there are generally two types of requests: (1) a request to perform some action; and (2) a request for information, e.g., a question. The first type has a response in which a unit of work is created. The second type has a response, e.g., a good answer to the question. For example, in some aspects, the answer may take the form of the AI constructing the answer from its extensive knowledge base(s), or matching the best existing answer by searching the internet or intranet or other publicly / privately available data sources.
[0114] Analyze request and response pairs
[0115] Figure 7 Depicts an exemplary DT for an example request regarding property taxes according to one aspect. Node labels are relationships and arrow lines point to satellites. The core is a solid line. Figure 7 The following text is depicted.
[0116] Request: "My husbands'grandmother gave him his grandfather's truck. Shesigned the title over but due to my husband having unpaid fines on hislicense, he was not able to get the truck put in his name. I wanted to put in my name and paid the property tax and got insurance for the truck. By the time it came to sending off the title and getting the tag, I didn't have the money to do so. Now, due to circumstances, I am not going to be able to afford the truck. I went to the insurance place and was refused a refund. I am just wondering that since I am not going to have a tag on this truck, is it possible to get the property tax refunded?”
[0117] Response: "The property tax is assessed on property that you own. Just because you chose to not register it does not mean that you don't own it, so the tax is not refundable. Even if you have not titled the vehicle yet, you still own it within the boundaries of the tax district, so the tax is payable. Note that all states give you a limited amount of time to transfer title and pay the use tax. If you apply late, there will be penalties on top of the normal taxes and fees. You don't need to register it at the same time, but you absolutely need to title it within the period of time stipulated in state law."
[0118] Analysis of the above text reveals that "My husbands’ grandmother gave him his grandfather’s truck" is elaborated upon by "She signed the title over but due to my husband", which is in turn elaborated upon by "having unpaid fines on his license, he was not able to get the truck put in his name", which is further elaborated upon by "I wanted to put in my name", "and paid the property tax", and "and got insurance for the truck".
[0119] “My husband's grandmother gave him his grandfather's truck. She signed the title over but due to my husband having unpaid fines on his license, he was not able to get the truck put in his name. I wanted to put in my name and paid the property tax and got insurance for the truck.” It is elaborated as follows:
[0120] “I didn't have the money”, which is further elaborated by “to do so”. The latter is contrasted with “By the time”. “By the time” is elaborated by “it came to sending off the title” “and getting the tag”.
[0121] “My husband's grandmother gave him his grandfather's truck. She signed the title over but due to my husband having unpaid fines on his license, he was not able to get the truck put in his name. I wanted to put in my name and paid the property tax and got insurance for the truck. By the time it came to sending off the title and getting the tag, I didn't have the money to do so” is contrasted with the following:
[0122] “Now, due to circumstances”, which is elaborated by “I am not going to be able to afford the truck”. The latter is elaborated as follows:
[0123] “I went to the insurance place”
[0124] “and was refused a refund”.
[0125] “My husbands’ grandmother gave him his grandfather’s truck. She signed the title over but due to my husband having unpaid fines on his license, he was not able to get the truck put in his name. I wanted to put in my name and paid the property tax and got insurance for the truck. By the time it came to sending off the title and getting the tag, I didn't have the money to do so. Now, due to circumstances, I am not going to be able to afford the truck. I went to the insurance place and was refused a refund” through the following elaboration:
[0126] “I am just wondering that since I am not going to have a tag on this truck, is it possible to get the property tax refunded?”
[0127] “I am just wondering” is attributed to:
[0128] “that” and “is it possible to get the property tax refunded?” are the same unit, and the condition of the latter is “since I am not going to have a tag on this truck”.
[0129] As you can see, the main topic of this question is "Property tax on a car." The problem involves a contradiction: on the one hand, all property is taxable, but on the other hand, ownership is somewhat incomplete. A good response must both address the main topic of the question and clarify the inconsistency. To do this, the responder makes a stronger statement about the necessity of taxing all property, regardless of its registration status. This example is a member of the positive training set in our Yahoo! Answers evaluation domain. The main topic of this question is "Property tax on a car." The problem involves a contradiction: on the one hand, all property is taxable, but on the other hand, ownership is somewhat incomplete. A good answer / response must both address the main topic of the question and clarify the inconsistency. The reader can observe that because the question involves a contrasting rhetorical relation, the answer must match it with a similar relation to be convincing. Otherwise, the answer will appear incomplete, even to those who are not domain experts.
[0130] Figure 8 Depicts a method according to certain aspects of the present invention. Figure 7 Example responses to the questions presented in [ ]. The central core is "the property tax is assessed on property," which is elaborated by "that you own." "The property tax is assessed on property that you own" is also a core, which is elaborated by "Just because you chose to not register it does not mean that you don't own it, so the tax is not refundable. Even if you have not titled the vehicle yet, you still own it within the boundaries of the tax district, so the tax is payable. Note that all states give you a limited amount of time to transfer title and pay the use tax."
[0131] Core: "The property tax is assessed on property that you own. Just because you chose to not register it does not mean that you don't own it, so the tax is not refundable. Even if you have not titled the vehicle yet, you still own it within the boundaries of the tax district, so the tax is payable. Note that all states give you a limited amount of time to transfertitle and pay the use tax." By the following elaboration: "there will be penalties on top of the normal taxes and fees", which has the condition: "If you apply late", which is illustrated by the contrast between "but you absolutely need to title it within the period of time stipulated in statelaw" and "You don't need to register it at the same time".
[0132] By comparison Figure 7 DT and Figure 8 DT, we can determine the response ( Figure 8 ) and request( Figure 7 In some aspects of the present invention, the above framework is used, at least in part, to determine the DT used for a request / response and the rhetorical consistency between DTs.
[0133] Figure 9 FIGURE 1 illustrates a discourse tree for an official answer according to one aspect. Figure 9As drawn, the official answer or mission statement states: "The Investigative Committee of the Russian Federation is the mainfederal investigating authority which operates as Russia's Anti-corruptionagency and has statutory responsibility for inspecting the police forces,combating police corruption and police misconduct, and is responsible for conducting investigations into local authorities and federal governmentalbodies".
[0134] Figure 10 FIGURE 1 illustrates a discourse tree for an original answer according to one aspect. Figure 10 Charted, another, perhaps more honest, answer states: "Investigative Committee of the Russian Federation is suggested to fight corruption. However, top-rank officers of the InvestigativeCommittee of the Russian Federation are charged with creation of a criminal community. Not only that, but their involvement in large bribes, moneylaundering, obstruction of justice, abuse of power, extortion, and racketeering has been reported. Due to the activities of these officers, dozens of High-profile cases including the ones against criminal lords had been ultimately ruined.”
[0135] The choice of answer depends on the context. The rhetorical structure allows to distinguish between "official", "value-correct", template-based answers and "actual", "original", "reported from the field" or "controversial" answers. (See Figure 9 and Figure 10 Sometimes, the question itself can provide a clue as to the expected answer category. If the question is phrased as a factual or definitional question, without any secondary meaning, then the first category of answers is appropriate. Otherwise, if the question has a "tell me what it actually is" meaning, then the second category is appropriate. Generally speaking, after extracting the rhetorical structure from the question, it is easier to select appropriate answers with similar, matching, or complementary rhetorical structures.
[0136] The official answer is based on elaboration and association, which is neutral with respect to possible controversies contained in the text (see Figure 9 ). At the same time, the original answer includes contrastive relations. This relation is extracted between the phrases that the agent is expected to do and the phrases that the agent is found to have done.
[0137] Classification of request-response pairs
[0138] Rhetorical consistency application 112 can determine whether a given answer or response (such as an answer obtained from answer database 105 or a public database) is responsive to a given question or request. More specifically, rhetorical consistency application 112 analyzes whether a request and response pair is correct or incorrect by determining one or both of (i) correlation or (ii) rhetorical consistency between the request and response. Rhetorical consistency can be analyzed without considering correlation, which can be treated orthogonally.
[0139] The rhetorical consistency application 112 can use different methods to determine the similarity between question-answer pairs. For example, the rhetorical consistency application 112 can determine the degree of similarity between individual questions and individual answers. Alternatively, the rhetorical consistency application 112 can determine a measure of similarity between a first pair including a question and an answer and a second pair including a question and an answer.
[0140] For example, the rhetorical consistency application 112 uses a rhetorical consistency classifier 120 that is trained to predict matching or non-matching answers. The rhetorical consistency application 112 can process two pairs at a time, e.g.,<q1,a1> and<q2,a2> The rhetorical consistency application 112 compares q1 with q2 and a1 with a1, thereby generating a combined similarity score. Such a comparison allows determining whether an unknown question / answer pair contains the correct answer by evaluating its distance to another question / answer pair with a known label. In particular, an unlabeled pair can be<q2,a2> The process is performed so that instead of "guessing" the correctness based on words or structures shared by q2 and a2, both q2 and a2 can be compared with the marked pairs based on such words or structures.<q2,a2> Since this method targets the domain-independent classification of answers, it can only exploit the structural cohesion between questions and answers, but not the "meaning" of the answers.
[0141] In one aspect, the rhetorical consistency application 112 uses training data 125 to train the rhetorical consistency classifier 120. In this manner, the rhetorical consistency classifier 120 is trained to determine the similarity between question and answer pairs. This is a classification problem. The training data 125 can include a positive training set and a negative training set. The training data 125 includes matching request-response pairs in the positive dataset and arbitrary or less relevant or appropriate request-response pairs in the negative dataset. For the positive dataset, various fields are selected with different acceptance criteria indicating whether the answer or response is appropriate for the question.
[0142] Each training dataset includes a set of training pairs. Each training set includes a question exchange utterance tree representing a question, an answer exchange utterance tree representing an answer, and an expected degree of complementarity between the question and the answer. Using an iterative process, rhetorical consistency application 112 provides training pairs to rhetorical consistency classifier 120 and receives the degree of complementarity from the model. Rhetorical consistency application 112 calculates a loss function by determining the difference between the determined degree of complementarity and the expected degree of complementarity for a particular training pair. Based on the loss function, rhetorical consistency application 112 adjusts the internal parameters of the classification model to minimize the loss function.
[0143] Acceptance criteria may vary depending on the application. For example, for community question answering, automated question answering, automated and manual customer support systems, social network communication, and written content by individuals such as consumers about their experience with a product (such as reviews and complaints), acceptance criteria may be low. In scientific texts in the form of FAQs, professional social networks (such as "stackoverflow"), professional news, health and legal documents, RR acceptance criteria may be high.
[0144] Communication Discourse Tree (CDT)
[0145] The rhetorical consistency application 112 can create, analyze, and compare communication discourse trees. Communication discourse trees are designed to combine rhetorical information with speech act structure. CDT contains arcs that are labeled with expressions for communication actions. By combining communication actions, CDT enables modeling of RST relations and communication actions. CDT is a simplification of the parse jungle. See Galitsky, B, Ilvovsky, D. and Kuznetsov SO. Rhetoric Map of an Answer to Compound Queries Knowledge Trail Inc. ACL 2015, 681-686. ("Galitsky 2015"). The parse jungle is a combination of parse trees of sentences with discourse-level relationships between words and parts of these sentences in one graph. By incorporating labels that identify speech acts, learning of communication discourse trees can occur on a richer set of features than just the rhetorical relationships and grammar of basic discourse units (EDUs).
[0146] In this example, a dispute between three parties regarding the cause of the crash of a commercial airliner (i.e., Malaysia Airlines Flight 17) is analyzed. An RST representation of the arguments presented is constructed. In this example, three conflicting agents—a Dutch investigator, the Russian Federation Investigative Committee, and a self-proclaimed republic—exchanged their opinions on the matter. This example illustrates a contentious conflict in which each party does its best to blame the other. To sound more convincing, each party not only presents its own claims but also responds by refuting the other party's claims. To achieve this, each party attempts to match the style and discourse of the other party's claims.
[0147] Figure 11 Illustrated is a communication utterance tree for a first agent's statement according to an aspect. Figure 11 A communication discourse tree 100 is depicted, which represents the following text: "Dutch accident investigators say that evidence points to pro-Russian rebels as being responsible for shooting down plane. The report indicates where the missile was fired from and identifies who was incontrol of the territory and pins the downing of MH17 on the pro-Russian rebels."
[0148] As from Figure 11 As can be seen, the non-terminal nodes of the CDT are rhetorical relations, and the terminal nodes are the basic discourse units (phrases, sentence fragments) that are the subjects of these relations. Some arcs of the CDT are labeled with expressions for communicative actions, including the actor agent and the subject of these actions (what is being communicated). For example, the core node (on the left) for elaborating the relationship is labeled say(Dutch, evidence), and the satellite is labeled responsible(rebels, shooting down). These labels are not intended to express that the subjects of the EDU are evidence and shooting down, but to enable the CDT to match other CDTs in order to find similarities between them. In this case, simply linking these communicative actions through rhetorical relations without providing information about the communicative discourse would be too limited to represent the structure of what is being communicated and how it is being communicated. RR's requirement for having identical or coordinated rhetorical relations is too weak, and therefore requires consistency in the CDT labels of the arcs on top of the matching nodes.
[0149] The straight edges of the graph are syntactic relations, and the curved arcs are discourse relations such as anaphora, identical entities, subentities, rhetorical relations, and communicative moves. The graph contains much richer information than just the combination of parse trees for each sentence. In addition to the CDT, the parse jungle can also be generalized at the level of words, relations, phrases, and sentences. Speech moves are logical predicates that express the agents involved in each speech act and its subject. As proposed by frameworks such as VerbNet, the arguments of the logical predicates are formed according to their respective semantic roles. See Karin Kipper, Anna Korhonen, Neville Ryant, Martha Palmer, A Large-scale Classification of English Verbs, Language Resources and Evaluation Journal, 42(1), pp.21-40, Springer Netherland, 2008 and / or Karin Kipper Schuler, Anna Korhonen, Susan W. Brown, VerbNet overview, extensions, mappings and apps, Tutorial, NAACL-HLT 2009,Boulder,Colorado.
[0150] Figure 12Illustrated is a communication utterance tree for a second agent's statement according to an aspect. Figure 12 A communication discourse tree 1200 is depicted, which represents the following text: "The Investigative Committee of the Russian Federation believes that the plane was hit by a missile, which was not produced in Russia. The committee cites an investigation that established the type of the missile."
[0151] Figure 13 Illustrated is a communication utterance tree for a third agent's statement according to an aspect. Figure 13 A communication discourse tree 1300 is depicted, representing the following text: "Rebels, the self-proclaimed Republic, deny that they controlled the territory from which the missile was allegedly fired. It became possible only after three months after the tragedy to say if rebels controlled one or another town."
[0152] As can be seen from the exchange utterance trees 1100-1300, the responses are not arbitrary. The responses talk about the same entities as the original text. For example, the exchange utterance trees 1200 and 1300 are related to the exchange utterance tree 1100. The responses support the different estimates and opinions about these entities and the actions of these entities.
[0153] More specifically, the replies of the involved agents need to reflect the communication discourse of the first seed message. As a simple observation, since the first agent uses attributions to convey its claims, the other agents must follow this set of claims, either providing their own attributions or attacking the validity of the proponent's attributions, or both. To capture the various characteristics required to preserve the communication structure of the seed message in subsequent messages, pairs of corresponding CDTs can be learned.
[0154] To verify request-response consistency, utterance relations or speech acts (communication actions) alone are usually not sufficient. Figure 11-13As can be seen from the examples drawn, the discourse structure of the interactions between agents and the types of interactions are useful, but the domain of the interactions (e.g., military conflict) or the subjects (i.e., entities) of these interactions need not be analyzed.
[0155] Indicates rhetorical relationships and communicative actions
[0156] To compute similarities between abstract structures, two approaches are frequently used: (1) representing these structures in a numerical space and expressing similarities as numbers, which is a statistical learning approach, or (2) using structural representations without a numerical space, such as trees and graphs, and expressing similarities as maximum common substructures. Expressing similarity as a maximum common substructure is called generalization.
[0157] Learning communicative actions facilitates the expression and understanding of arguments. Computational verb dictionaries support the acquisition of action entities and provide a rule-based formalism for expressing their meaning. Verbs convey the semantics of the event they describe and the relationships between participants in that event, projecting grammatical structures that encode this information. Verbs, especially communicative action verbs, can vary widely and exhibit rich semantic behavior. In response, verb classification helps learning systems cope with this complexity by organizing verbs into groups that share core semantic properties.
[0158] VerbNet is a lexicon that identifies the semantic roles and syntactic patterns of verbs in each class and makes explicit the connection between the syntactic patterns and the underlying semantic relations that can be inferred for all members of the class. See Karin Kipper, Anna Korhonen, Neville Ryant, and Martha Palmer, Language Resources and Evaluation, Vol. 42, No. 1 (March 2008), 21. Each syntactic frame or verb signature of a class has a corresponding semantic representation that details the semantic relations between the participants in the event process.
[0159] For example, the verb amuse is part of a cluster of similar verbs that have similar argument (semantic role) structures, such as amaze, anger, arouse, disturb, and irritate. The roles of the arguments of these communicative actions are as follows: Experience (experiencer, usually a living entity), Stimulus, and Result. Each verb can have a meaning category that is distinguished by syntactic features used for how the verb appears in a sentence or frame. For example, the frame for amuse is as follows, using the following key noun phrase (NP), noun (N), communicative action (V), verb phrase (VP), adverb (ADV):
[0160] NP V NP. Example: "The teacher amused the children." Syntax: Stimulus VExperiencer. Clause: amuse(Stimulus,E,Emotion,Experiencer), cause(Stimulus,E),emotional_state(result(E),Emotion,Experiencer).
[0161] NP V ADV-Middle. Example: "Small children amuse quickly." Syntax: Experiencer V ADV. Clauses: amuse(Experiencer, Prop):-, property(Experiencer, Prop), adv(Prop).
[0162] NP V NP-PRO-ARB. Example "The teacher amused". Syntax: Stimulus V.amuse(Stimulus,E,Emotion,Experiencer):.cause(Stimulus,E),emotional_state(result(E),Emotion,Experiencer).
[0163] NP.cause V NP. Example "The teacher's dolls amused the children." Syntax: Stimulus<+genitive>('s)V Experiencer.amuse(Stimulus,E,Emotion,Experiencer):.cause(Stimulus,E),emotional_state(during(E),Emotion,Experiencer).
[0164] NP V NP ADJ. Example "This performance bored me totally". Syntax: Stimulus VExperiencer Result.amuse(Stimulus,E,Emotion,Experiencer).cause(Stimulus,E),emotional_state(result(E),Emotion,Experiencer),Pred(result(E),Experiencer).
[0165] Communication actions can be characterized into multiple clusters, such as:
[0166] Verbs with predicate complement (appoint, characterize, dub, declare, conjecture, masquerade, orphan, captain, consider, classify), verbs of perception (see, sight, peer).
[0167] Mental state verbs (amuse, admire, marvel, appeal), desire verbs (want, long).
[0168] Verbs of judgment (judgment), verbs of evaluation (assess, estimate), verbs of search (hunt, search, stalk, investigate, rummage, ferret), verbs of social interaction (correspond, marry, meet, battle), verbs of communication (transfer (message), inquire, interrogate, tell, manner (speaking), talk, chat, say, complain, advise, confess, lecture, overstate, promise), verbs of avoidance (avoid), verbs of measurement (register, cost, fit, price, bill), verbs of aspect (begin, complete, continue, stop, establish, sustain).
[0169] The aspects described herein offer advantages over statistical learning models. Compared to statistical solutions, aspects using a classification system can provide verbs or verb-like structures that are determined to lead to target features (such as rhetorical consistency). For example, statistical machine learning models express similarity as numbers, which can make it difficult to interpret.
[0170] Represents a request-response pair
[0171] Representing request-response pairs facilitates classification-based operations based on pairs. In an example, request-response pairs can be represented as a parse jungle. A parse jungle is a representation of parse trees for two or more sentences, with discourse-level relationships between sentence parts and words in a graph. See Galitsky (2015). The topic similarity between questions and answers can be expressed as a common subgraph of the parse jungle. The greater the number of common graph nodes, the higher the similarity.
[0172] Figure 14 Illustrated is a parsing forest according to an aspect. Figure 14 A parse jungle 1400 is depicted, including a parse tree 1401 and a parse tree 1402 for a corresponding response.
[0173] Parse tree 1401 represents the problem: "I just had a baby and it looks more like the husband I had my baby with. However it does not look like me at all and I amscared that he was cheating on me with another lady and I had her kid. Thischild is the best thing that has ever happened to me and I cannot imagine giving my baby to the real mom."
[0174] Response 1402 indicates the response "Marital therapists advise on dealing with a child being born from an affair as follows. One option is for the husband to avoid contact but just have the basic legal and financial commitments. Another option is to have the wife fully involved and have the baby fully integrated into the family just like a child from a previous marriage."
[0175] Figure 14 Represents a greedy approach for representing linguistic information about a paragraph of text. The straight edges of the graph are syntactic relations, and the curved arcs are discourse relations such as anaphora, same entity, subentity, rhetorical relations, and communicative moves. Solid arcs are used for same entity / subentity / anaphora relations, and dashed arcs are used for rhetorical relations and communicative moves. Oval labels in the straight edges represent syntactic relations. Lemmas are written in the boxes of the nodes, and lemma forms are written to the right of the nodes.
[0176] Parse jungle 1400 contains much richer information than simply the combination of parse trees for individual sentences. Navigating the graph along the edges of syntactic relations and the arcs of discourse relations allows a given parse jungle to be transformed into a semantically equivalent form to match other parse jungles, thereby performing the task of assessing textual similarity. To form a complete formal representation of a paragraph, as many links as possible are expressed. Each discourse arc generates a pair of jungle phrases that may be potential matches.
[0177] The topical similarity between the seed (request) and the response is expressed as a common subgraph of the parse forest. These are visualized as connected clouds. The greater the number of common graph nodes, the higher the similarity. For rhetorical consistency, the common subgraph does not need to be as large as it is in a given text. However, the rhetorical relationship and communicative actions of the seed and response are relevant and require correspondence.
[0178] Generalization of communicative actions
[0179] The similarity between two communication actions A1 and A2 is defined as abstract verbs that have common features between A1 and A2. Defining the similarity of two verbs as abstract verb-like structures supports inductive learning tasks, such as rhetorical consistency assessment. In the example, the similarity between the following two common verbs (agree and disagree) can be generalized as follows: agree^disagree=verb(Interlocutor,Proposed_action,Speaker), where the Interlocutor is the person who proposes the Proposed_action to the Speaker and to whom the Speaker conveys the Speaker's response. The Proposed_action is the action that the Speaker will perform if the Speaker accepts or rejects the request or suggestion, and the Speaker is the person to whom a specific action has been proposed and who responds to the request or suggestion made.
[0180] In another example, the similarity between the verbs agree and explain is expressed as follows: agree^explain=verb(Interlocutor,*,Speaker). The subject of a communicative action is generalized within the context of the communicative action and not with respect to other "physical" actions. Thus, aspects generalize individual occurrences of communicative actions along with the corresponding subject.
[0181] Furthermore, a sequence of communicative actions representing a conversation can be compared with other such sequences of similar conversations. In this way, the meaning of individual communicative actions and the dynamic discourse structure of the conversation (as opposed to its static structure reflected through rhetorical relations) are represented. Generalization occurs at every level of the composite structural representation. Tokens of communicative actions are generalized with tokens, and their semantic roles are generalized with corresponding semantic roles.
[0182] Textual authors use communicative acts to indicate the structure or conflict of a conversation. See Searle, JR. 1969, Speechacts: an essay in the philosophy of language. London: Cambridge University Press. The subject is generalized within the context of these acts, not relative to other "physical" acts. Thus, individual occurrences of communicative acts, along with their subjects and their pairs, are generalized into discourse "steps."
[0183] The generalization of communicative actions can also be considered from the perspective of matching verb frameworks (such as VerbNet). Communicative links reflect the discourse structure associated with the participation (or mention) of more than one agent in a text. These links form sequences that connect words used in communicative actions (multiple words or verbs that implicitly indicate a person's communicative intention).
[0184] A communicative action consists of an actor, one or more agents towards whom the action is being taken, and a phrase that characterizes the action. A communicative action can be described as a function of the form: verb(agent, subject, cause), where the verb characterizes some type of interaction between the agents involved (e.g., explanation, confirmation, reminder, disagreement, rejection, etc.), the subject refers to the information being conveyed or the object being described, and the cause refers to the motivation or explanation given to the agent.
[0185] A scenario (marked as a directed graph) is a subgraph of the parse jungle G = (V, A), where V = {action1, action2…action n} is a finite set of vertices corresponding to communication actions, and A is a finite set of labeled arcs (ordered pairs of vertices) classified as follows:
[0186] Each arc action i ,action j ∈A sequence corresponding to references to the same subject (e.g., s j =s i ) or two actions of different subjects v i ,agi ,s i ,c i and v j ,ag j ,s j ,c j Time priority of each arc action i ,action j ∈A cause Corresponding to action i and action j The attack relationship between i Causes and actions j The subject or cause of the conflict.
[0187] The subgraph of the parse jungle associated with the interaction scenario between agents has some obvious features. For example, (1) all vertices are ordered by time so that all vertices (except the initial and terminal vertices) have one incoming arc and one outgoing arc, (2) for A sequence arcs, at most one incoming arc and only one outgoing arc is admissible, and (3) for A cause Arcs, a given vertex can have many outgoing arcs, as well as many incoming arcs. The involved vertices can be associated with different agents or the same agent (i.e., when he contradicts himself). To compute the similarity between the parse jungle and its communication actions, inductive subgraphs, subgraphs with the same configuration of similar arc labels and strict correspondence of vertices, are analyzed.
[0188] By analyzing the arcs of the communication actions of the parsing jungle, the following similarities exist: (1) a communication action whose subject comes from T1 is compared with another communication action whose subject comes from T2 (without using the communication action arc), and (2) a pair of communication actions whose subject comes from T1 is compared with another pair of communication actions from T2 (with using the communication action arc).
[0189] Generalize two different communication actions based on their properties. See (Galitsky et al. 2013). Figure 14As can be seen in the example in question, one communicative action from T1, cheating(husband, wife, another lady), can be compared with a second communicative action from T2, avoid(husband, contact(husband, another lady)). The generalization results in communicative_action(husband, *), which introduces a constraint on A of the form: if a given agent (=husband) is mentioned as the subject of a CA in Q, then (s)he should also be the subject of (possibly another) CA in A. While two communicative actions can always be generalized, this is not true for their subjects: if their generalization result is empty, then the generalization result of the communicative action with these subjects is also empty.
[0190] Generalization of RST relations
[0191] Some relations between discourse trees can be generalized, such as arcs representing the same type of relations (representational relations, such as comparisons; subject relations, such as conditions; and multi-core relations, such as lists) can be generalized. The core or the situation represented by the core is indicated by "N". The satellite or the situation presented by the satellite is indicated by "S". "W" indicates the author. "R" indicates the reader (audience). Situations are propositions, completed actions or ongoing actions, as well as communication actions and states (including beliefs, wishes, approvals, explanations, reconciliations and others). The generalization of two RST relations with the above parameters is expressed as: rst1(N1,S1,W1,R1)^rst2(N2,S2,W2,R2)=(rst1^rst2)(N1^N2,S1^S2,W1^W2,R1^R2).
[0192] The text in N1, S1, W1, and R1 is generalized as a phrase. For example, rst1 ^ rst2 can be generalized as follows: (1) If relation_type(rst1) ! = relation_type(rst2), then the generalization is empty. (2) Otherwise, the signature of the rhetorical relation is generalized as a sentence: sentence(N1, S1, W1, R1) ^ sentence(N2, S2, W2, R2). See Iruskieta, Mikel, Iria da Cunha, and Maite Taboada. A qualitative comparison method for rhetorical structures: identifying different discourse structures in multilingual corpora. Lang Resources & Evaluation. June 2015, Vol. 49, No. 2.
[0193] For example, the meaning of rst-background^rst-enablement = (S increases R's ability to understand elements in N) ^ (R understands S increases R's ability to perform actions in N) = increase-VB the-DT ability-NN of-IN R-NNto-IN.
[0194] Since the relations rst-background^rst-enablement are different, the RST relation part is empty. Then, generalize the expression that is the language definition of the corresponding RST relation. For example, for each word or placeholder for a word such as agent, if the word is the same in each input phrase, keep the word (and its POS), if the word is different between these phrases, remove the word. The resulting expression can be interpreted as the common meaning between the formally obtained definitions of two different RST relations.
[0195] Figure 14The two arcs drawn between the question and the answer in Figure 1 show an example of a generalization based on the RST relation "RST-contrast". For example, "I just had a baby" is an RST contrast to "it does not look like me" and is related to "husband to avoid contact", which is an RST-contrast to "have the basic legal and financial commitments". As can be seen, the answer does not have to be similar to the verb phrase of the question, but the rhetorical structure of the question and answer is similar. Not all phrases in the answer have to match phrases in the question. For example, the unmatched phrases have certain rhetorical relations to phrases in the answer that are related to phrases in the question.
[0196] Building a communication discourse tree
[0197] Figure 15 An exemplary process for constructing a communication discourse tree according to one aspect is illustrated. The rhetorical consistency application 112 can implement the process 1500. As discussed, the communication discourse tree enables improved search engine results.
[0198] At block 1501, the process 1500 involves accessing a sentence comprising segments. At least one segment comprises a verb and a word, and each word comprises a role of the word within the segment, and each segment is a basic unit of discourse. For example, the rhetorical consistency application 112 accesses a sentence such as a sentence about Figure 13 Sentences such as "Rebels, the self-proclaimed Republic, deny that they controlled the territory from which the missile was allegedly fired" are described.
[0199] Continuing with this example, rhetorical consistency application 112 determines that the sentence includes several segments. For example, the first segment is "rebels deny." The second segment is "that they controlled the territory." The third segment is "from which the missile was allegedly fired." Each segment contains a verb, for example, "deny" in the first segment and "controlled" in the second segment. However, segments do not necessarily need to contain verbs.
[0200] At block 1502, process 1500 involves generating a discourse tree representing rhetorical relationships between sentence segments. The discourse tree includes nodes, each non-terminal node representing a rhetorical relationship between two of the sentence segments, and each terminal node of the discourse tree is associated with one of the sentence segments.
[0201] Continuing with this example, the rhetorical consistency application 112 generates Figure 13 The discourse tree shown. For example, the third fragment, “from which the missile was allegedly fired,” states “that they controlled the territory.” The second and third fragments together relate to the attribution of what happened, i.e., the attack could not have been by the rebels because they did not control the territory.
[0202] At block 1503, process 1500 involves accessing multiple verb signatures. For example, rhetoric consistency application 112 accesses a list of verbs in, for example, VerbNet. Each verb is matched or related to a verb in the segment. For example, for the first segment, the verb is "deny." Accordingly, rhetoric consistency application 112 accesses a list of verb signatures related to the verb "deny."
[0203] As discussed, each verb signature includes the verb of the segment and one or more topic roles. For example, a signature includes one or more of a noun phrase (NP), a noun (N), a communicative action (V), a verb phrase (VP), or an adverb (ADV). The topic roles describe the relationship between the verb and the related words. For example, "the teacher amused the children" has a different signature than "small children amuse quickly." For the first segment, the verb "deny," the rhetorical consistency application 112 accesses a list of verb signatures or frames that match "deny." The list is "NP V NP tobe NP," "NP Vthat S," and "NP V NP."
[0204] Each verb signature includes a topic role. A topic role refers to the role of a verb in a sentence fragment. Rhetorical consistency application 112 determines the topic role in each verb signature. Examples of topic roles include actor, agent, asset, attribute, beneficiary, cause, location, destination, source, destination, source, location, experiencer, scope, instrument, material and product, material, product, patient, predicate, recipient, stimulus, stem, time, or subject.
[0205] At block 1504, the process 1500 involves determining, for each of the verb signatures, the number of topic roles of the respective signature that match the roles of the words within the segment. For the first segment, the rhetorical consistency application 112 determines that the verb "deny" has only three roles: "agent," "verb," and "stem."
[0206] At block 1505, process 1500 involves selecting a particular verb signature from the verb signatures based on the particular verb signature having the greatest number of matches. For example, referring again to Figure 13 In the first segment “the rebels deny...that theycontrol the territory”, deny matches the verb signature deny “NP V NP”, and control matches control(rebel, territory). The verb signatures are nested, resulting in a nested signature “deny(rebel, control(rebel, territory))”.
[0207] Fictional Discourse Tree
[0208] Certain aspects described herein use a fictitious discourse tree to improve question-answering (Q / A) recall for complex, multi-sentence, convergent questions. By augmenting the answer's discourse tree with tree fragments obtained from an ontology, the aspects obtain a canonical discourse representation of the answer that is independent of the thought structure of a given author.
[0209] As discussed, a discourse tree (DT) can represent how a text is organized, particularly using text paragraphs. Discourse-level analysis can be used in many natural language processing tasks where learning language structure may be essential. The DT outlines the relationships between entities. There are different ways of introducing entities and associated attributes within a text. Not all of the rhetorical relationships that exist between these entities appear in the DT of a given paragraph. For example, a rhetorical relationship between two entities in another body of text may be outside the DT of a given paragraph.
[0210] Therefore, in order to make the question and answer fully relevant to each other, a more complete discourse tree for the answer is desired. Such a discourse tree would include all rhetorical relationships between the entities involved. To achieve this, the initial or best matching discourse tree of the answer that at least partially solves the question is augmented with certain rhetorical relationships that are identified as missing from the answer.
[0211] These missing rhetorical relations can be obtained from a text corpus or by searching an external database such as the internet. Therefore, rather than relying on an ontology that may have definitions of entities missing from the candidate answers, the parties mine the rhetorical relations between these entities online. This process avoids the need to construct an ontology and can be implemented on top of a conventional search engine.
[0212] More specifically, the baseline requirement for an answer A to be relevant to a question Q is that the entities of A (En) cover the entities of Q: Naturally, some answer entities EA are not explicitly mentioned in Q, but are required to provide a complete answer A. Therefore, the logical flow of Q through A can be analyzed. However, since establishing a relationship between En can be challenging, an approximate logical flow of Q and A can be modeled. This approximate logical flow can be expressed in domain-independent terms EnDT-Q ~ EnDT-A and later verified and / or expanded.
[0213] There may be different flaws in the initial answer used to answer the question. For example, one case is when some entities E are not explicitly mentioned in Q, but are assumed. Another case is when some entities in A used to answer Q do not appear in A, but more specific entities or more general entities appear in A. In order to determine that some more specific entities do solve the question from Q, some external or additional source (referred to as fictitious EnDT-A in this article) can be used to establish these mutual relations. Such a source contains information about the internal mutual relations between En, which are omitted in Q and / or A, but are assumed to be known to the peers. Therefore, for a computer-implemented system, the following knowledge is expected at the discourse level:
[0214] EnDT-Q ~ EnDT-A + fictional EnDT-A.
[0215] For the purpose of discussion, the following examples are introduced. The first example is the question "What is an advantage of an electric car?" and the corresponding answer is "No need to for gas". In order to establish that a specific answer is appropriate for a specific question, the generalized entity advantage and the general noun entity car are used. More specifically, the explicit entities in A{need, gas} are linked. A fragment of a possible fictional EnDT-A is shown: […Noneed…-Elaborate-Advantage]…[gas–Enablement-engine]…[engine–Enablement-car]. Only evidence of the existence of these rhetorical links is required.
[0216] In the second example, a fictitious discourse tree is used to improve search. The user asks the autonomous agent for information about "a faulty brake switch can affect the automatic transmission Munro". Existing search engines may identify certain keywords that are not found in a given search result. However, it is possible to indicate how these unidentified keywords are related to the search results by finding documents in which these keywords are rhetorically connected to the keywords that appear in the query. This feature naturally improves the relevance of the answer on the one hand and provides the user with interpretability about how their keywords are addressed in the answer.
[0217] The fictitious discourse tree enables the search engine to interpret missing keywords in the search results. In the default search, munro is missing. However, by attempting to rhetorically connect munro to the entities in the question, the autonomous agent application 102 learns that Munro is the inventor of the automatic transmission. Figure 16 An example process that may be used by the autonomous agent application 102 is depicted.
[0218] Figure 16 An exemplary process 1600 for constructing a fictitious discourse tree according to one aspect is illustrated. The autonomous agent application 102 can implement the process 1600. The process 1600 explains how a discourse tree (DT) facilitates improved matching between questions and answers. For example, to verify that answer A (or an initial answer) is appropriate for a given question Q, the autonomous agent application 102 verifies that DT-A and DT-Q are consistent, and then optionally augments DT-A with fragments of other DTs to ensure that all entities in Q are addressed in the augmented DT-A. For the purposes of this discussion, Figure 17 Process 1600 is discussed.
[0219] Figure 17 Depicted are example discourse trees of questions, answers, and two fictitious discourse trees according to one aspect. Figure 17 Depicted are a question utterance tree 1710 , an answer utterance tree 1720 , an additional answer utterance tree 1730 , an additional answer utterance tree 1740 , links 1750 - 1754 , and nodes 1170 - 1174 .
[0220] The question discourse tree 1710 is formed by the following questions: “[When driving the cruise control][the engine will turn off][when I want to accelerate,][although the check engine light was off.][I have turned on the ignition][and listen for the engine pump running][to see][ifit is building up vacuum.][Could there be a problem with the brake sensor under the dash?][Looks like there could be a little play in the plug.]”. For the purposes of discussion, square brackets [ ] indicate basic discourse units.
[0221] return Figure 16 At block 1601, the process 1600 involves constructing a question discourse tree (DT-Q) from a question. A question may include segments (or basic discourse units). Each word in the DT-Q indicates a role of the respective word within the segment.
[0222] The autonomous agent application 102 generates a question discourse tree. The question discourse tree represents rhetorical relationships between segments and includes nodes. Each non-terminal node represents a rhetorical relationship between two segments. Each terminal node in the discourse tree is associated with one of the segments.
[0223] The question utterance tree may include one or more question entities. The autonomous agent application 102 may identify these entities. Figure 17 As shown, the question contains entities such as "CRUISE CONTROL", "CHECK ENGINE LIGHT", "ENGINE PUMP", "VACUUM", "BRAKE SENSOR", etc. As discussed later, the answer addresses some of these entities, while some are not addressed.
[0224] At block 1602, the process 1600 accesses initial answers from a text corpus. The initial answers may be obtained from existing documents, such as files or databases. The initial answers may also be obtained by searching local or external systems. For example, the autonomous agent application 102 may obtain an additional set of answers by submitting a search query. The search query may be derived from the question, for example, by formulating one or more keywords identified in the question. The autonomous agent application 102 may determine a relevance score for each of the multiple answers. The answer score or ranking indicates the degree of match between the question and the corresponding answer. The autonomous agent application 102 selects an answer with a score above a threshold as the initial answer.
[0225] Continuing with the example, the autonomous agent application 102 accesses the following initial answer "[A faulty brake switch canaffect the cruise control.] torqueconverter is unlocking transmission.]”
[0226] At block 1603, the process 1600 involves constructing an answer utterance tree including answer entities from the initial answer. The autonomous agent application 102 forms the utterance tree for the initial answer. At block 1603, the autonomous agent application 102 performs similar steps as described with respect to block 1601.
[0227] Continuing with this example, the autonomous agent application 102 generates an answer utterance tree 1720, which represents the text that is the initial answer to the question. The answer utterance tree 1720 includes some of the question entities, specifically, "CRUISE CONTROL," "CHECK ENGINE LIGHT," and "TORQUE CONVERTER."
[0228] At block 1604, process 1600 involves determining, using a computing device, that a score indicating relevance of an answer entity to a question entity is below a threshold. An entity in question Q may be resolved (e.g., by one or more entities in the answer) or unresolved, i.e., no entity in answer A resolves a particular entity. Unresolved entities in the question are indicated by E0.
[0229] Different methods can be used to determine how relevant a particular answer entity is to a particular question entity. For example, in one aspect, the autonomous agent application 102 can determine whether an exact textual match of the question entity appears in the answer entity. In other cases, the autonomous agent application 102 can determine a score indicating the proportion of keywords of the question entity that appear in the answer entity.
[0230] There may be different reasons why entities are not resolved. As shown in the figure, E0-Q includes {ENGINE PUMP, BRAKESENSOR, and VACUUM}. For example, any answer A is not completely relevant to question Q because the answer omits some of the entities E0. Alternatively, the answer uses different entities instead. E0-Q may be omitted in answer A. To verify the latter possibility, background knowledge is used to find entities E linked to both E0-Q and EA. img .
[0231] For example, it may not be clear how EA=TORQUE CONVERTER connects to Q. To verify the connection, a text snippet about torque converters from Wikipedia is obtained and DT-A is constructed. img1 (1730). Aspects may determine that the torque converter is connected to the engine via the stated rhetorical relationship.
[0232] Therefore, EA = Torque Converter is indeed relevant to the question, as indicated by the vertical blue arc. This determination can be made without building an offline ontology of linked entities and learning the relationships between them. Instead, the utterance-level context is used to confirm that A includes relevant entities.
[0233] Continuing with this example, the autonomous agent application 102 determines that other entities in the question utterance tree 1710, specifically those referring to "ENGINE PUMP," "VACUUM," "BRAKE SENSOR," and "TORQUE CONVERTER," are not resolved in the answer utterance tree 1720. Accordingly, these correspondences between EQ and EA are illustrated by links 1750 and 1751, respectively.
[0234] At block 1605, the process 1600 involves creating additional utterance trees based on the text corpus. The additional utterance trees are created from the text. In some cases, the autonomous agent application 102 selects an appropriate text from the one or more texts based on a scoring mechanism.
[0235] For example, the autonomous agent application 102 may access an additional set of answers. The autonomous agent application 102 may generate an additional set of fictitious utterance trees. Each additional utterance tree corresponds to a corresponding additional answer. The autonomous agent application 102 calculates a score for each additional utterance tree, the score indicating the number of question entities that include mappings to one or more answer entities in the corresponding additional utterance tree. The autonomous agent application 102 selects the additional utterance tree with the highest score from the set of additional utterance trees. Once the fictitious DT-A img DT-A is extended to measure search relevance as the inverse of unsolved E0-Q.
[0236] For example, the autonomous agent application 102 obtains the candidate set A s Then, for A s Each candidate A in c , the autonomous agent application 102 performs the following operations:
[0237] (a) Construction of DT-A c ;
[0238] (b) Establishing the mapping EQ->EA c ;
[0239] (c) Identify E0-Q;
[0240] (d) From E0-Q and E0-A c (Entities not in E0-Q) form a query;
[0241] (e) Obtain search results for query d) from B and construct the fictitious DTs-A c ;as well as
[0242] (f) Calculate the remaining fraction |E0|.
[0243] The autonomous agent application 102 may then select A having the best score.
[0244] In some cases, machine learning methods can be used to<EDT-Q,EDT-A> A pair is classified as correct or incorrect. The example training set includes both good (positive) and bad (negative) Q / A pairs. Therefore, a DT kernel learning method (SVM TK, Joty and Moschitti 2014, Galitsky 2017) was chosen, which applies SVM learning to the set of all sub-DTs of the DT of the Q / A pair. Tree kernel family methods are not very sensitive to errors in parsing (grammar and rhetoric) because incorrect subtrees are mostly random and unlikely to be common across different elements of the training set.
[0245] DT can be represented by a vector V of integer counts of each subtree type (regardless of its ancestry):
[0246] V(T) = (number of type 1 subtrees, ...).
[0247] Given two tree segments DT1 and DT2, the tree kernel function K(EDT1,EDT2)≤V(EDT1) and
[0248]
[0249] where n1∈N1,n2∈N2, where N1 and N2 are the sets of all nodes in DT1 and DT2 respectively.
[0250] I1(n) is the indicator function.
[0251] I1(n) = {1, a subtree of type i is present and rooted at the node; otherwise 0}
[0252] Continuing with this example, it is not clear how the EQ "ENGINE PUMP" in the question is resolved in the initial answer. Therefore, the autonomous agent application 102 can determine additional resources that can resolve the ENGINE PUMP entity.
[0253] At block 1606, the process 1600 involves determining that the additional utterance tree includes a rhetorical relationship connecting the question entity to the answer entity. Continuing with the example, Figure 17 As shown, node 1770 identifies the entity "VACUUM," and node 1171 identifies the entity "ENGINE." Therefore, nodes 1172 and 1173, both representing the rhetorical relationship "elaboration," are related to nodes 1170 and 1171. In this way, nodes 1172 and 1773 provide the missing link between "ENGINE" in the question and "VACUUM," which is in the question but not previously resolved.
[0254] The autonomous agent application 102 connects the additional discourse tree to the EA. As shown in the figure, DT-Aimg2 VACUUM and ENGINE are connected via elaboration. Thus, the combined DT-A consists of the real DT-A plus DT-A img1 and DT-A img2 By employing background knowledge in a domain-independent manner, both real and imaginary DTs are necessary to prove that the answer is relevant.
[0255] At block 1607, process 1600 involves extracting a subtree of the additional discourse tree that includes the question entity, the answer entity, and the rhetorical relationship, thereby generating a fictitious discourse tree. In some cases, the autonomous agent application 102 may extract a subtree or portion of the additional discourse tree that relates the question entity to the answer entity. The subtree includes at least one node. In some cases, the autonomous agent application 102 may integrate the subtree from the fictitious discourse tree into the answer discourse tree by connecting the node to the answer entity or another entity.
[0256] At block 1608, the process 1600 involves outputting an answer represented by a combination of the answer utterance tree and the fictitious utterance tree. The autonomous agent application 102 may combine the subtree identified at block 1607 with the answer utterance tree. The autonomous agent application 102 may then output an answer corresponding to the combined tree.
[0257] Experimental results
[0258] Traditional Q / A datasets for factual and non-factual questions, as well as SemEval and neural Q / A evaluations, are not suitable because the questions are ranked and not complex enough to observe the potential contribution of discourse-level analysis. For evaluation, two convergent Q / A sets are formed:
[0259] 1. The Yahoo! Answers (Webscope 2017) set of question-answer pairs covering a wide range of topics. We selected 3,300 user questions consisting of three to five sentences from the 140k set. Most questions had fairly detailed answers, so we did not filter answers by sentence length.
[0260] 2. Car repair dialogue including 9,300 Q / A pairs describing car problems and suggestions on how to correct them.
[0261] For each of these sets, we form positive pairs from actual Q / A pairs and similar-entities Forming Negative Pairs: EA similar-entities There is a strong overlap with EA, but A similar-entities There is no real correct, comprehensive and definitive answer. Therefore, Q / A is simplified to a classification task measured by the precision and recall of associating Q / A pairs to the class of the correct pair.
[0262]
[0263] Evaluation of Q / A Accuracy
[0264] The first two rows in Table 1 show the baseline performance of Q / A and show that in complex domains, the transformation from keywords to matched entities brings more than 12% performance improvement. The bottom three rows show the Q / A accuracy when applying discourse analysis. Ensuring rule-based correspondence between DT-A and DT-Q provides a 12% increase relative to the baseline, and using a fictitious DT provides a further 10% increase. Finally, moving from rule-based to machine-learned Q / A correspondence (SVM TK) gives a performance gain of about 7%.
[0265] Supplementing fictional discourse trees with rhetorical consistency classifiers
[0266] By using the exchange discourse tree, the rhetorical consistency application 112 can determine the complementarity between two sentences. For example, the rhetorical consistency application 112 can determine the degree of complementarity between the discourse tree of the question and the discourse tree of the initial answer, between the discourse tree of the question and the discourse tree of the additional or candidate answer, or between the discourse tree of the answer and the discourse tree of the additional answer. In this way, the autonomous agent application 102 ensures that complex questions are addressed through answers that are complete in terms of rhetorical consistency or style.
[0267] In an example, the rhetorical consistency application 112 constructs a question communication discourse tree from the question and constructs an answer communication discourse tree from the initial answer. The rhetorical consistency application 112 determines a question communication discourse tree for the question statement. The question discourse tree may include a root node. For example, referring again to Figure 13 and Figure 15 , an example question sentence is "are rebels responsible for the downing of the flight". The rhetoric classification application 102 can use Figure 15 Described process 1500. The example question has a root node of "elaboration".
[0268] The rhetorical consistency application 112 determines a second communication discourse tree for the answer statement. The answer communication discourse tree may include a root node. Continuing with the above example, the rhetorical consistency application creates a communication discourse tree such as Figure 13 As shown, the tree also has a root node labeled "elaboration".
[0269] The rhetorical consistency application 112 correlates the communication discourse trees by identifying that the question root node and the answer root node are the same. The rhetorical consistency application 112 determines that the question communication discourse tree and the answer communication discourse tree have the same root node. The resulting correlated communication discourse tree is as follows: Figure 17as shown, and can be labeled as a "request-response pair."
[0270] The rhetorical consistency application 112 calculates the degree of complementarity between the question exchange discourse tree and the answer exchange discourse tree by applying a predictive model to the merged discourse trees. Various machine learning techniques can be used. In one aspect, the rhetorical consistency application 112 trains and uses a rhetorical consistency classifier 120. For example, the rhetorical consistency application 112 can define positive and negative categories for request-response pairs. The positive category includes rhetorically correct request-response pairs, and the negative category includes related but rhetorically external request-response pairs. For each request-response pair, the rhetorical consistency application 112 can construct a CDT by parsing each sentence and obtaining the verb signature of the sentence fragment. The rhetorical consistency application 112 provides the associated communication discourse tree pairs to the rhetorical consistency classifier 120, which in turn outputs the degree of complementarity.
[0271] Rhetorical consistency application 112 determines that the degree of complementarity is above a threshold, and then identifies the question and answer statements as complementary. Rhetorical consistency application 112 may use the degree of complementarity threshold to determine whether the question-answer pair is sufficiently complementary. For example, if the classification score is greater than the threshold, then rhetorical consistency application 112 may use the answer. Alternatively, rhetorical consistency application 112 may discard the answer and access answer database 105 or a public database for another candidate answer, and repeat as needed.
[0272] In another aspect, the rhetorical consistency application 112 applies jungle kernel learning to the representation. Jungle kernel learning can occur in place of the classification-based learning described above. The rhetorical consistency application 112 constructs a parse jungle pair for the parse tree of the request-response pair. The rhetorical consistency application 112 applies discourse parsing to obtain a discourse tree pair for the request-response pair. The rhetorical consistency application 112 aligns the basic discourse units of the discourse tree request-response and the parse tree request-response. The rhetorical consistency application 112 merges the basic discourse units of the discourse tree request-response and the parse tree request-response.
[0273] Related work
[0274] At any point in a discourse, some entities are considered more important than others (appearing in the core of the DT) and are therefore expected to exhibit different properties. In central theory (Grosz et al., 1995; Poesio et al., 2004), entity importance determines how they are realized in speech, including the pronominal relationships between them. In other discourse theories, entity importance can be defined in terms of topicality (Prince 1978) and cognitive accessibility (Gundel et al., 1993).
[0275] Barzilay and Lapata (2008) automatically abstract text into a collection of entity transition sequences and record distributional, syntactic, and reference information about utterance entities. The authors formulate consistency assessment as a learning task and show that their entity-based representation is well-suited for ranking-based generation and text classification tasks.
[0276] (Nguyen and Joty 2017) proposed a local coherence model based on a convolutional neural network that operates on a distributed representation of entity transitions in a grid representation of text. The local coherence model can model sufficiently long entity transitions and can incorporate entity-specific features without losing generalization ability. Kuyten et al. (2015) developed a search engine that exploits the discourse structure in documents to overcome the limitations associated with bag-of-words document representations in information retrieval. The system does not solve the rhetorical coordination problem between Q and A, but given a Q, the search engine can retrieve related A and separate sentences from A that describe some rhetorical relationship to the query.
[0277] Answering questions in this research area is a significantly more complex task than factual QA, such as the Stanford QA Database (Rajpurkar et al., 2016), which only involves one or two entities and their parameters. To answer "how to solve the problem" questions, it is necessary to maintain a logical flow connecting the entities in the question. Since some entities in the Q are inevitably omitted, these entities may need to be recovered from some background knowledge text about these omitted entities and the entities presented in the Q. In addition, the logical flow needs to complement the logical flow of the Q.
[0278] Domain-specific ontologies, such as those related to the mechanics of automobiles, are difficult and expensive to construct. In this work, we propose an alternative approach based on domain-independent, discourse-level analysis. More specifically, we address the unresolved aspects of DT-A by finding text snippets in background knowledge corpora, such as Wikipedia. This eliminates the need for an ontology that must maintain relationships between the involved entities.
[0279] The fictitious DT features of the proposed Q / A system provide a significant improvement in the accuracy of answering complex convergent questions. However, compared to the baseline focusing on relevance, DT for answer style matching improves Q / A accuracy by more than 10%, and relying on fictitious DT improves it by another 10%.
[0280] Aspects described herein analyze the complementary relationship between DT-A and DT-Q, thereby significantly reducing the learning feature space, making it reasonable to learn from available datasets of limited size, such as car maintenance lists.
[0281] Figure 18 A simplified diagram of a distributed system 1800 for implementing one of these aspects is depicted. In the illustrated aspect, the distributed system 1800 includes one or more client computing devices 1802, 1804, 1806, and 1808 configured to execute and operate client applications, such as web browsers, proprietary clients (e.g., Oracle Forms), etc., over one or more networks 1810. A server 1812 can be communicatively coupled to the remote client computing devices 1802, 1804, 1806, and 1808 via the network 1810.
[0282] In various aspects, server 811 can be suitable for running one or more services or software applications provided by one or more components of the system.Services or software applications can include non-virtual and virtual environments.Virtual environments can include environments for virtual events, exhibitions, simulators, classrooms, shopping trading places and enterprises, no matter whether they are two-dimensional or three-dimensional (3D) representations, page-based logical environments or other forms.In some aspects, these services can be provided to the users of client computing devices 1802, 1804, 1806 and / or 1808 as web-based services or cloud services or under software as a service (SaaS) model.The users of operating client computing devices 1802, 1804, 1806 and / or 1808 can utilize one or more client applications to interact with server 1812 in turn, to utilize the services provided by these components.
[0283] In the configuration depicted in the figure, software components 1818, 1820, and 1822 of distributed system 1800 are shown as being implemented on server 1812. In other aspects, one or more components of distributed system 1800 and / or the services provided by these components may also be implemented by one or more of client computing devices 1802, 1804, 1806, and / or 1808. Users operating the client computing devices can then utilize one or more client applications to use the services provided by these components. These components may be implemented in hardware, firmware, software, or a combination thereof. It should be appreciated that a variety of different system configurations are possible, which may differ from distributed system 1800. Therefore, the aspects shown in the figure are an example of a distributed system for implementing an aspect system and are not intended to be limiting.
[0284] Client computing devices 1802, 1804, 1806, and / or 1808 may be portable handheld devices (e.g., Cellular phones, computing tablets, personal digital assistants (PDAs), or wearable devices (e.g., Google head-mounted displays), which run systems such as Microsoft Windows and / or software for various mobile operating systems (such as iOS, Windows Phone, Android, BlackBerry 18, Palm OS, etc.), and enabling Internet, email, short message service (SMS), or other communication protocols. The client computing device may be a general-purpose personal computer, including, for example, a computer running various versions of Microsoft Apple The client computing device may be a personal computer and / or laptop computer running any of a variety of commercially available Alternatively or additionally, the client computing devices 1802, 1804, 1806, and 1808 may be any other electronic device capable of communicating over the network(s) 1810, such as a thin client computer, an Internet-enabled gaming system (e.g., with or without a PC or other computer). A gesture input device (Microsoft Xbox game console) and / or a personal messaging device.
[0285] Although exemplary distributed system 1800 is shown with four client computing devices, any number of client computing devices may be supported. Other devices (such as devices with sensors, etc.) may interact with server 1812.
[0286] The network(s) 1810 in the distributed system 1800 may be any type of network familiar to those skilled in the art that supports data communications using any of a variety of commercially available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internet Message Exchange), AppleTalk, and the like. By way of example only, the network(s) 1810 may be a local area network (LAN), such as a LAN based on Ethernet, Token Ring, or the like. The network(s) 1810 may be a wide area network and the Internet. It may include virtual networks, including but not limited to virtual private networks (VPNs), intranets, extranets, public switched telephone networks (PSTNs), infrared networks, wireless networks (e.g., in accordance with the Institute of Electrical and Electronics Engineers (IEEE) 802.18 protocol suite, and / or any other wireless protocol); and / or any combination of these and / or other networks.
[0287] The server 1812 may be composed of one or more general-purpose computers, dedicated server computers (including, for example, PC (personal computer) servers, The server 1812 may be a server, a mid-range server, a mainframe computer, a rack-mounted server, etc.), a server farm, a server cluster, or any other suitable arrangement and / or combination. The server 1812 may include one or more virtual machines running a virtual operating system or other computing architecture involving virtualization. A flexible pool of one or more logical storage devices may be virtualized to maintain virtual storage devices for the server. The server 1812 may use software-defined networking to control the virtual network. In various aspects, the server 1812 may be suitable for running one or more services or software applications described in the foregoing disclosure. For example, the server 1812 may correspond to a server for performing the processing described above according to aspects of the present disclosure.
[0288] The server 1812 can run an operating system including any of the operating systems discussed above, as well as any commercially available server operating system. The server 1812 can also run any of a variety of additional server applications and / or middle-tier applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, Servers, database servers, etc. Exemplary database servers include, but are not limited to, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), etc.
[0289] In some implementations, server 1812 may include one or more applications to analyze and integrate data feeds and / or event updates received from users of client computing devices 802, 804, 806, and 808. By way of example, data feeds and / or event updates may include, but are not limited to, feed, Real-time updates and continuous data streams received from one or more third-party information sources may include real-time events related to sensor data applications, financial quote machines, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc. Server 1812 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client computing devices 1802, 1804, 1806, and 1808.
[0290] Distributed system 1800 may also include one or more databases 1814 and 1816. Databases 1814 and 1816 may reside in various locations. As an example, one or more of databases 1814 and 1816 may reside on a non-transient storage medium local to server 1812 (and / or residing in server 1812). Alternatively, databases 1814 and 1816 may be located away from server 1812 and communicate with server 1812 via a network-based connection or a dedicated connection. In one set of aspects, databases 1814 and 1816 may reside in a storage area network (SAN). Similarly, any necessary files for executing the functions possessed by server 1812 may be appropriately stored locally on server 1812 and / or stored remotely. In one set of aspects, databases 1814 and 1816 may include a relational database suitable for storing, updating, and retrieving data in response to SQL formatted commands, such as a database provided by Oracle.
[0291] Figure 19 1 is a simplified block diagram of one or more components of a system environment 1900 according to aspects of the present disclosure, through which services provided by one or more components of the aspect system can be provided as cloud services. In the illustrated aspect, the system environment 1900 includes one or more client computing devices 1904, 1906, and 1908 that can be used by users to interact with a cloud infrastructure system 1902 that provides cloud services. The client computing devices can be configured to operate client applications, such as web browsers, proprietary client applications (e.g., Oracle Forms), or some other application, that can be used by users of the client computing devices to interact with the cloud infrastructure system 1902 to use the services provided by the cloud infrastructure system 1902.
[0292] It should be appreciated that the cloud infrastructure system 1902 depicted in the figure can have other components in addition to those depicted. Furthermore, the aspects shown in the figure are merely one example of a cloud infrastructure system that can incorporate aspects of the present invention. In some other aspects, the cloud infrastructure system 1902 can have more or fewer components than shown in the figure, can combine two or more components, or can have a different configuration or arrangement of components.
[0293] Client computing devices 1904 , 1906 , and 1908 may be similar devices to those described above with respect to 1002 , 1004 , 1006 , and 1008 .
[0294] While the exemplary system environment 1900 is shown with three client computing devices, any number of client computing devices may be supported. Other devices such as devices with sensors, etc. may interact with the cloud infrastructure system 1902.
[0295] The network(s) 1910 can facilitate the communication and exchange of data between the client computing devices 1904, 1906, and 1908 and the cloud infrastructure system 1902. Each network can be any type of network familiar to those skilled in the art that can support data communication using any of a variety of commercially available protocols, including those described above for the network(s) 1810.
[0296] Cloud infrastructure system 1002 may include one or more computers and / or servers, which may include those described above with respect to server 1812 .
[0297] In some aspects, the services provided by the cloud infrastructure system may include a number of services available on demand to users of the cloud infrastructure system, such as online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database processing, managed technical support services, and the like. The services provided by the cloud infrastructure system can be dynamically scaled to meet the needs of users of the cloud infrastructure system. A specific instantiation of a service provided by a cloud infrastructure system is referred to herein as a "service instance." In general, any service that is available to a user from a cloud service provider's system via a communication network (such as the Internet) is referred to as a "cloud service." Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own local servers and systems. For example, a cloud service provider's system can host applications, and users can subscribe to and use the applications on demand via a communication network such as the Internet.
[0298] In some examples, services in a computer network cloud infrastructure may include protected computer network access to storage devices, hosted databases, hosted web servers, software applications, or other services provided by the cloud provider to users, or as otherwise known in the art. For example, a service may include password-protected access to remote storage devices on the cloud via the Internet. As another example, a service may include a hosted relational database and scripting language middleware engine based on web services for private use by networked developers. As another example, a service may include access to an email software application hosted on the cloud provider's website.
[0299] In certain aspects, the cloud infrastructure system 1902 may include a suite of application, middleware, and database service offerings delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. An example of such a cloud infrastructure system is the Oracle Public Cloud provided by the present assignee.
[0300] Large amounts of data (sometimes referred to as big data) can be hosted and / or manipulated by infrastructure systems at many levels and scales. This data can include datasets that are so large and complex that they are difficult to process using typical database management tools or traditional data processing applications. For example, using personal computers or their rack-based counterparts can be difficult to store, retrieve, and process terabytes of data. Data of this size can be difficult to process using the latest relational database management systems and desktop statistics and visualization packages. They can require massively parallel processing software running on thousands of server computers, exceeding the structure of commonly used software tools, to capture, organize, manage, and process the data within a tolerable elapsed time.
[0301] Analysts and researchers can store and manipulate very large data sets to visualize large amounts of data, detect trends, and / or otherwise interact with the data. Dozens, hundreds, or thousands of processors linked in parallel can operate on such data to present the data or simulate external forces on the data or the things it represents. These data sets may involve structured data (e.g., data organized in a database or otherwise organized according to a structured model) and / or unstructured data (e.g., emails, images, data blobs (binary large objects), web pages, complex event processing). By leveraging one aspect's ability to relatively quickly focus more (or fewer) computing resources on a single target, cloud infrastructure systems can be better utilized to perform tasks on large data sets based on the needs of enterprises, government agencies, research organizations, private individuals, like-minded individuals or organizations, or other entities.
[0302] In various aspects, the cloud infrastructure system 1002 can be adapted to automatically provision, manage, and track customer subscriptions to services offered by the cloud infrastructure system 1902. The cloud infrastructure system 1002 can provide cloud services via different deployment models. For example, services can be provided according to a public cloud model, in which the cloud infrastructure system 1002 is owned by the organization selling the cloud service (e.g., owned by Oracle), and the services are available to the general public or businesses in different industries. As another example, services can be provided according to a private cloud model, in which the cloud infrastructure system 1002 operates only for a single organization and can provide services to one or more entities within that organization. Cloud services can also be provided according to a community cloud model, in which the cloud infrastructure system 1002 and the services provided by the cloud infrastructure system 1002 are shared by several organizations in a related community. Cloud services can also be provided according to a hybrid cloud model, which is a combination of two or more different models.
[0303] In some aspects, the services provided by the cloud infrastructure system 1002 may include one or more services provided under the Software as a Service (SaaS) category, the Platform as a Service (PaaS) category, the Infrastructure as a Service (IaaS) category, or other service categories including hybrid services. A customer may subscribe to one or more services provided by the cloud infrastructure system 1902 via a subscription order. The cloud infrastructure system 1002 then performs processing to provide the services in the customer's subscription order.
[0304] In some aspects, the services provided by the cloud infrastructure system 1002 may include, but are not limited to, application services, platform services, and infrastructure services. In some examples, application services may be provided by the cloud infrastructure system via a SaaS platform. The SaaS platform may be configured to provide cloud services that fall into the SaaS category. For example, a SaaS platform may provide the ability to build and deliver on-demand application suites on an integrated development and deployment platform. The SaaS platform may manage and control the underlying software and infrastructure used to provide SaaS services. By utilizing the services provided by the SaaS platform, customers may utilize applications executed on the cloud infrastructure system. Customers may obtain application services without the need for the customer to purchase separate licenses and support. A variety of different SaaS services may be provided. Examples include, but are not limited to, services that provide solutions for sales performance management, enterprise integration, and business agility to large organizations.
[0305] In some aspects, platform services can be provided by a cloud infrastructure system via a PaaS platform. The PaaS platform can be configured to provide cloud services that fall into the PaaS category. Examples of platform services can include, but are not limited to, services that enable organizations (such as Oracle) to integrate existing applications on a shared public architecture and to fully utilize the shared services provided by the platform to build new applications. The PaaS platform can manage and control the underlying software and infrastructure used to provide PaaS services. Customers can obtain PaaS services provided by the cloud infrastructure system without the need for customers to purchase separate licenses and support. Examples of platform services include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), etc.
[0306] By utilizing the services provided by the PaaS platform, customers can adopt programming languages and tools supported by the cloud infrastructure system and also control the deployed services. In some aspects, the platform services provided by the cloud infrastructure system may include database cloud services, middleware cloud services (e.g., Oracle Fusion Middleware Services), and Java cloud services. In one aspect, the database cloud service may support a shared service deployment model that enables organizations to pool database resources and supply databases as a service to customers in the form of a database cloud. In the cloud infrastructure system, the middleware cloud service can provide customers with a platform for developing and deploying various business applications, and the Java cloud service can provide customers with a platform for deploying Java applications.
[0307] A variety of infrastructure services can be provided by IaaS platforms in cloud infrastructure systems. Infrastructure services facilitate the management and control of underlying computing resources (such as storage, network, and other basic computing resources) for customers to utilize the services provided by SaaS and PaaS platforms.
[0308] In certain aspects, the cloud infrastructure system 1002 may also include infrastructure resources 1930 for providing resources for providing various services to customers of the cloud infrastructure system. In one aspect, the infrastructure resources 1930 may include a combination of pre-integrated and optimized hardware (such as servers, storage devices, and networking resources) to perform the services provided by the PaaS platform and the SaaS platform.
[0309] In some aspects, resources in the cloud infrastructure system 1002 can be shared by multiple users and dynamically reallocated as needed. Furthermore, resources can be allocated to users in different time zones. For example, the cloud infrastructure system 1002 can enable a first group of users in a first time zone to utilize the cloud infrastructure system's resources for a specified number of hours, and then enable the same resources to be reallocated to another group of users in a different time zone, thereby maximizing resource utilization.
[0310] In certain aspects, a plurality of internal shared services 1932 may be provided that are shared by different components or modules of the cloud infrastructure system 1902 and services provided by the cloud infrastructure system 1902. These internal shared services may include, but are not limited to: security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelisting services, high availability, backup and recovery services, cloud-enabled services, email services, notification services, file transfer services, and the like.
[0311] In certain aspects, the cloud infrastructure system 1902 can provide comprehensive management of cloud services (e.g., SaaS, PaaS, and IaaS services) in the cloud infrastructure system. In one aspect, cloud management functionality can include, among other things, the ability to provision, manage, and track customer subscriptions received by the cloud infrastructure system 1902.
[0312] In one aspect, as depicted in the figure, cloud management functionality may be provided by one or more modules, such as order management module 1920, order orchestration module 1922, order provisioning module 1924, order management and monitoring module 1926, and identity management module 1928. These modules may include or be provided using one or more computers and / or servers, which may be general purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0313] In example operation 1934, a customer using a client computing device (such as client computing device 1904, 1906, or 1908) may interact with cloud infrastructure system 1902 by requesting one or more services provided by cloud infrastructure system 1902 and placing a subscription order for the one or more services offered by cloud infrastructure system 1902. In certain aspects, the customer may access a cloud user interface (UI) (cloud UI 1919, cloud UI 1914, and / or cloud UI 1916) and place a subscription order via these UIs. Order information received by cloud infrastructure system 1902 in response to the customer placing the order may include information identifying the customer and the one or more services offered by cloud infrastructure system 1902 to which the customer wishes to subscribe.
[0314] After a customer places an order, order information is received via cloud UIs 1010 , 1014 and / or 1014 .
[0315] At operation 1936, the order is stored in order database 1918. Order database 1918 may be one of several databases operated by cloud infrastructure system 1902 and in conjunction with other system elements.
[0316] At operation 1938, the order information is forwarded to the order management module 1920. In some cases, the order management module 1920 may be configured to perform billing and accounting functions related to the order, such as verifying the order and, upon verification, booking the order.
[0317] At operation 1940, information about the order is transmitted to order orchestration module 1922. Order orchestration module 1922 can use the order information to orchestrate the provisioning of services and resources for the order placed by the customer. In some cases, order orchestration module 1922 can use the services of order provisioning module 1924 to orchestrate the provisioning of resources to support the subscribed services.
[0318] In certain aspects, the order orchestration module 1922 enables management of the business processes associated with each order and applies business logic to determine whether the order should proceed to provisioning. At operation 1942, upon receiving an order for a new subscription, the order orchestration module 1922 sends a request to the order provisioning module 1924 to allocate resources and configure those resources required to fulfill the subscription order. The order provisioning module 1924 enables allocation of resources for the services ordered by the customer. The order provisioning module 1924 provides an abstraction layer between the cloud services provided by the cloud infrastructure system 1902 and the physical implementation layer for provisioning resources for providing the requested services. Thus, the order orchestration module 1922 can be isolated from implementation details such as whether services and resources are actually provisioned immediately or pre-provisioned and allocated / assigned only upon request.
[0319] At operation 1944 , once the services and resources are provisioned, a notification of the provided services may be sent to the customer on the client computing device 1904 , 1906 , and / or 1908 via the order provisioning module 1924 of the cloud infrastructure system 1902 .
[0320] At operation 1946, the order management and monitoring module 1926 can manage and track the customer's subscription order. In some cases, the order management and monitoring module 1926 can be configured to collect usage statistics for the services in the subscription order, such as the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and system downtime.
[0321] In certain aspects, the cloud infrastructure system 1902 can include an identity management module 1928. The identity management module 1928 can be configured to provide identity services, such as access management and authorization services within the cloud infrastructure system 1902. In some aspects, the identity management module 1928 can control information about clients that wish to utilize services provided by the cloud infrastructure system 1902. Such information can include information authenticating the identities of these clients and information describing which actions these clients are authorized to perform with respect to various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). The identity management module 1928 can also include management of descriptive information about each client and how and by whom this descriptive information can be accessed and modified.
[0322] Figure 20 An exemplary computer system 2000 is shown in which various aspects of the present invention may be implemented. System 2000 may be used to implement any of the computer systems described above. As shown, computer system 2000 includes a processing unit 2004 that communicates with multiple peripheral subsystems via a bus subsystem 2002. These peripheral subsystems may include a processing acceleration unit 2006, an I / O subsystem 2008, a storage subsystem 2018, and a communication subsystem 2024. Storage subsystem 2018 includes a computer-readable storage medium 2022 and system memory 2010.
[0323] The bus subsystem 2002 provides a mechanism for allowing the various components and subsystems of the computer system 2000 to communicate with each other by intention. Although the bus subsystem 2002 is schematically shown as a single bus, the alternative aspect of the bus subsystem can utilize multiple buses. The bus subsystem 2002 can be any of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, and a local bus using any various bus architectures. For example, this architecture can include an industry standard architecture (ISA) bus, a microchannel architecture (MCA) bus, an enhanced ISA (EISA) bus, a video electronics standards association (VESA) local bus, and a peripheral component interconnect (PCI) bus, which can be implemented as a Mezzanine bus manufactured by the IEEE P2086.1 standard.
[0324] The processing unit 2004 that can be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers) controls the operation of the computer system 2000. One or more processors can be included in the processing unit 2004. These processors can include single-core or multi-core processors. In some aspects, the processing unit 2004 can be implemented as one or more independent processing units 2032 and / or 2034, wherein each processing unit includes a single-core or multi-core processor. In other aspects, the processing unit 2004 can also be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0325] In various aspects, the processing unit 2004 can execute various programs in response to program code and can maintain multiple concurrently executing programs or processes. At any given time, some or all of the program code to be executed can reside in (one or more) processing units 2004 and / or in the storage subsystem 2018. Through appropriate programming, (one or more) processing units 2004 can provide the various functions described above. The computer system 2000 can additionally include a processing acceleration unit 2006, which can include a digital signal processor (DSP), a special-purpose processor, etc.
[0326] I / O subsystem 2008 may include user interface input devices and user interface output devices. User interface input devices may include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touch screen incorporated into a display, a scroll wheel, a click wheel, a dial, buttons, switches, a keypad, an audio input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices may include, for example, motion sensing and / or gesture recognition devices such as Microsoft Motion sensors that enable users to control devices such as Microsoft 360 game controller. The user interface input device may also include an eye gesture recognition device, such as detecting eye activity from the user (e.g., a "wink" when taking a picture and / or making a menu selection) and translating the eye gesture to the input device (e.g., Google ) in the input Google In addition, the user interface input device may include an input device that enables the user to communicate with the voice recognition system (e.g., Navigator) interactive voice recognition sensing device.
[0327] The user interface input device may also include, but is not limited to, a three-dimensional (3D) mouse, a joystick or pointing stick, a game panel and a drawing board, and audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and sight tracking devices. In addition, the user interface input device may include, for example, a medical imaging input device such as computed tomography, magnetic resonance imaging, positron emission tomography, medical ultrasound equipment. The user interface input device may also include, for example, an audio input device such as a MIDI keyboard, a digital musical instrument, etc.
[0328] The user interface output devices may include a display subsystem, indicator lights, or non-visual displays such as audio output devices, among others. The display subsystem may be a cathode ray tube (CRT), a flat panel device such as one using a liquid crystal display (LCD) or plasma display, a projection device, a touch screen, and the like. In general, the use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from the computer system 2000 to a user or to another computer. For example, user interface output devices may include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, voice output devices, and modems.
[0329] Computer system 2000 may include a storage subsystem 2018 containing software elements, shown currently located in system memory 2010. System memory 2010 may store program instructions that may be loaded and executed on processing unit 2004, as well as data generated during execution of these programs.
[0330] Depending on the configuration and type of computer system 2000, system memory 2010 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that are immediately accessible to and / or currently being operated and executed by processing unit 2004. In some implementations, system memory 2010 may include multiple different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS), which contains basic routines that help transfer information between elements within computer system 2000, such as during startup, may typically be stored in ROM. By way of example, but not limitation, system memory 2010 also shows application programs 2012, program data 2014, and an operating system 2016. Application programs 2012 may include client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc. As an example, operating system 2016 may include various versions of Microsoft Apple and / or Linux operating systems, various commercially available or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google operating systems, etc.) and / or such as iOS, Phone, OS, 10OS and OS operating system such as mobile operating system.
[0331] The storage subsystem 2018 may also provide a tangible computer-readable storage medium for storing basic programming and data structures that provide some aspects of the functionality. Software (programs, code modules, instructions) that provide the above-described functionality when executed by the processor may be stored in the storage subsystem 2018. These software modules or instructions may be executed by the processing unit 2004. The storage subsystem 2018 may also provide a repository for storing data used in accordance with the present invention.
[0332] The storage subsystem 2018 may also include a computer-readable storage media reader 2020, which may be further connected to computer-readable storage media 2023. Together with the system memory 2010, and optionally in conjunction with the system memory 2010, the computer-readable storage media 2022 may comprehensively represent remote, local, fixed, and / or removable storage devices plus storage media for temporarily and / or more permanently containing, storing, transmitting, and retrieving computer-readable information.
[0333] The computer-readable storage medium 2022 containing the code or portions of the code may also include any suitable media known or used in the art, including storage media and communication media, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage and / or transmission of information. This may include tangible, non-transitory computer-readable storage media such as RAM, ROM, electrically erasable programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or other tangible computer-readable media. When specified, this may also include non-tangible, transitory computer-readable media such as data signals, data transmissions, or any other medium that can be used to send the desired information and that can be accessed by the computing system 2000.
[0334] As examples, computer-readable storage media 2022 may include a hard drive that reads from or writes to non-removable nonvolatile magnetic media, a magnetic disk drive that reads from or writes to a removable nonvolatile magnetic disk, and a magnetic disk drive that reads from or writes to a removable nonvolatile optical disk, such as a CD ROM, a DVD, and a DVD. An optical drive that reads or writes to a removable non-volatile optical disk (disc or other optical media). Computer readable storage media 2022 may include, but is not limited to, Drives, flash memory cards, universal serial bus (USB) flash drives, secure digital (SD) cards, DVDs, digital video tapes, and the like. The computer-readable storage medium 2022 may also include solid-state drives (SSDs) based on non-volatile memory (such as SSDs based on flash memory, enterprise flash drives, solid-state ROMs, and the like), SSDs based on volatile memory (such as solid-state RAM, dynamic RAM, static RAM), DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM-based and flash memory-based SSDs. The disk drives and their associated computer-readable media may provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer system 2000.
[0335] The communication subsystem 2024 provides an interface to other computer systems and networks. The communication subsystem 2024 is used as an interface for receiving data from other systems and sending data from the computer system 2000 to other systems. For example, the communication subsystem 2024 can enable the computer system 2000 to be connected to one or more devices via the Internet. In some aspects, the communication subsystem 2024 may include a radio frequency (RF) transceiver component, a global positioning system (GPS) receiver component and / or other components for accessing wireless voice and / or data networks (for example, using cellular phone technology, advanced data network technologies such as 3G, 4G or EDGE (Enhanced Data Rates for Global Evolution), WiFi (IEEE 802.10 series standards), or other mobile communication technologies, or any combination thereof). In some aspects, in addition to or in place of a wireless interface, the communication subsystem 2024 can provide a wired network connection (for example, Ethernet).
[0336] In some aspects, the communication subsystem 2024 may also receive incoming communications in the form of structured and / or unstructured data feeds 2026 , event streams 2028 , event updates 2030 , and the like on behalf of one or more users who may use the computer system 2000 .
[0337] As an example, the communication subsystem 2024 may be configured to receive unstructured data feeds 2026 in real time from users of social media networks and / or other communication services, such as feed, Updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0338] Additionally, the communication subsystem 2024 may also be configured to receive data in the form of continuous data streams, which may include event streams 2028 and / or event updates 2030, which may be continuous or unbounded in nature, without a clear end to real-time events. Examples of applications that generate continuous data may include, for example, sensor data applications, financial quote machines, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, and the like.
[0339] The communication subsystem 2024 may also be configured to output structured and / or unstructured data feeds 2026 , event streams 2028 , event updates 2030 , etc. to one or more databases that may be in computer communication with one or more streaming data source computers coupled to the computer system 2000 .
[0340] Computer system 2000 may be of various types, including a handheld portable device (e.g., Cellular phones, computing tablets, PDAs), wearable devices (e.g., Google head-mounted display), PC, workstation, mainframe, kiosk, server rack, or any other data processing system.
[0341] Due to the ever-changing nature of computers and networks, the description of the computer system 2000 depicted in the figure is intended only as a specific example. Many other configurations with more or fewer components than the system depicted in the figure are possible. For example, customized hardware may also be used and / or specific elements may be implemented with hardware, firmware, software (including applets) or a combination thereof. In addition, connections to other computing devices such as network input / output devices may also be employed. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will recognize other ways and / or methods for implementing various aspects.
[0342] In the foregoing description, various aspects of the present invention have been described with reference to their specific aspects, but those skilled in the art will recognize that the present invention is not limited thereto. Each feature and aspect of the foregoing invention can be used individually or in combination. In addition, without departing from the broader spirit and scope of this description, aspect can be used in any number of environments and applications other than those described herein. Accordingly, this description and the accompanying drawings should be considered to be illustrative rather than restrictive.
Claims
1. A computer-implemented method comprising: establishing a mapping between a first entity in a first plurality of entities in a first discourse tree and a second entity in a second plurality of entities in a second discourse tree, the mapping establishing a relevance of the second entity to the first entity, wherein the discourse tree represents rhetorical relations between basic discourse units; In response to determining that a third entity in the first plurality of entities is not addressed by any entity in the second plurality of entities, generating a fictitious utterance tree by combining the additional utterance tree with the second utterance tree; determining a first communication discourse tree from the first discourse tree, wherein the communication discourse tree is a discourse tree having one or more verb signatures, each verb signature including a topic role indicating a role of a word in a corresponding basic discourse unit; determining a second communication discourse tree from the fictitious discourse tree; calculating a rhetorical consistency level between the first communication discourse tree and the second communication discourse tree by applying the prediction model to the first communication discourse tree and the second communication discourse tree; and In response to determining that the rhetorical consistency level is above a threshold, text corresponding to the fictitious discourse tree is output.
2. The method according to claim 1, wherein Creating a mapping involves: determining entity relevance scores for additional entities in the second plurality of entities; and In response to determining that the entity relevance score is greater than a threshold, the additional entity is selected as a second entity.
3. The method according to claim 1, wherein The fictitious discourse tree includes nodes representing rhetorical relations, and the method further includes integrating the fictitious discourse tree into a second discourse tree by connecting the nodes to a second entity.
4. The method according to claim 1, wherein Generating the fictitious discourse tree includes: calculating a relevance score for each of a plurality of additional utterance trees by applying the additional prediction model to the first utterance tree and the corresponding additional utterance tree, wherein the relevance score indicates a relevance of the first utterance tree to the corresponding additional utterance tree; and An additional utterance tree having a highest relevance score is selected from the plurality of additional utterance trees as the additional utterance tree.
5. The method according to claim 1, further comprising: Accessing a sentence comprising a plurality of basic discourse units, wherein at least one basic discourse unit comprises a verb and a plurality of words, each word comprising a role of the word within the basic discourse unit; and At least one of a first discourse tree, a second discourse tree, or the additional discourse tree is generated, wherein the generated tree represents a rhetorical relationship between the plurality of basic discourse units.
6. The method according to claim 1, further comprising: Constructing at least one of the first communication utterance tree or the second communication utterance tree by matching each segment having a verb with a verb signature, wherein matching each segment having a verb with a verb signature comprises: accessing a plurality of verb signatures, wherein each verb signature comprises the verb of the segment and a series of topic roles, wherein the topic roles describe a relationship between the verb and related words; For each of the plurality of verb signatures, determining a plurality of topic roles in the corresponding signature that match roles of words in the segment; selecting a particular verb signature from the plurality of verb signatures based on the particular verb signature including a highest number of matches; and The particular verb signature is associated with the segment.
7. The method according to claim 1, wherein The applying includes providing a first exchange utterance tree and a second exchange utterance tree to the prediction model, and receiving the rhetorical consistency level from the prediction module.
8. A system comprising: a non-transitory computer-readable medium storing computer-executable program instructions; and a processing device communicatively coupled to the non-transitory computer-readable medium for executing the computer-executable program instructions, wherein executing the computer-executable program instructions configures the processing device to perform operations comprising: establishing a mapping between a first entity in a first plurality of entities in a first discourse tree and a second entity in a second plurality of entities in a second discourse tree, the mapping establishing a relevance of the second entity to the first entity, wherein the discourse tree represents rhetorical relations between basic discourse units; In response to determining that a third entity in the first plurality of entities is not addressed by any entity in the second plurality of entities, generating a fictitious utterance tree by combining the additional utterance tree with the second utterance tree; determining a first communication discourse tree from the first discourse tree, wherein the communication discourse tree is a discourse tree having one or more verb signatures, each verb signature including a topic role indicating a role of a word in a corresponding basic discourse unit; determining a second communication discourse tree from the fictitious discourse tree; calculating a rhetorical consistency level between the first communication discourse tree and the second communication discourse tree by applying the prediction model to the first communication discourse tree and the second communication discourse tree; and In response to determining that the rhetorical consistency level is above a threshold, text corresponding to the fictitious discourse tree is output.
9. The system according to claim 8, wherein: Creating a mapping involves: determining entity relevance scores for additional entities in the second plurality of entities; and In response to determining that the entity relevance score is greater than a threshold, the additional entity is selected as a second entity.
10. The system according to claim 8, wherein: The fictitious discourse tree includes nodes representing rhetorical relations, and the system further includes integrating the fictitious discourse tree into a second discourse tree by connecting the nodes to a second entity.
11. The system according to claim 8, wherein Generating the fictitious discourse tree includes: calculating a relevance score for each of a plurality of additional utterance trees by applying the additional prediction model to the first utterance tree and the corresponding additional utterance tree, wherein the relevance score indicates a relevance of the first utterance tree to the corresponding additional utterance tree; and An additional utterance tree having a highest relevance score is selected from the plurality of additional utterance trees as the additional utterance tree.
12. The system of claim 8, wherein executing the computer-executable program instructions configures the processing device to perform operations comprising: accessing a sentence comprising a plurality of basic discourse units, wherein at least one basic discourse unit comprises a verb and a plurality of words, each word comprising a role of the word within the basic discourse unit; and At least one of a first discourse tree, a second discourse tree, or the additional discourse tree is generated, wherein the generated tree represents a rhetorical relationship between the plurality of basic discourse units.
13. The system of claim 8, wherein executing the computer-executable program instructions configures the processing device to perform operations comprising: Constructing at least one of the first communication utterance tree or the second communication utterance tree by matching each segment having a verb with a verb signature, wherein matching each segment having a verb with a verb signature comprises: accessing a plurality of verb signatures, wherein each verb signature comprises the verb of the segment and a series of topic roles, wherein the topic roles describe a relationship between the verb and related words; For each of the plurality of verb signatures, determining a plurality of topic roles in the corresponding signature that match roles of words in the segment; selecting the particular verb signature from the plurality of verb signatures based on the particular verb signature including a highest number of matches; and The particular verb signature is associated with the segment.
14. The system of claim 8, wherein executing the computer-executable program instructions configures the processing device to perform operations comprising: The first communication utterance tree and the second communication utterance tree are merged by identifying that a first root node of the first communication utterance tree and a second root node of the second communication utterance tree are the same.