Techniques for generating multimodal discourse trees

Multimodal discourse trees integrate text and numerical data through rhetorical and causal relationships, addressing the separation issue in current systems and enhancing the performance of automated agents and machine learning models.

JP2026086423APending Publication Date: 2026-05-26ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
ORACLE INT CORP
Filing Date
2026-01-14
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Current systems struggle to effectively combine text and numerical data in a meaningful way for tasks such as automated question answering and machine learning classification, as they often treat these data types separately, leading to inaccurate or incomplete analysis.

Method used

The generation and utilization of multimodal discourse trees (MMDTs) that integrate text and numerical data through discourse analysis, creating a hierarchical structure that includes rhetorical relationships and causal links, enabling enhanced classification and navigation of complex data sets.

Benefits of technology

MMDTs provide a more comprehensive understanding of interconnected data, improving the accuracy and capability of autonomous agents and machine learning models in tasks like question answering and classification by integrating textual and numerical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026086423000001_ABST
    Figure 2026086423000001_ABST
Patent Text Reader

Abstract

This invention provides a method, device, and storage medium for generating and utilizing a multimodal discourse tree (MMDT). [Solution] The method generates an Extended Discourse Tree (EDT) from a text corpus (e.g., from a Discourse Tree (DT) or Communication DT (CDT)), links data records (e.g., records containing numerical data) to the Extended Discourse Tree, and generates a Multimedia Discourse Tree (MMDT). The MMDT links any suitable text / records from heterogeneous sources. For example, entities identified from the basic discourse units of the EDT are matched with entities in the data records. The method also identifies causal links between EDTs and / or between data records, identifies rhetorical relationships for each entity match / causal link match, and merges the data records with the EDT to generate the MMDT.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims the interests of U.S. Provisional Application No. 63 / 166,568, filed on 26 March 2021, entitled “Multimodal Discourse Tree for Health Management and Security,” the contents of which U.S. Provisional Application No. 63 / 166,568 is incorporated herein by reference in its entirety for all purposes.

[0002] Technical field This disclosure relates to linguistics in general. More specifically, this disclosure relates to generating and using multi-modal discourse trees (MMDTs) (for example, to classify input text, generate answers to questions, navigate text, etc.). [Background technology]

[0003] background Significant increases in processor speed and memory capacity are leading to a growing number of computer-based linguistic applications. For example, computer-based analysis of linguistic discourse facilitates numerous applications, such as automated agents that can answer questions received from user devices, or machine learning applications that can classify input. In these and other contexts, current systems may provide text data and potentially related numerical data (e.g., logs, maps, data records) separately. In some cases, it may be beneficial to utilize both text and various potentially related numerical data to classify similar data and / or answer subsequent related questions. However, combining such data in a meaningful way is difficult. [Overview of the project] [Means for solving the problem]

[0004] Brief Overview Techniques for generating and utilizing multimodal discourse trees (MMDTs) are disclosed (for example, for classifying input text, generating answers to questions, navigating text, etc.).

[0005] In some embodiments, a method is disclosed that can be implemented by a computer to generate a multimodal discourse tree (for example, to classify input, navigate text, identify answers to subsequent questions, etc.). The method may include the step of obtaining a corpus of text and one or more data records separate from the corpus of text. The method may further include the step of generating an extended discourse tree for the corpus of text. In some embodiments, the extended discourse tree includes a plurality of discourse trees. Each discourse tree may include a plurality of nodes, where each terminal node of the discourse tree corresponds to a fragment of text, and each non-terminal node of the discourse tree indicates a rhetorical relationship between the nodes of the discourse tree. In some embodiments, the extended discourse tree includes further links between the plurality of discourse trees indicating yet another rhetorical relationship between the nodes of each discourse tree. The method may further include the step of identifying entity matches between a set of basic discourse units of the plurality of discourse trees and the one or more data records. In some embodiments, the entity matches include a first entity identified from the basic discourse units and an entity identified from the data records. It can be identified by comparing it with a second entity. The above method may further include the step of running a causal link identification algorithm for identifying one or more causal links between two data records from the one or more data records. The above method may further include the step of determining a corresponding rhetorical relationship for each identified entity match and each of the identified one or more causal links. The above method may further include the step of generating a node for each identified entity match and each identified causal link for the extended discourse tree. The above method may further include the step of creating a multimodal discourse tree by linking the respective nodes generated for each entity match and each causal link to the respective nodes of the extended discourse tree, at least in part based on the determined corresponding rhetorical relationships.

[0006] In some embodiments, the step of generating the extended discourse includes i) generating a first discourse tree from a first text of the text corpus, the first discourse tree corresponding to a first part of the first text; the step of generating the extended discourse further includes ii) generating a second discourse tree from a second text of the text corpus, the second discourse tree corresponding to a second part of the second text; and the step of generating the extended discourse further includes iii) linking the first and second discourse trees using the specific rhetorical relationships in response to determining specific rhetorical relationships between the respective basic discourse units of the first and second discourse trees.

[0007] In some embodiments, the first discourse tree and the second discourse tree are communication discourse trees that include the respective verb signatures generated for each basic discourse unit of the first discourse tree and the second discourse tree.

[0008] In some embodiments, the step of identifying the entity further includes comparing the second entity identified from the data record with a third entity identified from the second data record, wherein the entity refers to one of (i) a person, (ii) a company, (iii) a place, (iv) the name of a document, or (v) a date or time.

[0009] In some embodiments, the step of identifying the entity match further includes the step of identifying the entity from a predefined ontology.

[0010] In some embodiments, the method may further include a step of classifying subsequent inputs based at least in part on the multimodal discourse trees, the step of classifying subsequent inputs comprising: i) generating a training dataset comprising a plurality of multimodal discourse trees, each multimodal discourse tree corresponding to a text corpus and a set of data records, and each multimodal discourse tree associated with a label corresponding to a classification; the step of classifying subsequent inputs further comprises: ii) training a machine learning model to classify inputs based at least in part on the training dataset and a supervised learning algorithm; iii) generating a corresponding multimodal discourse tree from the subsequent inputs, each subsequent input comprising a set of texts and a set of data records; and the step of classifying subsequent inputs further comprises: iv) classifying the subsequent inputs based at least in part on providing the corresponding multimodal discourses generated from the subsequent inputs as input to the machine learning model and receiving an output from the machine learning model indicating a classification of the subsequent inputs.

[0011] In some embodiments, the method further includes the step of navigating the corpus of text using the multimodal discourse tree, the step of navigating the corpus of text including the step of accessing the multimodal discourse tree and the step of determining from the multimodal discourse tree a first basic discourse unit to respond to a query from a user device, the first basic discourse unit corresponding to a first node of the first discourse tree of the multimodal discourse tree, and the step of navigating the corpus of text further includes the step of determining from the multimodal discourse tree a set of navigation options, the set of navigation options being (i) a first rhetorical relationship between the first node of the first discourse tree and a second node of the first discourse tree and (ii) a second rhetorical relationship between the first node and a third node of the second discourse tree of the multimodal discourse tree, or (iii) the first The steps of navigating the corpus of text include at least two of a third rhetorical relationship between the first node of the discourse tree and the fourth node of the multimodal discourse tree associated with the corresponding data record, and further include presenting at least two of the first, second, or third rhetorical relationships to the user device, and in response to receiving further user input including a selection of the first, second, or third rhetorical relationship, i) presenting a second basic discourse unit corresponding to the second node, at least based on a decision on the selection corresponding to the first rhetorical relationship; ii) presenting a third basic discourse unit corresponding to the third node, at least based on a decision on the selection corresponding to the first rhetorical relationship; or iii) presenting at least a portion of the corresponding data record, at least based on a decision on the selection corresponding to the third rhetorical relationship.

[0012] The exemplary methods discussed herein can be implemented on a system and / or device comprising one or more processors and / or can be stored as instructions on a non-temporary computer-readable medium. Various aspects of this disclosure can be implemented using a computer program product that includes a computer program / instruction that, when executed by the processor, causes the processor to perform one of the methods disclosed herein. [Brief explanation of the drawing]

[0013] [Figure 1] This figure shows an example of a computing environment for classifying input using a multimodal discourse tree (MMDT), according to at least one embodiment of the present disclosure. [Figure 2] This figure shows an exemplary flow of a method performed by the computing device of Figure 1 (for example, by the multimodal discourse tree (MMDT) generation module of Figure 1) according to at least one embodiment. [Figure 3] This figure shows an exemplary discourse tree according to at least one embodiment. [Figure 4] This figure shows a flow chart of a method for generating a communication discourse tree according to at least one embodiment. [Figure 5] This figure shows an example of an extended discourse tree according to at least one embodiment. [Figure 6] This is a flowchart of an example process for creating an extended discourse tree according to at least one embodiment. [Figure 7] This figure shows the relationships between text units of a document at various levels of granularity, according to at least one embodiment. [Figure 8] This figure shows an example of a multimodal discourse tree according to at least one embodiment. [Figure 9]Another diagram showing the multi-modal conversation tree of FIG. 8, along with the relationships between text units of a document at various levels of granularity and the relationships between those text units and data records associated therewith, according to at least one embodiment. [Figure 10] A diagram showing the flow of a method for generating a multi-modal conversation tree, according to at least one embodiment. [Figure 11] A schematic diagram of a distributed system for implementing one of the aspects, according to at least one embodiment. [Figure 12] A simplified block diagram of one or more components of a system environment in which services provided by one or more components of a system according to at least one embodiment according to one aspect of the present disclosure can be provided as cloud services. [Figure 13] A diagram showing an exemplary computing subsystem in which various aspects can be implemented, according to at least one embodiment.

Mode for Carrying Out the Invention

[0014] Detailed Description Aspects of the present disclosure relate to generating and utilizing a multi-modal conversation tree (MMDT) to classify an input. A multi-modal conversation tree (MMDT) refers to a conversation tree that includes associated data (values, sets, and / or records) related to parts of the conversation tree (e.g., basic conversation units). The MMDT includes relationships between a data set and the conversation tree and relationships between data within a given data set. Examples of rhetorical relationships (also referred to herein as "rhetorical relatedness") that can exist between data values, sets, and records include reason, cause, enablement, contrast, and temporal sequence.

[0015] Discourse analysis plays a crucial role in constructing the logical structure of ideas expressed in text. Discourse trees are a means of hierarchically formalizing textual discourse, specifying rhetorical relationships between phrases and sentences (discourse units). A discourse tree (DT) is a compromise between a complete logical representation, like a logical form, and the informal representation of the original text. Learning DTs has led to numerous applications in content generation, summarization, machine translation, and question answering. A limitation of using DTs in common data analysis tasks is that they are designed to represent textual discourse rather than causal relationships between components of abstract data items.

[0016] The techniques disclosed herein utilize discourse analysis to generate various discourse trees from a corpus of text. For example, a discourse tree may be generated for each text in a corpus of documents (e.g., a corpus of related / related content). In some embodiments, a discourse tree (DT) may be extended with communicative behaviors (e.g., verb signatures) to generate a communicative discourse tree (CDT). Rhetorical relationships between units of the CDT (e.g., basic discourse units) may be identified to generate an extended discourse tree (EDT). An EDT may include combinations of CDTs that have identified relationships between different units of CDTs and across different levels of granularity (e.g., between sentences in a document, between paragraphs in a document, between sentences / paragraphs in different documents, and / or between entire documents).

[0017] A multimodal discourse tree (MMDT) can be generated from an EDT by combining it with associated numerical data through discourse abstraction, thereby formalizing discourse across various types of numerical data, as well as text and text documents in a corpus. An MMDT can be generated such that the same rhetorical relationships that apply between text fragments also apply between data values, sets, and records. An MMDT can be a discourse tree containing associated data (values, sets, and / or numerical data records) related to parts of an EDT (e.g., basic discourse units). The MMDT can also contain relationships between associated data and It can be generated to include the relationships between the portion of the accompanying data and the portion of the corpus text. The resulting MMDT can then be used to drive automated responses in the field of autonomous agents and / or to classify inputs. In one example, a discourse tree may be used to represent the reasoning structure necessary to reach a conclusion in a legal case. An extended discourse tree can be generated from this discourse tree, and the extended discourse tree can be extended with related information such as telephone conversations, information about people's movements, transaction records / financial records, and internal relationships between pieces of information (e.g., data from a specific case).

[0018] The disclosed technology improves existing solutions by better understanding the entirety of multiple relevant sources. MMDT can be used to enhance the capabilities of interactive chatbots (e.g., autonomous agent applications). Traditional discourse trees represent the flow of an author's thoughts at the paragraph or multi-paragraph level. Traditional discourse trees become quite inaccurate when applied to larger text fragments or documents. These discourse trees can be used to generate augmented discourse trees, which can then generate multimodal discourse trees that can function as representations of interconnected documents and corresponding data records covering a given topic.

[0019] Specific definition The term "rhetorical structure theory" is used in this specification. In such cases, it is a field of research and study that has provided a theoretical foundation that can enable the analysis of discourse consistency.

[0020] When used in this specification, “discourse tree” or “DT” refers to a structure that expresses rhetorical relationships between sentences that form part of a sentence.

[0021] When used in this specification, “rhetorical relationship,” “rhetorical relevance,” “consistency relationship,” or “discourse relationship” describes how two segments of discourse are logically connected to one another. Examples of rhetorical relationships include detail, contrast, and attribute.

[0022] Where used in this specification, a “sentence fragment” or “fragment” is a part of a sentence that can be separated from the rest of the sentence. A fragment is a basic discourse unit. For example, in the sentence “Dutch accident investigators say that evidence points to pro-Russian rebels as being responsible for shooting down the plane,” the two fragments are “Dutch accident investigators say that evidence points to pro-Russian rebels” and “as being responsible for shooting down the plane.” "To" can contain a verb, but it does not necessarily have to contain a verb.

[0023] Where used in this specification, “signature” or “frame” refers to the characteristics of a verb in a fragment. Each signature may contain one or more subject roles. For example, in the fragment “Dutch accident investigators say that evidence points to pro-Russian rebels,” the verb is “say,” and the verb “say” is used in this way. When used in a specific way, the signature can be "agent verb topic". In this case, "investigators" is the agent and "evidence" is the topic.

[0024] Where used in this specification, “subject role” refers to a signature component used to describe the role of one or more words. Following the example above, “agent” and “topic” are subject roles.

[0025] When used in this specification, "nuclearity" refers to which text segment, fragment, or span is closer to the center of the writer's intent. A nucleus is a more central span, while a satellite is less central.

[0026] When used in this specification, "coherency" refers to two rhetorical relationships. It refers to things that link together.

[0027] When used in this specification, "communication discourse tree" or "CDT" refers to a discourse tree supplemented with communicative behavior. Communicative behavior is a collaborative action performed by individuals based on mutual deliberation and discussion. Therefore, a communication discourse tree combines rhetorical information with communicative behavior.

[0028] As used in this specification, a "communicative verb" is a verb that indicates communication. For example, the verb "deny" is a communicative verb.

[0029] When used in this specification, "communication behavior" refers to an action performed by one or more agents, and the subject matter of the agents.

[0030] When used in this specification, "entity" refers to an independent and distinct entity. Examples include objects, places, and people. Furthermore, an entity may be a subject or topic such as "electric vehicle," "brake," or "France."

[0031] Moving on to the drawings, Figure 1 shows an example of a computing environment for correcting raw text generated using deep learning techniques, according to at least one embodiment. In the example shown in Figure 1, the computing environment 100 includes one or more of a computing device 102 and a user device 104. The computing device 102 can implement an application (e.g., application 106). In some embodiments, application 106 engages in conversation with the user device 104 and uses one or more of the techniques disclosed herein. Using the above, an autonomous agent (e.g., a chatbot) can interact in response to input provided by user device 104. As another example, application 106 may be part of the machine learning component of computing device 102. Application 106 may be configured to generate any number of suitable multimodal discourse trees (MMDTs) from text data (e.g., a corpus of text data obtained from text data 116) and data records (e.g., numerical data obtained from record data store 118). Examples of computing device 102 are the distributed system 1000 and client computing devices 1002, 1004, 1006, and 1008 in Figure 10.

[0032] The user device 104 may be any mobile device such as a mobile phone, smartphone, tablet, laptop, or smartwatch. As shown, the user device 104 has a user interface 108. The user interface 108 may be configured to accept input from the user (e.g., via a keyboard, microphone input, mouse input, touchscreen, etc.) and provide data corresponding to that input to the computing device 102. In some embodiments, the user interface 108, or another component of the user device 104, may be configured to capture voice input and convert that voice input to text before sending it to the computing device 102. The user device 104 then sends the text to the computing device It may be configured to transmit to 102. In other embodiments, the text may be retrieved using a user interface (not shown) provided by the computing device 102 and / or from a data store (not shown) accessible to the computing device 102. Suitable text examples include electronic text sources such as text files, Portable Document Format (PDF)® documents, and rich text documents. In some cases, preprocessing may be performed on the input text to remove unwanted characters or formatting fields. The input text may be organized using one or more structural or organizational approaches such as sections, paragraphs, and pages.

[0033] In some embodiments, the user device 104 and the computing device 102 may be connected communicatively via a network 110. The network 110 may be any suitable public or private network, including the Internet, a local area network, a virtual private network, etc.

[0034] In some embodiments, the computing device 102 may include a classifier 112. The classifier 112 may be any suitable machine learning model trained using training data 114 to provide outputs (e.g., answers, generated text, etc.) in response to inputs (e.g., questions or other data posed in a user interface 108). The training data 114 may include any suitable data on which the classifier 112 can be trained (e.g., good / bad answer question / answer pairs, examples of good and / or bad text generation, etc.). In some cases, entities in the text may be matched using an ontology 117. The ontology 117 may be a domain-specific ontology (e.g., finance, law, business, medicine, science, etc.). The ontology 117 may include, among several features, formal specifications of various entities and their relationships. The ontology 117 may be used to identify synonymous words and / or phrases. In some embodiments, the ontology 117 can be predefined, and / or application 106 can construct at least a portion of the ontology 117 from an external source.

[0035] The MMDT generation module 120 may be configured to generate any suitable number of MMDTs from the text data 116 and the record data store 118. In some embodiments, these MMDTs can serve as training data for classifying future inputs. As an example, text in police records (e.g., an example of text data) may be merged with numerical-based evidence (e.g., call records, financial transaction data, location data obtained from a mobile phone, etc.) to generate MMDTs related to data associated with a particular criminal case. These MMDTs may be associated with a label (e.g., extortion) and used as labeled training data for a classifier 112 (e.g., an example of a machine learning model). The classifier 112 may be trained using various examples of such training data (e.g., some examples showing text / data records showing extortion, some examples showing other text / data records showing extortion, etc.) along with any suitable supervised learning algorithm to configure the classifier 112 to classify subsequent inputs (e.g., text and data records received from a user device 104 and associated with the case) as either indicating an extortion crime or not indicating an extortion crime. In some embodiments, the generated MMDT can be used to answer questions relating to the event in which the MMDT was generated. The operation for generating and using the MMDT will be discussed in more detail with reference to Figures 2 to 9.

[0036] Figure 2 shows an exemplary flow of Method 200 performed by the computing device 102 of Figure 1 (for example, by the MMDT generation module 120) according to at least one embodiment.

[0037] Method 200 may begin with 201, in which a corpus of documents may be obtained. In some embodiments, the corpus of documents may relate to a specific topic. For example, the corpus of documents may include any suitable text from police records relating to a particular crime. The documents may have any suitable length and / or format.

[0038] In 202, discourse trees can be generated for each text in the corpus, at least partially based on rhetorical structure theory. The techniques for generating these discourse trees are discussed below in relation to Figure 3.

[0039] Figure 3 shows a text encoding 300 of a discourse tree (referred to as “discourse tree 300”) generated from an instance of text, according to at least one embodiment. As just one example, discourse tree 300 may be generated from the following text: “I avoid flu shots because I am allergic to eggs. Most flu shots produced today use an egg-based manufacturing process that leaves trace amounts of egg protein behind. I do not need a flu shot since I got it last year. I believe it is not necessary for me, because the vaccine is not 100% effective. Also, I never get the flu, so I do not need a vaccine. (I have an egg allergy, so I don't need the flu shot.) Most influenza vaccines currently in production use an egg-based manufacturing process, which leaves trace amounts of egg protein. I don't need a flu shot because I got one last year. I don't think I need one because vaccines aren't 100% effective. Also, I've never had the flu, so I don't need a vaccine.) Discourse tree 300 can be generated from the above input text, at least partially based on rhetorical structure theory.

[0040] Rhetorical Structure Theory and Discourse Trees Linguistics is the scientific study of language. For example, linguistics can include sentence structure (syntactic), such as subject-verb-object; sentence meaning (semantics), such as the difference between "dog bites man" and "man bites dog"; and even what speakers do during a conversation, i.e., discourse analysis or analysis of language beyond the scope of a sentence.

[0041] The theoretical foundation of discourse (Rhetorical Structure Theory (RST)) is based on "Rhetorical structure theory: A Theory of Mann, William and Thompson, Sandra." Text-Interdisciplinary Journal for the Study of Discourse This may be due to (8(3):243-281: 1988). Just as the syntax and semantics of programming language theory have helped enable modern software compilers, RST has helped enable discourse analysis. More specifically, RST assumes structural blocks at at least two levels. The two levels include a first level such as nucleality and rhetorical relations, and a second level of structure or schema. Discourse parsers or other computer software can parse (analyze) text into discourse trees. These parsers / software can be configured to identify various rhetorical relations between EDUs in the discourse tree.

[0042] Rhetorical structure theory models the logical structure of a text (the structure used by the writer) based on the relationships between parts of the text. RST (Rhetorical Structure Theory) uses a discourse tree to analyze the hierarchy of the text. Text coherence is simulated by forming a layered, interconnected structure. Rhetorical relationships are divided into equal and subordinate classes. These relationships are maintained across two or more text spans, thus achieving coherence. These text spans are called basic discourse units (EDUs). Clauses within a sentence and sentences within a text are logically connected by the author. The meaning of a given sentence relates to the meanings of the preceding and following sentences. This logical relationship between clauses is called the text's coherence structure. RST is one of several theories of discourse based on a tree-like discourse structure, a discourse tree (DT). The leaves of the DT correspond to EDUs (sequential atomic text spans). Adjacent EDUs are connected by coherence relationships (e.g., attributes, sequences) that form higher-level discourse units. These units are also subordinate to this relational link. EDUs linked by relationships are further distinguished based on their relative importance. The nucleus is the core part of the relationship, and the satellites are the peripheral parts. As mentioned above, both topic and rhetorical agreement are analyzed to determine accurate request-response pairs. When a speaker answers a question, such as a phrase or sentence, the speaker's response must address the topic of the question. If the question is implicitly formed, the topic can be addressed through the message's seed text. The expected response is not only to maintain the status quo but also to be consistent with a generalized understanding of this seed.

[0043] Rhetorical relationships As described above, the embodiments described herein use a discourse tree (and / or a communicative discourse tree) that includes rhetorical relationships. Rhetorical relationships can be described in various ways. Some rhetorical relationships are listed below. However, this list is not intended to be exclusive.

[0044] [Table 1]

[0045] Several empirical studies have shown that the vast majority of texts are constructed using nuclear-satellite relationships. While this is based on the premise that [certain conditions are met], other relationships do not involve a finite selection of the kernel. Examples of such relationships are shown below.

[0046] [Table 2]

[0047] Returning to Figure 3, the discourse tree 300 may be generated based on the parsing of the input text (e.g., using a discourse parser and / or software). The discourse tree 300 may contain any preferred combination of the rhetorical relations (also referred to as "rhetorical relationships") described above. For example, the discourse tree 300 may contain rhetorical relationships 302-322. The discourse tree 300 may further contain several text units (e.g., basic discourse units (EDUs) 324-346). Each rhetorical relation represents a rhetorical relationship between two or more EDUs. As an example, rhetorical relation 314 may represent a rhetorical relationship (e.g., "explanation") between EDU 334 and EDU 336.

[0048] Any preferred method for constructing the discourse tree 300 can be used. One exemplary method for constructing the discourse tree 300 may include the following actions:

[0049] (1) Divide the discourse text into multiple units according to (a) and (b) below. (a) The unit size may vary depending on the purpose of the analysis.

[0050] (b) Typically the unit is a section. (2) Examine each unit and each adjacent unit. Is a relationship maintained between them (for example, identified at least partially based on a predetermined set of rules)? (3) If a relationship is maintained, mark that relationship.

[0051] (4) If a relationship is not maintained, that unit may be at the boundary of a higher-level relationship. Focus on the relationships maintained between larger units (spans).

[0052] (5) Continue until all units in the text are understood. In some embodiments, the MMDT generation module 120 may generate a discourse tree for each text in the corpus. Discourse trees may be generated at varying degrees of granularity. For example, a DT may be generated for each sentence, for each paragraph, for each document, etc. Each discourse tree may contain nodes, where each non-terminal node (or edge) represents a rhetorical relationship between two fragments (e.g., sentence fragments), and each terminal node of the discourse tree is associated with one of the fragments.

[0053] Returning to Figure 2, once a DT has been generated for the corpus text (e.g., police reports and / or text information related to crime cases) (according to any preferred granularity), method 200 may proceed to 203, in which a Communicative Discourse Tree (CDT) may be generated for each text in the corpus. Figure 4 illustrates in more detail the technique for generating CDTs from text.

[0054] Figure 4 shows a flow diagram of a method 400 for generating a communication discourse tree according to at least one embodiment. In some embodiments, using a communication discourse tree to generate queries enables improvements in search engine results.

[0055] In block 401, process 400 includes accessing a sentence containing a fragment (for example, a sentence from the corpus obtained in 201 of Figure 2). At least one fragment contains a verb and one or more words, each word containing the role of a word within the fragment, and each fragment is a basic discourse unit.

[0056] In some embodiments, the MMDT generation module 120 in Figure 1 may identify that a sentence contains several fragments. Each fragment may contain a verb, but may not contain a verb.

[0057] In block 402, process 400 includes generating a discourse tree that represents the rhetorical relationships between sentence fragments. For example, the MMDT generation module 120 may generate one or more discourse trees, as discussed in association with Figure 3. Each discourse tree may contain nodes, each non-terminal node representing a rhetorical relationship between two sentence fragments, and each terminal node of the discourse tree is associated with one of the sentence fragments.

[0058] In block 403, process 400 includes accessing multiple verb signatures. For example, the MMDT generation module 120 (e.g., VerbNet, predefined) A list of verbs can be accessed (from a list of verbs that have been created, for example). Each verb can match or relate to a verb in a fragment. For example, if the fragment contains the verb "deny", the MMDT generation module 120 can access a list of verb signatures related to the verb "deny".

[0059] Each verb signature may contain the verb of the fragment and one or more topic roles. For example, a signature may contain one or more of the following: a noun phrase (NP), a noun (N), a communicative action (V), a verb phrase (VP), or an adverb (ADV). The topic role describes the relationship between the verb and the associated words. For example, "the teacher amused the children" It has a different signature than "small children amuse quickly". In the case of the first fragment, the verb "deny", application 102 accesses a list of frames or verb signatures for verbs that match "deny". The list is "NP V NP to These are "be NP", "NP V that S", and "NP V NP".

[0060] Each verb signature contains a subject role. The subject role refers to the role of the verb in the sentence fragment. In some embodiments, the MMDT generation module 120 determines the subject role in each verb signature. An exemplary subject role is "actor". (Actor), "agent", "asset", "attribute" (gender), "beneficiary", "cause", "location destination source", "destination", "source", "location", "experiencer", "extent", "instrument", "material and product", "material", "product" Product, patient, predicate, recipient, stimulus, theme, time, or topic Includes (k).

[0061] In block 404, process 400 performs the following for each verb signature: This includes determining several subject roles of each signature that match the role of the word in the fragment. For example, if the fragment of the raw text sentence contains the verb "deny", the MMDT generation module 120 will determine that the verb "deny" is "agent", It can be determined that it has only three roles: "verb" and "theme".

[0062] In block 405, process 400 includes selecting a particular verb signature from among the verb signatures based on the verb signature with the highest number of matches. For example, the fragment "the rebels deny... that they control the territory" may be matched with the verb signature "deny" "NP V NP", and "control" may be matched with control (rebel, territory). Verb signatures are nested, and as a result This yields a nested signature, "deny (rebel, control (rebel, territory))". Each selected verb signature is associated with its corresponding fragment, completing the communicative discourse tree.

[0063] Returning to 204 in Figure 2, once a communication discourse tree (CDT) is generated for the text corpus in the manner described above in relation to Figures 3 and 4, the MMDT generation module 120 can construct an extended discourse tree (EDT) for the text corpus. Figures 5 and 6 illustrate in more detail the techniques for generating an EDT for the text corpus.

[0064] Figure 5 shows an example of an extended discourse tree according to at least one embodiment. Figure 5 shows an extended discourse tree 500. As shown, the extended discourse tree 500 includes groups 500, 520, 530, 540, and 550. Each group includes a document (e.g., from a text corpus) and a discourse tree generated from that document. For example, group 510 includes a discourse tree 511 and a document 512, group 520 includes a discourse tree 521 and a document 522, and so on. The discourse tree shown in Figure 5 may be an example of a DT generated as described above in Figure 2, 202, and / or the discourse tree in Figure 5 may be an example of a CDT generated as described above in Figure 2, 203.

[0065] For example, in addition to links within specific discourse trees such as discourse trees 511, 521, 531, 541, and 551, the extended discourse tree 500 includes inter-discourse tree links 561-564 and associated inter-document links 571-574. As further illustrated with reference to Figure 6, the MMDT generation module 120 can construct discourse trees 511-515, where discourse tree 511 represents document 512, discourse tree 521 represents document 522, and so on. The extended discourse tree 500 can be constructed by constructing a discourse tree (e.g., DT and / or CDT) for each paragraph or document (e.g., using process 400 for each tree).

[0066] Inter-discourse tree link 561 connects discourse trees 511 and 521, inter-discourse tree link 562 connects discourse trees 521 and 531, inter-discourse tree link 563 connects discourse trees 511 and 541, and inter-discourse tree link 564 connects discourse trees 521 and 551. Based on inter-discourse tree links 561 to 564, the MMDT generation module 120 creates inter-document links 571, 572, 573, and 574, which correspond to inter-discourse tree links 561, 562, 563, and 564, respectively. Documents 512, 522, 532, 542, and 552 can be navigated using inter-document links 571 to 574.

[0067] The MMDT generation module 120 determines one or more entities in the first discourse tree from among the discourse trees 511-515. Examples of entities include location, etc. This could be a person or a company. The MMDT generation module 120 then identifies identical entities present in other discourse trees. Based on the identified entities, the MMDT generation module 120 determines the rhetorical relationships between each matching entity. These decisions utilize the same discourse rules used to generate the discourse tree.

[0068] For example, the entity "San Francisco" appears in document 512, such as "San Francisco is in California." "San Francisco has a moderate climate but can be quite windy." While the climate in San Francisco is mild, it can be quite windy, as document 522 further explains, MMDT generation module 120 will use the entity "San Francisco" to generate data between If a rhetorical relationship is determined to be one of the "details," links 561 and 571 can be marked as "details." In some embodiments, the MMDT generation module 120 uses a discourse parser to syntactically parse a combination of EDUs (e.g., "San Francisco is in California" from document 512 and "San Francisco has a moderate climate but can be quite windy" from document 522) such that one is the other It can be identified that one is a detail of the other. In some embodiments, this relationship may be unidirectional, or multiple rhetorical relationships may be used to represent a bidirectional relationship (for example, to indicate that each of the two EDUs is a detail of the other). Following the examples above, the MMDT generation module 120 determines links 562-564 and the corresponding links 572-574 based on the determined rhetorical relationships. The MMDT generation module 120 combines the discourse trees of the document paragraphs to form an extended discourse tree 500.

[0069] By using links in the extended discourse tree 500, the MMDT generation module 120 can navigate between paragraphs within the same document or between documents such as documents 512 and 522. For example, if a user is interested in more information on a particular topic, the MMDT generation module 120 can navigate via a detailed rhetorical relationship from core to satellite within a paragraph or via a detailed rhetorical relationship hyperlink to a document that provides more specific information on that topic.

[0070] Conversely, if the user determines that a suggested topic is not necessarily required, they can return to a higher level of view of the document (e.g., from satellite to core, or from narrow document to broad document). The MMDT generation module 120 can then navigate detailing relationships in reverse order, i.e., from satellite to core within paragraphs or across documents. Similarly, the MMDT generation module 120 can facilitate other navigation options, such as relying on contrasting or conditional rhetorical relationships to explore controversial topics.

[0071] To construct rhetorical links between text fragments in different paragraphs or documents, the MMDT generation module 120 may use hypothetical text fragments or temporary paragraphs to identify relationships between entities from each text fragment of the original paragraph and perform coreference analysis and discourse syntactic analysis on the paragraphs. In some embodiments, the MMDT generation module 120 may utilize an ontology (e.g., ontology 117 in Figure 1) to identify relationships between entities based on the relationships provided in the ontology.

[0072] Figure 6 is a flowchart of an example of a process 600 for creating an extended discourse tree (e.g., the extended discourse tree 500 in Figure 5) according to at least one embodiment. The input to process 600 is a set of documents, and the output is an extended discourse tree encoded as a regular discourse tree with document identification labels for each node. For illustrative purposes, process 600 is described with respect to two documents, for example, documents 110a and 110b (e.g., the example corpus text obtained in 201 of Figure 2), but process 600 can be used with any number of documents.

[0073] In block 601, process 600 includes accessing the first and second documents. Examples of documents include texts, books, news articles, and other electronic text documents. In the examples provided in the field of law enforcement, the document may be any suitable text document corresponding to the alleged criminal case. The document may include police reports, statements by witnesses / victims / perpetrators, etc.

[0074] In one embodiment, the MMDT generation module 120 performs document analysis, including the generation of a document tree representing the sentence and phrase structures of the document. Rhetorical relationships associated with inter-document links can determine various navigation scenarios. By default, detail can be used. The MMDT generation module 120 can provide links to other documents associated by attribute relationships if the user is interested in questions such as "why" or "how". If the user objects to the originally presented document, or requests a document that provides a counterpoint to the current document, the MMDT generation module 120 can provide links to documents associated by contrast relationships.

[0075] In yet another embodiment, the MMDT generation module 120 obtains first and second documents. In block 602, process 600 includes creating a first discourse tree for the first paragraph of the first document. The MMDT generation module 120 accesses the paragraph from the first document. Each sentence in this paragraph contains a fragment or basic discourse unit. At least one fragment contains a verb. Each word in a fragment contains a role of the words within the fragment, such as function. The MMDT generation module 120 generates a discourse tree representing the rhetorical relationships between fragments according to the technique described above in relation to Figure 3. This discourse tree contains multiple nodes, each non-terminal node representing a rhetorical relationship between two fragments, and each terminal node associated with one of the fragments. The MMDT generation module 120 continues in this manner to build a set of discourse trees for each paragraph in the first document. Process 600 describes paragraphs as units of text, but other sizes of text (e.g., sentences, pages, chapters, etc.) can also be used.

[0076] In block 603, process 600 includes creating a second discourse tree for the second paragraph of the second document. In block 603, process 600 performs substantially the same steps for the second document as those performed for the first document in block 602. If process 600 creates extended discourse trees for three or more documents, process 600 performs the functions described in block 602 for multiple documents. Process 600 can iterate over all pairs of discourse trees in the set of discourse trees corresponding to a document. A pair of discourse trees can be represented as follows:

[0077] DT i and DT j ∈DTA In block 604, process 600 includes determining an entity and a corresponding first basic conversation unit from a first conversation tree. Various methods can be used, such as keyword processing (searching for one of a list of predefined keywords in the sentences of the first document), using a trained machine learning model, or searching Internet resources. The MMDT generation module 120 processes the conversation tree DT i and DT j to identify all noun phrases and named entities in it.

[0078] In one example, the MMDT generation module 120 extracts noun phrases from the conversation tree. Then, the MMDT generation module 120 uses a trained machine learning model to classify the noun phrases as (i) an entity or (ii) not an entity.

[0079] In block 605, process 600 includes determining a second basic conversation unit that matches the first basic conversation unit in a second conversation tree. More specifically, the MMDT generation module 120 calculates the overlap to identify a common entity E i between DT j and DT. The MMDT generation module 120 establishes the relationships between the occurrences of the entities in E i,j , such as being equal, being a sub-entity, or being a part. Then, the MMDT generation module 120 forms a paragraph-level rhetorical link R(E i,j ) for each entity pair occurrence in E i,j . i,j ) for each entity pair occurrence in E

[0080] In block 606, process 600 includes creating an extended conversation tree by linking the first conversation tree and the second conversation tree via the rhetorical relationship in response to determining the rhetorical relationship between the first basic conversation unit and the second basic conversation unit. More specifically, the MMDT generation module 120, for example, EDU(E i) and EDU(E j The rhetorical relationships for each rhetorical link are classified by merging text fragments such as ) to construct their DT and using the recognized relational labels for each rhetorical link.

[0081] In one embodiment, the MMDT generation module 120 combines a first basic discourse unit and a second basic discourse unit to form a temporary paragraph. The discourse navigation application 102 then applies discourse syntactic analysis to the temporary paragraph to determine the rhetorical relationships between the first basic discourse unit and the second basic discourse unit within the temporary paragraph.

[0082] In yet another embodiment, in response to not determining rhetorical relationships, the MMDT generation module 120 creates default rhetorical relationships of type details between a first basic discourse unit and a second basic discourse unit to link the first discourse tree and the second discourse tree.

[0083] In one embodiment, the MMDT generation module 120 performs automatic construction and classification of links between text spans across a document. Here, the following approaches can be used: lexical distance, lexical chaining, information extraction, and linguistic template matching. Lexical distance can use cosine similarity across pairs of sentences, and lexical chaining can be more robust by leveraging synonyms and superordinate concepts.

[0084] Extended discourse trees can form relationships between two or more documents at various levels of granularity. For example, they can determine relationships between basic discourse units, as described in relation to process 600. Furthermore, extended discourse trees can represent relationships between words, sentences, paragraphs, sections of documents, or entire documents. As shown, each individual graph represents a smaller level of each individual document. It consists of subgraphs. Links are shown that represent the logical connections between topics within a single document.

[0085] Figure 7 shows the relationships between text units of documents at various levels of granularity in one embodiment. Figure 7 shows discourse trees 701, 702, and 703, each corresponding to a separate document. Figure 7 also shows various inter-document links, such as word links 710 linking words in documents 702 and 703, paragraph / sentence links 711 linking paragraphs or sentences in documents 701 and 702, phrase links 712 linking phrases in documents 701 and 703, and inter-document links 713 linking documents 701 and 703. The MMDT generation module 120 can navigate between documents 701 and 703 using links 710 to 713.

[0086] An augmented discourse tree, such as the augmented discourse tree created by process 700, can be used to navigate other body texts of a document or text. Augmented discourse trees enable a variety of applications, such as autonomous agents, improved search and navigation, and question-answer tuning. In some embodiments, an EDT, such as the EDT created by process 700, can be used, at least in part, as training data to train a classifier (e.g., classifier 112 in Figure 1) to classify inputs (e.g., any preferred inputs including text and / or numerical-based data).

[0087] Returning to Figure 2, at 205, an accompanying data record may be obtained. In this example where police records are used, the accompanying data record may include any suitable data record. As just one example, the data record in this example may include call logs, location data, financial transactions (e.g., bank statements), web page visits, images, etc. In general, data records may be obtained from the same source or from different sources. These sources may be different from the source of the text data.

[0088] In 206, each data record may have at least a portion of a record that has been converted to a unified canonical format. For example, each data source may be converted to a unified canonical format having normalized named entities such as time, date, place, person's name, telephone number, and account number (if available). In some embodiments, the specific format and / or content of the converted data may depend on the context and a predefined set of rules. In some embodiments, this information may initially be provided in different formats or by different identifiers, but a unified identifier (e.g., name) can be identified using the set of rules and / or ontology 117, and the data of each record may be associated with an identifier (e.g., an originally included identifier and / or a unified identifier).

[0089] In 207, for each basic discourse unit (EDU) of the EDT (e.g., generated from a text corpus), several candidate phrases that may be associated with an accompanying data record can be identified. In some embodiments, identifying a particular candidate phrase (e.g., a particular EDU) may depend on a domain and a predefined protocol for identifying such candidate phrases.

[0090] In 208, for each candidate phrase, the MMDT generation module 120 may identify the entity of the candidate phrase. In some embodiments, the specific entity identified may depend on the context / domain and a predefined set of rules. In some embodiments, this information is initially presented in different forms or by different identifiers. Candidate phrases may be provided, but a unified identifier (e.g., a name) can be identified using a set of rules and / or ontology 117, and candidate phrases may be associated with identifiers (e.g., originally included identifiers and / or unified identifiers).

[0091] In 209, the MMDT generation module 120 may identify matching entities between data records and candidate phrases and / or between data records. For example, the MMDT generation module 120 may iterate over each entity in each data record and compare those entities with entities associated with other data records. Associations between two data records that have matching entities may be maintained over any number of suitable matches found. In some embodiments, a list of matching data records may be maintained. As another example, the MMDT generation module 120 may iterate over each candidate phrase and compare the entities in the candidate phrase with entities associated with each data record. If a match is found between a candidate phrase and a data record, an association indicating this match may be maintained. In some embodiments, a list of matching EDU / data record pairs may be maintained.

[0092] In 210, the MMDT generation module 120 may perform actions to determine causal links between data records and / or between EDUs (candidate phrases) and data records. Typically, multiple data records are interconnected and represent correlated events (some of which cause others to occur). An algorithm for identifying causal links (for example, when an event referenced in one data record causes an event referenced in another data record and / or between an EDU and a data record). For example, when person A calls person B and person A sends money to person B, a set of rules can be used to identify that the former event causes the latter event. In some embodiments, these rules can be based on the premise that if there is no match between values ​​in a data record, i.e., they share a value, then the earlier event causes the later event. Thus, a data record can be identified as the cause of an event described in another data record or text, and vice versa.

[0093] An algorithm for identifying causal links (e.g., a causal link identification algorithm) is provided. In some embodiments, two operators R(.) (reason) and C(.) (conclusion), as well as other negations, can be utilized. Two negation operators are required: ¬ (¬x indicates that x is false) for negating propositional expressions and - for negating R(.) and C(.). An argument is an expression of the form R(y):(-)C(x). An argument is the reason for concluding a claim. It has two main parts: the premise (reason) and the conclusion. Functions R and C play the roles of providing the reason and concluding, respectively. An argument can be interpreted as follows: its conclusion is true because it arises as a natural consequence from the premise according to a given concept. This concept refers to the nature of the link between them (e.g., the premise implicitly indicates the conclusion), and is formally identified by a colon in the definition. However, the conclusion may be true but the function is not true, and vice versa.

[0094] For example, R(y):C(x) corresponds to the output showing that "y is a reason to conclude x," and R(y):-C(x) corresponds to the output showing that "y is a reason not to conclude x." Processing nested arguments is crucial for finding the answers to break them down, because it is insufficient to process only the object-level layers of the argument or only the meta-level layers separately. Nested arguments are central to the processing of text and dialogue and must provide support for nested arguments and negations. Table 2 shows the various forms made possible by a predefined set of definitions. This document provides proof and negation of the statement (x, y, z, t are propositional expressions used to simplify the matter). The table is not mutually exclusive.

[0095] [Table 3]

[0096] The illustrative arguments in Table 2 relate to the functionality of credit cards. By default, credit cards function (are usable) especially when the account balance is positive. However, there are exceptions. For whatever reason, banks may refuse a transaction. These examples demonstrate that internal and external reasons R, as well as charges C, can potentially be identified using argument mining techniques. Furthermore, internal reasons and charges can be identified by argument mining techniques through recursion. Therefore, nested structures appear to be a more suitable target language for argumentation when the argumentation occurs in natural language dialogue and text.

[0097] Table 3 contains templates that the MMDT generation module 120 can use to extract logical atoms from EDU (e.g., candidate phrases), convert rhetorical relations into RC operators, and form logical representations of arguments. To do so, semantic representations can be constructed for representations of objects related to banking transactions ch(g). These semantic representations can be associated with EDU. Using the determined structure of the discourse tree, RC representations in L can be formed, and these RC representations undergo argument analysis in downstream components.

[0098] [Table 4]

[0099] [Table 5]

[0100] [Table 6]

[0101] The MMDT generation module 120 can generate two causal chains according to the rules above and / or below (for example, one from the EDU / fragment of a discourse tree generated from a text corpus, and the other from the text of a data record, the respective EDUs of two discourse trees, the respective EDUs generated from the text of two different data records, etc.) to determine whether the result of the first chain is caused or implicitly indicated by the latter or formally. A set of arguments and their negations may be provided as a set of formulas, some of which are identified below. The result operator is the smallest closure of a set of inference rules extended by one meta-rule. These predefined rules can be used to determine whether a causal link (e.g., the rhetorical relation "cause") is identified between an EDU and a data record and / or between data records. Some of these rules are listed below.

[0102] Metarules are arbitrary inference rules

[0103]

number

[0104] This indicates that it can be reversed. The inference rule inversion process occurs every time a negation occurs before the leftmost "R", and therefore, in the general case, the inference rule is 1 and i,j ∈ {0,1}.

[0105]

number

[0106] The reasons are interchangeable. The following rules may demonstrate mutual support.

[0107]

number

[0108] Another rule is to gather different reasons for the same conclusion in a single argument.

[0109]

number

[0110] Cautious monotonicity is the reason for the argument that it justifies any It means that it can be extended by the premises. "Cut" represents a kind of minimality of the reasons in the argument. .

[0111]

number

[0112] The following two rules describe nesting of R(.) and C(.): exportation shows how meta-arguments are simplified, and permutation shows that sequences of reasons are possible for some forms of meta-arguments.

[0113]

number

[0114] Recursion, monotonicity, and cut hold when the smallest inference relation follows the above rules, which means that, using the result relation, the operation of argumentation by the inference rules becomes well-founded. Let Δ be a set of arguments (or their negations). Let α and β be arguments.

[0115] If α ∈ Δ, then Δα (recursive). If Δβ ​​then Δ∪{α}β (monotonicity) If Δ∪{α}β and Δα, then Δβ (cut). In 212, the MMDT generation module 120 may perform any preferred actions to verify the matches and / or causal links identified in 209 and / or 210. This includes linking data records to each other and linking data records to EDTs (e.g., EDU / candidate phrases for the EDT). This may include iterating over the links. The MMDT generation module 120 may generate an MMDT by starting from the EDT and generating nodes and edges for each data record involved in the confirmed entity matches and / or causal links. For example, if an entity match is identified between two data records or between a data record and an EDU generated from a text corpus, and a node does not yet exist in the EDT for one (or both) data records, a new node may be generated (e.g., for each data record), and depending on the specific use case, the new node for one data record may be linked to the EDU generated from the corpus or the new node generated for the other data record. For example, if the EDU generated from the corpus contains entities that match entities found in the data record (e.g., they both contain terms found in ontology 117 in Figure 1, they both contain matching nouns such as "San Francisco", etc.), and the data record's node If a node does not yet exist, the MMDT generation module 120 may generate a new node for the EDT, and the node corresponding to the EDU may be linked to this new node using an edge associated with a rhetorical relationship (e.g., "detail"). In some embodiments, the MMDT generation module 120 may link entity matches to each other via one rhetorical relationship (e.g., "detail"), while causal links may be linked in the MMDT via different rhetorical relationships (e.g., "cause").

[0116] Figure 8 shows an example of a multimodal discourse tree 800 according to at least one embodiment. Figure 8 shows the extended discourse tree 500 of Figure 5, along with further edges corresponding to data records 802, 804, and 806.

[0117] For example, in addition to links between specific discourse trees such as discourse trees 511, 521, 531, 541, and 551, the extended discourse tree 500 includes inter-discourse tree links 561-564 and associated inter-document links 571-574. As further illustrated with reference to Figure 6, the MMDT generation module 120 can construct discourse trees 511-515. Discourse tree 511 represents document 512, discourse tree 521 represents document 522, and so on. The extended discourse tree 500 can be constructed by constructing a discourse tree (e.g., DT and / or CDT) for each paragraph or document. Discourse trees 511-515 can be constructed from corpus text.

[0118] Inter-discourse tree link 561 connects discourse trees 511 and 521, inter-discourse tree link 562 connects discourse trees 521 and 531, inter-discourse tree link 563 connects discourse trees 511 and 541, and inter-discourse tree link 564 connects discourse trees 521 and 551. Based on inter-discourse tree links 561 to 564, the MMDT generation module 120 creates inter-document links 571, 572, 573, and 574, which correspond to inter-discourse tree links 561, 562, 563, and 564, respectively. Documents 512, 522, 532, 542, and 552 can be navigated using inter-document links 571 to 574.

[0119] The MMDT generation module 120 determines one or more entities in the first discourse tree from among the discourse trees 511 to 515. Examples of entities include places, things, people, or companies. The MMDT generation module 120 then identifies identical entities present in other discourse trees. Based on the determined entities, the MMDT generation module 120 determines the rhetorical relationships between each matching entity.

[0120] As described above, the MMDT generation module 120 is one of the discourse trees 511-515. The MMDT generation module 120 may be configured to identify one or more entities in the EDU and data records of the first discourse tree. For example, the MMDT generation module 120 may identify entities referenced in the discourse tree 510 and data record 802. Based on the identified entities, the MMDT generation module 120 may generate links 808 (e.g., edges).

[0121] In some embodiments, the MMDT generation module 120 may be configured to identify one or more causal links within the EDU and data records of a first discourse tree among discourse trees 511-515, and / or between data records. For example, using the algorithm described above in relation to Figure 2, the MMDT generation module 120 may identify causal links between data records 802 and 804 and between discourse tree 541 (or the EDU of discourse tree 541) and data record 806. Based on this identification information, the MMDT generation module 120 may generate causal links 810 and 812, respectively.

[0122] By using links in the extended discourse tree 500, the MMDT generation module 120 can navigate between paragraphs in the same document or between documents in a text corpus, and between data records and discourse trees. The MMDT 800 does not necessarily have to include rhetorical relationships corresponding to links 808-812, at least initially.

[0123] Returning to Figure 2, at 214, the MMDT generation module 120 may perform actions to identify rhetorical relationships between phase / data record matches / links and data records / data record matches / links identified in 212. To identify rhetorical relationships between text fragments, the MMDT generation module 120 may use hypothetical text fragments or temporary paragraphs from each text fragment of the original paragraph to identify relationships between entities and perform coherence analysis and discourse parsing on the paragraph. In some embodiments, the MMDT generation module 120 may utilize an ontology (e.g., ontology 117 in Figure 1) to identify relationships between entities based on the relationships provided in the ontology. The identified rhetorical relationships may be associated with links 808-812 in Figure 8.

[0124] In 216, the MMDT generation module 120 may perform an operation to identify data records (e.g., data records 808-812) as corresponding to the core or satellite of the EDT. Any preferred predefined set of rules may be used to identify data records as corresponding to the core or satellite of the EDT (e.g., the candidate phrase / EDU to which the EDT is associated).

[0125] In 218, the generated MMDT may be converted to a normalized MMDT. This conversion may involve generating an EDU for each data record and attaching the generated EDU to the EDT according to the decision made in 216. A simple example is given below in which data records previously linked to a pair of text EDUs and associated with detailed rhetorical relationships may be added as EDUs. Rhetorical relationships (causes) may be inserted to reinforce the core according to a predefined set of rules. This process may be carried out for each data record until each data record is included as an EDU in the MMDT just completed.

[0126] [Table 7]

[0127] Figure 9 shows the relationships between text units of a document at various levels of granularity and the relationships between those text units and associated data records in an MMDT according to at least one embodiment. Figure 9 shows an MMDT900 which includes discourse trees 901, 902, and 903, each corresponding to a separate document (for example, examples of DTs or CDTs generated using the techniques described above in Figures 3 and 4, respectively). Figure 9 also shows various inter-document links, such as word links 910 which link words in documents 902 and 903, paragraph / sentence links 911 which link paragraphs or sentences in documents 901 and 902, phrase links 912 which link phrases in documents 901 and 903, and inter-document links 913 which link documents 901 and 903. The MMDT900 also includes data records 905-909. MMDT900 may include any preferred number of causal links (e.g., causal links 910-912) and / or entity links (e.g., entity links 913-915) identified in the embodiments described above in relation to Figure 2. In some embodiments, each of the data records 905-909 may be included as an EDU within the discourse trees 901-903 that link the data records 905-909. The MMDT generation module 120 can navigate between documents 901-903 and / or data records 905-909 using any preferred links.

[0128] Figure 10 shows a flow of method 1000 for generating a multimodal discourse tree according to at least one embodiment. In some embodiments, method 1000 may be performed by the computing device 102 in Figure 1 (e.g., application 106 in Figure 1, MMDT generation module 120 in Figure 1). The operations of method 1000 may be performed in any preferred order. In some embodiments, method 1000 may include more operations than those shown in Figure 10, or fewer operations than those shown in Figure 10.

[0129] Method 1000 may begin with 1001, in which a text corpus and one or more data records separate from the text corpus may be obtained (for example, from text data 116 and record data store 118 in Figure 1, respectively). One or more data records may be obtained from any preferred source. In some embodiments, the source of one or more data records may be different from the source of the text corpus.

[0130] In 1002, an extended discourse tree (e.g., EDT500 in Figures 5 and 8) may be generated for the text corpus. In some embodiments, the extended discourse tree includes multiple discourse trees (e.g., each being an example of DT300 in Figure 3 or CDT400 in Figure 4). In some embodiments, each discourse tree includes multiple nodes, and each of the discourse trees Terminal nodes correspond to text fragments (and / or data records), and each non-terminal node in the discourse tree indicates a rhetorical relationship between nodes in the discourse tree. In some embodiments, an extended discourse tree includes further links between multiple discourse trees that indicate yet another rhetorical relationship between nodes in each discourse tree.

[0131] In 1003, entity matches can be identified between a set of basic discourse units of multiple discourse trees and one or more data records (and / or between basic discourse units). In some embodiments, entity matches are identified by comparing a first entity identified from a basic discourse unit with a second entity identified from a data record.

[0132] In 1004, one or more causal links may be identified. For example, a causal link identification algorithm (as described, for example, in association with Figure 2) may be performed to identify one or more causal links between two data records out of one or more data records.

[0133] In 1005, for each identified entity match and each of the identified one or more causal links, a corresponding rhetorical relation. For example, a “detail” rhetorical relation may be identified for each entity match, and a “cause” rhetorical relation may be identified for each causal link.

[0134] In 1006, for each identified entity match and each identified causal link, a corresponding node may be generated for the extended discourse tree.

[0135] In 1007, a multimodal discourse tree (e.g., MMDT800 and 900 in Figures 8 and 9) can be created by linking each node generated for each entity match and each causal link to each node in the extended discourse tree, at least partially based on the determined corresponding rhetorical relationships.

[0136] In some embodiments, the step of generating the extended discourse includes i) generating a first discourse tree from a first text of the text corpus, the first discourse tree corresponding to a first part of the first text; the step of generating the extended discourse further includes ii) generating a second discourse tree from a second text of the text corpus, the second discourse tree corresponding to a second part of the second text; and the step of generating the extended discourse further includes iii) linking the first and second discourse trees using the specific rhetorical relationships in response to determining specific rhetorical relationships between the respective basic discourse units of the first and second discourse trees.

[0137] In some embodiments, the first discourse tree and the second discourse tree are communication discourse trees that include the respective verb signatures generated for each basic discourse unit of the first discourse tree and the second discourse tree.

[0138] In some embodiments, the step of identifying the entity further includes comparing the second entity identified from the data record with a third entity identified from the second data record, wherein the entity refers to one of (i) a person, (ii) a company, (iii) a place, (iv) the name of a document, or (v) a date or time.

[0139] In some embodiments, the step of identifying the entity match is predetermined The process further includes the step of identifying the above entities from the ontology.

[0140] In some embodiments, Method 1000 may further include the step of classifying subsequent inputs based at least in part on the multimodal discourse trees, the step of classifying subsequent inputs comprising: i) generating a training dataset comprising a plurality of multimodal discourse trees, each multimodal discourse tree corresponding to a text corpus and a set of data records, and each multimodal discourse tree associated with a label corresponding to a classification; the step of classifying subsequent inputs further comprises: ii) training a machine learning model to classify inputs based at least in part on the training dataset and a supervised learning algorithm; iii) generating a corresponding multimodal discourse tree from the subsequent inputs, each subsequent input comprising a set of texts and a set of data records; and the step of classifying subsequent inputs further comprises: iv) classifying the subsequent inputs based at least in part on providing the corresponding multimodal discourses generated from the subsequent inputs as input to the machine learning model and receiving an output from the machine learning model indicating the classification of the subsequent inputs.

[0141] In some embodiments, Method 1000 further includes the step of navigating the corpus of text using the multimodal discourse tree, the step of navigating the corpus of text including the step of accessing the multimodal discourse tree and the step of determining from the multimodal discourse tree a first basic discourse unit to respond to a query from a user device, the first basic discourse unit corresponding to a first node of the first discourse tree of the multimodal discourse tree, and the step of navigating the corpus of text further includes the step of determining from the multimodal discourse tree a set of navigation options, the set of navigation options being (i) a first rhetorical relationship between the first node of the first discourse tree and a second node of the first discourse tree and (ii) a second rhetorical relationship between the first node and a third node of the second discourse tree of the multimodal discourse tree, or (iii) the first The step of navigating the corpus of text includes at least two of a third rhetorical relationship between the first node of a discourse tree and the fourth node of the multimodal discourse tree associated with the corresponding data record, and further includes presenting at least two of the first, second, or third rhetorical relationships to the user device, and in response to receiving further user input including a selection of the first, second, or third rhetorical relationship, i) presenting a second basic discourse unit corresponding to the second node, at least based on at least a decision on the selection corresponding to the first rhetorical relationship; ii) presenting a third basic discourse unit corresponding to the third node, at least based on at least a decision on the selection corresponding to the first rhetorical relationship; or iii) presenting at least a portion of the corresponding data record, at least based on the decision on the selection corresponding to the third rhetorical relationship.

[0142] Exemplary computing system Figure 11 is a simplified diagram of a distributed system 1100 for realizing one of the embodiments described above. In the embodiment shown, the distributed system 1100 includes one or more client computing devices 1102, 1104, 1106, and 1108 configured to run and operate client applications such as web browsers and proprietary clients (e.g., Oracle Forms) via one or more networks 1110. A server 1112 may be coupled to communicate with the remote client computing devices 1102, 1104, 1106, and 1108 via the network 1110.

[0143] In various embodiments, the server 1112 may be adapted to run one or more services or software applications provided by one or more components of the system. These services or software applications may include non-virtual and virtual environments. The virtual environment may be a two-dimensional or three-dimensional (3D) representation, a page-based logical environment, or any other virtual event, tray, etc. This may include services used for show, simulators, classrooms, purchasing and trading of goods, and business activities. These services may be offered as web-based services, cloud services, or Software as a Service (S) services. Under the aS) model, it can be supplied to users of client computing devices 1102, 1104, 1106, and / or 1108. Users operating client computing devices 1102, 1104, 1106, and / or 1108 can use one or more client applications to interact with server 1112 and utilize the services provided by these components.

[0144] In the configuration shown in the figure, the software components 1118, 1120, and 1122 of the distributed system 1100 are shown to be implemented on server 1112. In other embodiments, one or more of the components of the distributed system 1100 and / or the services provided by these components may be implemented by one or more of the client computing devices 1102, 1104, 1106, and / or 1108. In this case, a user operating the client computing device may use one or more client applications to access the services provided by these components. These components may be implemented in hardware, firmware, software, or a combination thereof. It should be understood that various different system configurations that differ from the distributed system 1100 are possible. Therefore, the embodiments shown in the figure are examples of a distributed system for implementing one embodiment of the system and are not intended to be limiting.

[0145] Client computing devices 1102, 1104, 1106 and / or 1108 are handheld portable devices (e.g., iPhone®, mobile phones, iPad®, calculating tablets, personal digital assistants (PDAs)). The client computing device may be a Digital Assistant or a wearable device (e.g., a Google Glass® head-mounted display) that runs software such as Microsoft Windows® Mobile® and / or various mobile operating systems such as iOS®, Windows Phone, Android®, BlackBerry 10, PalmOS, and the Internet, email, Short Message Service (SMS), BlackBerry®, or other available communication protocols. The client computing device may be a general-purpose personal computer, which includes, for example, personal computers and / or laptop computers that run various versions of the Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems. The client computing device may be a workstation computer, which runs one of various commercially available UNIX® or UNIX-like operating systems, including, but not limited to, various GNU / Linux operating systems such as Google Chrome OS. Alternatively, or additionally, the client computing devices 1102, 1104, 1106, and 1108 may be thin client computers, internet-enabled gaming systems (e.g., Microsoft Xbox game consoles with or without Kinect® gesture input devices), and / or other electronic devices such as personal messaging devices that can communicate via network 1110.

[0146] The exemplary distributed system 1100 is shown to have four client computing devices, but any number of client computing devices may be supported. Other devices, such as devices with sensors, may interact with the server 1112.

[0147] Network 1110 in the distributed system 1100 may be any type of network familiar to those skilled in the art, capable of supporting data communication using any of a variety of commercially available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Packet Switching), AppleTalk, etc. For example, network 1110 may be a local area network (LAN), such as one based on Ethernet® or Token Ring. Network 1110 may also be a wide area network or the Internet. Network 1110 may include a virtual network, such as a virtual private network (VPN), intranet, extranet, or public switched telephone network. (PSTN: Public Switched Telephone Network), infrared network, wireless network Work (for example, the Institute of Electrical and Electronics (IEE)) E) Networks operating under any of the 802.9 protocols, Bluetooth®, and / or other radio protocols, and / or any combination thereof, and / or other networks, including but not limited to these.

[0148] Server 1112 may consist of one or more general-purpose computers, dedicated server computers (including, for example, PC (personal computer) servers, UNIX® servers, midrange servers, mainframe computers, rack-mount servers, etc.), server farms, server clusters, or other appropriate configurations and / or combinations. Server 1112 may include one or more virtual machines running a virtual operating system, or other computing architectures with virtualization. One or more flexible pools of logical storage can be virtualized to maintain virtual storage for the server. A virtual network can be controlled by Server 1112 using software-defined networking. In various embodiments, Server 1112 may be adapted to run one or more services or software applications described in the above disclosure. For example, Server 1112 may correspond to a server for performing the processing described above in the embodiments of this disclosure.

[0149] Server 1112 may run an operating system including any of the above and any commercially available server operating system. Server 1112 may also run any of a variety of additional server applications and / or middle-tier applications, including a Hypertext Transport Protocol (HTTP) server, a File Transfer Protocol (FTP) server, a Common Gateway Interface (CGI) server, a Java® server, and a database server. Examples of database servers include those from Oracle, Microsoft, and Sybase. Commercially available from companies such as Sybase and IBM (International Business Machines). This includes, but is not limited to, these items.

[0150] In some implementations, server 1112 may include one or more applications for analyzing and integrating data feeds and / or event updates received from users of client computing devices 1102, 1104, 1106, and 1108. For example, the data feeds and / or event updates may be received from one or more third-party sources and continuous data streams, such as Twitter®. The data feed and / or real-time events may include, but are not limited to, feeds, Facebook® updates, or real-time updates, as well as real-time events related to sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring. The server 1112 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client computing devices 1102, 1104, 1106, and 1108.

[0151] Furthermore, the distributed system 1100 may include one or more databases 1114 and 1116. Databases 1114 and 1116 may reside in various locations. For example, one or more of databases 1114 and 1116 may reside on a non-temporary storage medium local to (and / or located on) server 1112. Alternatively, databases 1114 and 1116 may be far away from server 1112 and communicate with server 1112 via a network-based connection or a dedicated connection. In one embodiment, databases 1114 and 1116 may reside on a Storage-Area Network (SAN). Similarly, any necessary files for performing functions originating from server 1112 may be stored locally on server 1112 and / or remotely as appropriate. In one embodiment, databases 1114 and 1116 may include relational databases, such as those provided by Oracle, adapted to store, update, and retrieve data in response to SQL format commands.

[0152] Figure 12 is a simplified block diagram of one or more components of a system environment 1200 that can supply services provided by one or more components of one embodiment of the system according to one embodiment of the present disclosure as cloud services. In the embodiment shown, the system environment 1200 includes one or more client computing devices 1204, 1206, and 1208 that can be used by a user to interact with a cloud infrastructure system 1202 that provides cloud services. The client computing devices may be configured to run client applications such as a web browser, a proprietary client application (e.g., Oracle Forms), or other applications that can be used by a user of the client computing device to interact with the cloud infrastructure system 1202 to use services provided by the cloud infrastructure system 1202.

[0153] It should be understood that the cloud infrastructure system 1202 shown in the figure may have components other than those shown. Furthermore, the embodiments shown in the figure are merely examples of cloud infrastructure systems that may incorporate embodiments of the present invention. In some embodiments, the cloud infrastructure system 1202 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different configurations or arrangements of components.

[0154] Client computing devices 1204, 1206, and 1208 may be similar to those described above for client computing devices 1102, 1104, 1106, and 1108.

[0155] Although the exemplary system environment 1200 is shown to have three client computing devices, any number of client computing devices may be supported. Other devices, such as devices with sensors, may interact with the cloud infrastructure system 1202.

[0156] Network 1210 can facilitate data communication and exchange between client computing devices 1204, 1206, and 1208 and the cloud infrastructure system 1202. Each network may be any type of network familiar to those skilled in the art, capable of supporting data communication using any of the various commercially available protocols, including those described above for network 1210.

[0157] The cloud infrastructure system 1202 may include one or more computers and / or servers that may include the above-described components for server 1212.

[0158] In certain embodiments, the services provided by the cloud infrastructure system may include a number of services available on demand to users of the cloud infrastructure system, such as online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database processing, and managed technical support services. The services provided by the cloud infrastructure system are dynamically scalable to meet the needs of its users. A specific instance of a service provided by the cloud infrastructure system is referred to herein as a “service instance.” Generally, any service available to users from a cloud service provider’s system via a communication network such as the Internet is referred to as a “cloud service.” Typically, in a public cloud environment, the servers and systems that make up the cloud service provider’s system are different from the customer’s own on-premises servers and systems. For example, the cloud service provider’s system may host an application, and users may order and use such application on demand via a communication network such as the Internet.

[0159] In some examples, services in a computer network cloud infrastructure may include storage, hosted databases, hosted web servers, secure computer network access to software applications, or other services provided to users by the cloud vendor or otherwise known in the art. For example, services may include password-protected access to remote storage on the cloud over the internet. Another example may include a web service-based hosted relational database and scripting language middleware engine for private use by networked developers. Yet another example may include access to an email software application hosted on the cloud vendor's website.

[0160] In certain embodiments, the cloud infrastructure system 1202 may include a set of applications, middleware, and database service offerings delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. One example is the Oracle Public Cloud provided by the assignee.

[0161] Massive amounts of data, sometimes referred to as big data, can be hosted and / or manipulated by infrastructure systems at numerous levels and different scales. The datasets such data can contain are so large and complex that they can be difficult to process using typical database management tools or conventional data processing applications. For example, terabytes of data may be difficult to store, retrieve, and process using personal computers or their rack-based counterparts. Data of this size may be difficult to run using modern relational database management systems and desktop statistics and visualization packages. They may require massively parallel processing software running thousands of server computers, going beyond the structure of commonly used software tools, to capture, curate, manage, and process the data within an acceptable timeframe.

[0162] To visualize massive amounts of data, detect trends, and / or interact with the data, analysts and researchers can store and process extremely large datasets. Dozens, hundreds, or thousands of processors linked in parallel can act on such data, enabling them to display such data, simulate external forces on the data, or simulate what it represents. These datasets may require structured data, such as data organized in databases or data following structured models, and / or unstructured data (e.g., emails, images, data blobs (binary large objects), web pages, complex event processing). By enhancing the ability of the architecture to concentrate more (or fewer) computing resources on a target relatively quickly, cloud infrastructure systems become better available for performing tasks on massive datasets based on requests from businesses, government agencies, research organizations, private individuals, groups of like-minded individuals or organizations, or other entities.

[0163] In various embodiments, the cloud infrastructure system 1202 may be adapted to automatically provision, manage, and track customer subscriptions to services supplied by the cloud infrastructure system 1202. The cloud infrastructure system 1202 may provide cloud services through various deployment models. For example, the cloud infrastructure system 1202 may be owned by an organization that sells cloud services (e.g., owned by Oracle) and services may be provided under a public cloud model where the services are available to the general public or various industrial enterprises. As another example, the cloud infrastructure system 1202 may be operated for a single organization only and services may be provided under a private cloud model where services are available to one or more entities within that organization. The cloud services may also be provided under a community cloud model where the cloud infrastructure system 1202 and the services provided by the cloud infrastructure system 1202 are shared by several organizations within a relevant community. The cloud services may also be provided under a hybrid cloud model, which is a combination of two or more different models.

[0164] In some embodiments, the services provided by the cloud infrastructure system 1202 include services in the Software as a Service (SaaS) category, Platform as a Service (PaaS) category, Infrastructure as a Service (IaaS) category, or hybrid services. This may include one or more services offered under the category. A customer may order one or more services provided by the cloud infrastructure system 1202 by subscription order. The cloud infrastructure system 1202 then performs processing to provide the services in the customer's subscription order.

[0165] In some embodiments, the services provided by the cloud infrastructure system 1202 may include, but are not limited to, application services, platform services, and infrastructure services. In some examples, application services may be provided by the cloud infrastructure system via a SaaS platform. The SaaS platform may be configured to provide cloud services that fall under the SaaS category. For example, the SaaS platform may provide the functionality to build and deliver a set of on-demand applications on an integrated development and deployment platform. The SaaS platform may manage and control the underlying software and infrastructure for providing SaaS services. By using the services provided by the SaaS platform, customers can utilize applications that run on the cloud infrastructure system. Customers can obtain application services without having to purchase separate licenses and support. A variety of different SaaS services may be provided. Examples include, but are not limited to, services that provide solutions for sales performance management, enterprise integration, and business flexibility for large organizations.

[0166] In some embodiments, platform services may be provided by a cloud infrastructure system via a PaaS platform. The PaaS platform may be configured to provide cloud services categorized as PaaS. Examples of platform services include, but are not limited to, services that enable organizations (such as Oracle) to integrate existing applications on a shared common architecture, and the ability to build new applications that leverage the shared services provided by the platform. The PaaS platform may manage and control the underlying software and infrastructure for providing PaaS services. Customers can obtain PaaS services provided by the cloud infrastructure system without having to purchase separate licenses and support. Examples of platform services include, but are not limited to, Oracle Java Cloud Service (JCS) and Oracle Database Cloud Service (DBCS).

[0167] By utilizing the services provided by the PaaS platform, customers can leverage programming languages ​​and tools supported by the cloud infrastructure system and control the deployed services. In some embodiments, the platform services provided by the cloud infrastructure system may include database cloud services, middleware cloud services (e.g., Oracle Fusion middleware services), and Java cloud services. In one embodiment, the database cloud service may support a shared service deployment model that enables an organization to pool database resources and supply database-as-a-service to customers in the form of a database cloud. The middleware cloud service may provide customers with a platform for developing and deploying various business applications on the cloud infrastructure system, while the Java cloud service provides support for the cloud infrastructure. We can provide customers with a platform for deploying Java applications within a structured system.

[0168] Various different infrastructure services may be provided by an IaaS platform in a cloud infrastructure system. Infrastructure services facilitate the management and control of basic computing resources such as storage and networking, as well as other fundamental computing resources for customers using services provided by SaaS and PaaS platforms.

[0169] In a particular embodiment, the cloud infrastructure system 1202 may include infrastructure resources 1230 for providing resources used to provide various services to customers of the cloud infrastructure system. In one embodiment, the infrastructure resources 1230 may include a pre-integrated optimal combination of hardware such as servers, storage, and networking resources for running services provided by the PaaS platform and SaaS platform.

[0170] In some embodiments, resources in the cloud infrastructure system 1202 may be shared by multiple users and dynamically reallocated according to demand. Resources may also be allocated to users in different time zones. For example, the cloud infrastructure system 1202 can maximize resource utilization by enabling a first group of users in a first time zone to utilize the resources of the cloud infrastructure system for a specified period of time, and by enabling the reallocation of the same resources to other groups of users located in different time zones.

[0171] In certain embodiments, several internal shared services 1232 may be provided, shared by various components or modules of the cloud infrastructure system 1202 and the services provided by the cloud infrastructure system 1202. These internal shared services may include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelisting services, high availability, backup and recovery services, services to enable cloud support, email services, notification services, and file transfer services.

[0172] In certain embodiments, the cloud infrastructure system 1202 may provide comprehensive management of cloud services (e.g., SaaS services, PaaS services, and IaaS services) within the cloud infrastructure system. In one embodiment, the cloud management function may include functions for provisioning, managing, and tracking customer subscriptions received by the cloud infrastructure system 1202.

[0173] In one embodiment, as shown in the figure, cloud management functionality may be provided by one or more modules, such as an order management module 1220, an order orchestration module 1222, an order provisioning module 1211, an order management and monitoring module 1210, and an identity management module 1228. These modules may include, or be provided by, one or more computers and / or servers, which may be general-purpose computers, dedicated server computers, server farms, server clusters, or other appropriate configurations and / or combinations.

[0174] In exemplary operation 1234, a customer using a client device such as client computing device 1204, 1206, or 1208 may interact with the cloud infrastructure system 1202 by requesting one or more services provided by the cloud infrastructure system 1202 and by placing an order for a subscription to one or more services provided by the cloud infrastructure system 1202. In certain embodiments, the customer may access a cloud user interface (UI), i.e., cloud UI 1212, cloud UI 1214, and / or cloud UI 1216, and place a subscription order through these UIs. The order information received by the cloud infrastructure system 1202 in response to the customer's order may include information identifying the customer and one or more services provided by the cloud infrastructure system 1202 that the customer intends to subscribe to.

[0175] After an order is placed by the customer, the order information is received via the cloud UI 1210, 1214 and / or 1216.

[0176] In operation 1236, the order is stored in the order database 1218. The order database 1218 may be one of several databases that are operated by the cloud infrastructure system 1202 and operate in conjunction with other system elements.

[0177] In operation 1238, order information is transferred to the order management module 1220. In some examples, the order management module 1220 may be configured to perform order-related invoicing and accounting functions, such as order confirmation and reservation of orders at the time of confirmation.

[0178] In operation 1240, information about the order is communicated to the order orchestration module 1222. The order orchestration module 1222 may use the order information to orchestrate the provisioning of services and resources for the order placed by the customer. In some examples, the order orchestration module 1222 may orchestrate the provisioning of resources to support services subscribed to using the services of the order provisioning module 1211.

[0179] In certain embodiments, the order orchestration module 1222 enables the management of business processes associated with each order and applies business logic to determine whether the order should proceed to provisioning. In operation 1242, upon receiving an order for a new subscription, the order orchestration module 1222 sends a request to the order provisioning module 1211 to allocate resources and configure those resources required to fulfill the subscription order. The order provisioning module 1211 enables the allocation of resources for the service ordered by the customer. The order provisioning module 1211 provides a level of abstraction between the cloud services provided by the cloud infrastructure system 1202 and the physical implementation layer used to provision resources to provide the requested service. Thus, the order orchestration module 1222 can be decoupled from implementation details such as whether services and resources are actually provisioned in execution or pre-provisioned and only allocated / assigned when requested.

[0180] In operation 1244, once the services and resources are provisioned, a notification of the provided services may be sent by the order provisioning module 1211 of the cloud infrastructure system 1202 to the customers on client computing devices 1204, 1206 and / or 1208.

[0181] In operation 1246, customer subscription orders may be managed and tracked by the order management and monitoring module 1210. In some examples, the order management and monitoring module 1210 may be configured to collect usage statistics about the service in the subscription order, such as the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and system downtime.

[0182] In certain embodiments, the cloud infrastructure system 1202 may include an identity management module 1228. The identity management module 1228 may be configured to provide identity services in the cloud infrastructure system 1202, such as access management and authorization services. In some embodiments, the identity management module 1228 may control information about customers who wish to use the services provided by the cloud infrastructure system 1202. Such information may include information that authenticates the identity of such customers and information that describes what actions those customers are authorized to perform on various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). The identity management module 1228 may also include managing descriptive information about each customer, as well as descriptive information about how and by whom this descriptive information may be accessed and modified.

[0183] Figure 13 shows an exemplary computing subsystem 1300 that can implement various embodiments. The computing subsystem 1300 can be used to implement any of the computing subsystems described above. As shown in the figure, the computing subsystem 1300 includes a processing unit 1304 that communicates with several peripheral subsystems via a bus subsystem 1302. These peripheral subsystems may include a processing acceleration unit 1306, an I / O subsystem 1308, a storage subsystem 1318, and a communication subsystem 1311. The storage subsystem 1318 includes a tangible computer-readable storage medium 1309 and a system memory 1310.

[0184] The bus subsystem 1302 provides a mechanism for various components and subsystems of the computing subsystem 1300 to communicate with each other as intended. Although the bus subsystem 1302 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 1302 may be one of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of the various bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus, which can be implemented as a mezzanine bus manufactured to the IEEE P1186.1 standard.

[0185] A processing unit 1304, which can be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of the computing subsystem 1300. The processing unit 1304 may include one or more processors. These processors may include single-core or multi-core processors. In certain embodiments, the processing unit 1304 may be implemented as one or more independent processing units 1332 and / or 1334, each containing a single-core or multi-core processor. In other embodiments, the processing unit 1304 may be implemented as a quad-core processing unit formed by integrating two dual-core processors onto a single chip.

[0186] In various embodiments, the processing unit 1304 may execute various programs in response to program code and may maintain multiple programs or processes running simultaneously. At any given time, some or all of the program code to be executed may reside in the processing unit 1304 and / or the storage subsystem 1318. Through suitable programming, the processing unit 1304 may provide the various functions described above. The computing subsystem 1300 may also include a processing acceleration unit 1306, which may include a digital signal processor (DSP), a special-purpose processor, and the like.

[0187] The I / O subsystem 1308 may include a user interface input device and a user interface output device. The user interface input device may include a pointing device such as a keyboard, mouse, or trackball, a touchpad, or a touchscreen, which are incorporated into a display, scroll wheel, click wheel, dial, button, switch, keypad, or audio input device, along with a voice command recognition system, microphone, and other types of input devices. The user interface input device may include, for example, a motion detection and / or gesture recognition device such as a Microsoft Kinect® motion sensor, which enables the user to interact with an input device, such as a Microsoft Xbox® 360 game controller, by controlling it through a natural user interface using gestures and spoken commands. The user interface input device may also include an eye gesture recognition device such as a Google Glass® blink detector, which detects eye movements from the user (e.g., blinking while taking a picture and / or selecting a menu) and translates the eye gesture into input to an input device (e.g., Google Glass®). Furthermore, the user interface input device may include a voice recognition detection device that enables the user to interact with a voice recognition system (e.g., Siri® Navigator) via voice commands.

[0188] Furthermore, user interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads and graphic tablets, audio / visual devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and gaze detection devices. User interface input devices may also include, for example, medical imaging input devices such as computed tomography, magnetic resonance imaging, positional emission tomography, and medical ultrasound equipment. Additionally, user interface input devices may include, for example, audio input devices such as MIDI keyboards and digital musical instruments.

[0189] User interface output devices may include non-visual displays such as display subsystems, indicator lights, or audio output devices. Display subsystems include cathode ray tubes (CRTs) and liquid crystal displays (LCDs). This may include flat panel devices such as those using an al Display or plasma display, projection devices, touchscreens, etc. Generally, the use of the term “output device” is intended to include all feasible types of devices and mechanisms for outputting information from the computing subsystem 1300 to a user or other computer. For example, user interface output devices may include, but are not limited to, a variety of display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, car navigation systems, plotters, audio output devices, and modems.

[0190] The computing subsystem 1300 may include a storage subsystem 1318 having software elements that are currently shown to be located in the system memory 1310. The system memory 1310 may store program instructions that can be loaded and executed on the processing unit 1304, and data generated during the execution of these programs.

[0191] Depending on the configuration and type of the computing subsystem 1300, the system memory 1310 may be volatile (such as Random Access Memory (RAM)) and / or non-volatile (such as Read-Only Memory (ROM) or flash memory). RAM typically contains data and / or program modules that are immediately accessible to the processing unit 1304, and / or data and / or program modules currently operating and executing by the processing unit 1304. In some implementations, the system memory 1310 may include several different types of memory, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). In some implementations, during startup... This includes a basic input / output system (BIOS) containing basic routines that help transfer information between elements within the computing subsystem 1300. However, it can typically be stored in ROM. As an example and not limited to, system memory 1310 may include application programs 1312, program data 1, etc., such as client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc. 314 and Operating System 1316 are also shown. For example, Operating System 1316 is a Microsoft Windows®, Apple Macintosh® and / or Linux operating system, various commercially available UNIX® or UNIX-like operating systems (various GNU / Linux operating systems, Google This includes, but is not limited to, Chrome® OS, and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® 10 OS, and Palm® OS.

[0192] Furthermore, the storage subsystem 1318 may provide a tangible, computer-readable storage medium for storing basic programming and data structures that provide several embodiments of functionality. Software (programs, code modules, instructions) that provides the above-mentioned functionality when executed by the processor may be stored in the storage subsystem 1318. These software modules or instructions may be executed by the processing unit 1304. The storage subsystem 1318 may also provide a repository for storing data used in accordance with the present invention.

[0193] Furthermore, the memory subsystem 1318 is further connected to the computer-readable storage medium 1309. It may include a computer-readable storage medium reader 1320 that can be connected to it. Together and optionally in combination with the system memory 1310, the computer-readable storage medium 1309 may comprehensively represent remote, local, fixed, and / or removable storage devices, in addition to storage media for temporarily and / or permanently storing, storing, transmitting, and retrieving computer-readable information.

[0194] The computer-readable storage medium 1309 containing code or a portion of code may include any suitable medium known or used in the art. Such medium includes, but is not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing and / or transmitting information. This may include tangible, non-temporary computer-readable storage media such as RAM, ROM, Electronically Erasable Programmable ROM (EEPROM), flash memory or other memory technologies, CD-ROM, Digital Versatile Disk (DVD), or other optical storage devices, magnetic cassettes, magnetic tapes, magnetic disk storage devices or other magnetic storage devices, or other tangible computer-readable media. Where specified, this may also include intangible, temporary computer-readable media such as data signals, data transmissions, or any other medium that can be used to transmit desired information and is accessible by the computing subsystem 1300.

[0195] As an example, computer-readable storage media 1309 may include hard disk drives that read from or write to non-removable non-volatile magnetic media, magnetic disk drives that read from or write to removable non-volatile magnetic disks, and optical disk drives that read from or write to removable non-volatile optical disks such as CD-ROMs, DVDs, and Blu-ray® discs or other optical media. Computer-readable storage media 1309 may include, but are not limited to, zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD discs, and digital videotapes. Furthermore, computer-readable storage media 1309 may include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory such as solid-state ROMs, and solid-state SSDs may include those based on volatile memory such as RAM, dynamic RAM, and static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. Disk drives and their associated computer-readable media may provide computer-readable instructions, data structures, program modules, and other data to the computing subsystem 1300.

[0196] The communication subsystem 1311 provides interfaces with other computing subsystems and networks. The communication subsystem 1311 acts as an interface for receiving data from other systems and transmitting data from computing subsystem 1300 to other systems. For example, the communication subsystem 1311 may enable computing subsystem 1300 to connect to one or more devices via the Internet. In some embodiments, the communication subsystem 1311 provides radio frequency (RF) access to radio voice and / or data networks (using, for example, cellular technologies such as 3G, 4G, or EDGE (enhanced data rates for global evolution), advanced data network technologies). ) Transceiver components, WiFi (IEEE 902.9 family standard, or other mobile communication technologies, or any combination thereof), Global Positioning System (GPS: GL obal Positioning System) receiver component, and / or other components This may include a wireless interface. In some embodiments, the communication subsystem 1311 may provide a wired network connection (e.g., Ethernet) in addition to, or instead of, the wireless interface.

[0197] In some embodiments, the communication subsystem 1311 may receive input communications on behalf of one or more users who may use the computing subsystem 1300, in the form of structured and / or unstructured data feeds 1326, event streams 1328, event updates 1330, etc.

[0198] As an example, the communication subsystem 1311 handles Twitter® feeds, Facebook® updates, and Rich Site Summary (RSS) feeds. It may be configured to receive data feeds 1326, such as web feeds, in real time from users of social media networks and / or other communication services, and / or to receive real-time updates from one or more third-party sources.

[0199] In addition, the communication subsystem 1311 may be configured to receive data in the form of a continuous data stream. This data may include an event stream 1328 and / or event updates 1330 of real-time events, which may be continuous or have no boundaries in an essentially definite-end state. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring.

[0200] Furthermore, the communication subsystem 1311 may be configured to output structured and / or unstructured data feeds 1326, event streams 1328, event updates 1330, etc., to one or more databases that can communicate with one or more streaming data source computers coupled to the computing subsystem 1300.

[0201] The computing subsystem 1300 may be one of a variety of types, including handheld portable devices (e.g., iPhone® mobile phones, iPad® computing tablets, PDAs), wearable devices (e.g., Google Glass® head-mounted displays), PCs, workstations, mainframes, kiosks, server racks, or other data processing systems.

[0202] Due to the constantly changing nature of computers and networks, the description of the computing subsystem 1300 shown in the figure is intended only as a specific example. Many other configurations are possible, having more or fewer components than the system shown in the figure. For example, customized hardware may be used, and / or certain elements may be implemented, in hardware, firmware, software (including applets), or combinations. Furthermore, connections to other computing devices, such as network input / output devices, may be utilized. Based on the disclosures and teachings provided herein, those skilled in the art will understand other means and / or methods for realizing various embodiments. Various embodiments of this disclosure may be realized using computer program products that include computer programs / instructions that, when executed by a processor, cause the processor to perform any of the methods disclosed herein.

[0203] While the embodiments of the invention are described in the above specification with reference to specific embodiments, those skilled in the art will recognize that the invention is not limited thereto. The various features and embodiments of the invention described above may be used individually or in combination. Furthermore, embodiments may be used in many environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings should be considered illustrative rather than restrictive.

Claims

1. A computer-implemented method for generating a multimodal discourse tree, The steps include obtaining a text corpus and one or more data records separate from the text corpus, The method, implemented by the computer, further comprises the steps of generating an extended discourse tree for the corpus of the text, wherein the extended discourse tree comprises a plurality of discourse trees, each discourse tree comprises a plurality of nodes, each terminal node of the discourse tree corresponds to a fragment of the text, each non-terminal node of the discourse tree indicates a rhetorical relationship between the nodes of the discourse tree, and the extended discourse tree comprises further links between the plurality of discourse trees indicating further rhetorical relationships between the nodes of each discourse tree, and the method implemented by the computer further comprises The method, implemented by the computer, includes the step of identifying entity matches between a set of basic discourse units of the plurality of discourse trees and one or more data records, wherein the entity matches are identified by comparing a first entity identified from the basic discourse unit with a second entity identified from the data record, and further includes the step of identifying entity matches between a set of basic discourse units of the plurality of discourse trees and one or more data records, wherein the entity matches are identified by comparing a first entity identified from the basic discourse unit with a second entity identified from the data record, and the method implemented by the computer further includes The steps include identifying one or more causal relationship links between two of the one or more data records, For each identified entity match and each of the identified one or more causal links, the steps include determining the corresponding rhetorical relationship, For the extended discourse tree, the steps include generating a node for each identified entity match and each identified causal link, A computer-implemented method comprising the steps of creating the multimodal discourse tree by linking the respective nodes generated for each entity match and each causal link to the respective nodes of the extended discourse tree, at least in part based on the determined corresponding rhetorical relationships.

2. The step of generating the extended discourse tree is: The step of generating an extended discourse tree includes the step of generating a first discourse tree from a first text of the text corpus, wherein the first discourse tree corresponds to a first portion of the first text, and the step of generating an extended discourse tree further includes the step of generating an extended discourse tree. The step of generating an extended discourse tree includes the step of generating a second discourse tree from a second text of the corpus of the aforementioned text, wherein the second discourse tree corresponds to a second portion of the second text, and the step of generating the extended discourse tree further includes the step of generating an extended discourse tree A computer-implemented method according to claim 1, comprising the step of linking the first and second discourse trees using the specific rhetorical relationships in response to determining specific rhetorical relationships between the respective basic discourse units of the first and second discourse trees.

3. The computer-implemented method according to claim 2, wherein the first discourse tree and the second discourse tree are communication discourse trees, each containing a verb signature generated for each basic discourse unit of the first discourse tree and the second discourse tree.

4. The step of identifying the entity match further includes comparing the second entity identified from the data record with a third entity identified from the second data record, wherein the entity refers to one of (i) a person, (ii) a company, (iii) a place, (iv) a document name, (v) a date or time, (vi) a transaction, or (vii) an activity, as implemented by the computer according to claim 2. How to do it.

5. The computer-based method according to claim 1, further comprising the step of identifying the entity from a predetermined ontology.

6. The process further includes a step of classifying subsequent inputs based at least partially on the multimodal discourse tree, wherein the step of classifying subsequent inputs is: The steps include generating a training dataset containing multiple multimodal discourse trees, each multimodal discourse tree corresponding to a text corpus and a set of data records, each multimodal discourse tree associated with a label corresponding to a classification, and further classifying the subsequent input. A step of training a machine learning model to classify inputs based at least partially on the aforementioned training dataset and supervised learning algorithm, The step of generating a corresponding multimodal discourse tree from the subsequent input, wherein the subsequent input includes a corresponding text corpus and a corresponding set of data records, and the step of classifying the subsequent input further includes: A computer-implemented method according to claim 1, comprising the step of classifying the subsequent input, at least in part, by providing the corresponding multimodal discourse tree generated from the subsequent input as input to the machine learning model and receiving an output from the machine learning model indicating a classification label for the subsequent input.

7. The process further includes the step of navigating the corpus of text using the multimodal discourse tree, the step of navigating the corpus of text is: The steps include accessing the aforementioned multimodal discourse tree, The steps include determining a first basic discourse unit from the multimodal discourse tree to respond to a query from a user device, wherein the first basic discourse unit corresponds to a first node of a first discourse tree in the multimodal discourse tree, and further navigating the text corpus, The steps include determining a set of navigation options from the multimodal discourse tree, the set of navigation options including at least two of (i) a first rhetorical relationship between the first node of the first discourse tree and the second node of the first discourse tree, (ii) a second rhetorical relationship between the first node and the third node of the second discourse tree of the multimodal discourse tree, or iii) a third rhetorical relationship between the first node of the first discourse tree and the fourth node of the multimodal discourse tree associated with the corresponding data record, and the steps of navigating the corpus of text further include The steps include presenting at least two of the first, second, or third rhetorical relationships to the user device, In response to receiving further user input, including the selection of the first rhetorical relationship, the second rhetorical relationship, or the third rhetorical relationship, A step of presenting a second basic discourse unit corresponding to the second node, at least in part on the decision of the selection corresponding to the first rhetorical relation, A step of presenting a third basic discourse unit corresponding to the third node, at least in part, based on the selection decision corresponding to the first rhetorical relationship, or A computer-implemented method according to claim 1, comprising the step of presenting at least a portion of the corresponding data records based at least in part on the selection decision corresponding to the third rhetorical relationship.

8. A computing device, One or more processors, The computing device comprises one or more memories for storing computer-executable instructions for generating a multimodal discourse tree, wherein, when the instructions are executed by the one or more processors, the computing device shall Obtaining a text corpus and one or more data records separate from the said text corpus, The instruction causes the computer to generate an extended discourse tree for the corpus of the text, the extended discourse tree comprising a plurality of discourse trees, each discourse tree comprising a plurality of nodes, each terminal node of the discourse tree corresponding to a fragment of the text, each non-terminal node of the discourse tree indicating a rhetorical relationship between the nodes of the discourse tree, and the extended discourse tree comprising further links between the plurality of discourse trees indicating further rhetorical relationships between the nodes of each discourse tree, and the instruction, when executed by the one or more processors, causes the computing device to The instruction causes the system to identify entity matches between a set of basic discourse units in the plurality of discourse trees and one or more data records, wherein the entity matches are identified by comparing a first entity identified from the basic discourse unit with a second entity identified from the data record, and the instruction, when executed by one or more processors, causes the computing device to: Identifying one or more causal links between two of the aforementioned one or more data records, For each identified entity match and each of the identified one or more causal links, determine the corresponding rhetorical relationship, For the aforementioned extended discourse tree, a node is generated for each identified entity match and each identified causal link, A computing device that causes the device to create the multimodal discourse tree by linking the respective nodes generated for each entity match and each causal link to the respective nodes of the extended discourse tree, at least partially based on the determined corresponding rhetorical relationships.

9. Performing the operation to generate the extended discourse tree further involves the computing device, The process involves generating a first discourse tree from a first text in the corpus of the aforementioned text, the first discourse tree corresponding to a first portion of the first text, and further performing the operation to generate the extended discourse tree. The process involves generating a second discourse tree from a second text in the corpus of the aforementioned text, the second discourse tree corresponding to a second portion of the second text, and further performing the operation to generate the extended discourse tree. The computing device according to claim 8, which, in response to determining specific rhetorical relationships between the basic discourse units of the first and second discourse trees, generates a link between the first and second discourse trees using the rhetorical relationships.

10. The computing device according to claim 9, wherein the first discourse tree and the second discourse tree are communication discourse trees, each containing a verb signature generated for each basic discourse unit of the first discourse tree and the second discourse tree.

11. Identifying the aforementioned entity match further causes the computing device to compare the second entity identified from the data record with the third entity identified from the second data record, wherein the entity is (i) a person, (ii) a company, (iii) a place, (iv) a document name, or (v) a date or The computing device according to claim 8, wherein refers to one of the time periods.

12. The computing device according to claim 8, wherein the entity matching further comprises identifying the entity from a predetermined ontology.

13. The computing device classifies subsequent inputs based at least partially on the multimodal discourse tree, and classifying the subsequent inputs is performed by the computing device, The system generates a training dataset containing multiple multimodal discourse trees, each multimodal discourse tree corresponding to a text corpus and a set of data records, and each multimodal discourse tree is associated with a label corresponding to a classification, and the subsequent input is further classified by the computing device. Training a machine learning model to classify inputs based at least partially on the aforementioned training dataset and supervised learning algorithm, The computing device is made to generate a corresponding multimodal discourse tree from the subsequent input, the subsequent input including a corresponding text corpus and a corresponding set of data records, and classifying the subsequent input is performed by the computing device. The computing device according to claim 8, which provides the corresponding multimodal discourse tree generated from the subsequent input as input to the machine learning mode, causing the machine learning model to classify the subsequent input, at least in part, on the basis of receiving an output from the machine learning model indicating a classification label for the subsequent input.

14. The computing device uses the multimodal discourse tree to navigate the corpus of text, and navigating the corpus of text further involves Accessing the aforementioned multimodal discourse tree, This includes determining a first basic discourse unit from the multimodal discourse tree to respond to a query from a user device, wherein the first basic discourse unit corresponds to a first node of the first discourse tree in the multimodal discourse tree, and further includes navigating the text corpus. Navigating the text corpus further includes determining a set of navigation options from the multimodal discourse tree, the set of navigation options including at least two of (i) a first rhetorical relationship between the first node of the first discourse tree and the second node of the first discourse tree, (ii) a second rhetorical relationship between the first node and the third node of the second discourse tree of the multimodal discourse tree, or iii) a third rhetorical relationship between the first node of the first discourse tree and the fourth node of the multimodal discourse tree associated with the corresponding data record, and further includes navigating the text corpus. Presenting at least two of the first, second, or third rhetorical relationships to the user device, In response to receiving further user input, including the selection of the first rhetorical relationship, the second rhetorical relationship, or the third rhetorical relationship, To present a second basic discourse unit corresponding to the second node, at least in part, based on the selection decision corresponding to the first rhetorical relationship, Presenting a third basic discourse unit corresponding to the third node, at least in part, based on the selection decision corresponding to the first rhetorical relationship, or The computing device according to claim 8, further comprising presenting at least a portion of the corresponding data records based at least in part on the selection decision corresponding to the third rhetorical relationship.

15. A non-temporary computer-readable storage medium for storing instructions for generating a multimodal discourse tree, wherein, when the instructions are executed by one or more processors of a computing device, the computing device... Obtaining a text corpus and one or more data records separate from the said text corpus, The instruction causes the computer to generate an extended discourse tree for the corpus of the text, the extended discourse tree comprising a plurality of discourse trees, each discourse tree comprising a plurality of nodes, each terminal node of the discourse tree corresponding to a fragment of the text, each non-terminal node of the discourse tree indicating a rhetorical relationship between the nodes of the discourse tree, and the extended discourse tree comprising further links between the plurality of discourse trees indicating further rhetorical relationships between the nodes of each discourse tree, and the instruction further causes the computer to generate an extended discourse tree for the corpus of the text, the extended discourse tree comprising a plurality of discourse trees, each discourse tree comprising a plurality of nodes, each terminal node of the discourse tree corresponding to a fragment of the text, each non-terminal node of the discourse tree indicating a rhetorical relationship the discourse tree, and the extended discourse tree comprising further links between the plurality of discourse trees indicating further rhetorical relationships between the nodes of each discourse tree, and when the instruction is executed by the one or more processors of the computing device, the computing device The instruction causes the system to identify entity matches between a set of basic discourse units in the plurality of discourse trees and one or more data records, wherein the entity match is identified by comparing a first entity identified from the basic discourse unit with a second entity identified from the data record, and the instruction, when executed by one or more processors of the computing device, causes the computing device to: Identifying one or more causal links between two of the aforementioned one or more data records, For each identified entity match and each of the identified one or more causal links, determine the corresponding rhetorical relationship, For the aforementioned extended discourse tree, a node is generated for each identified entity match and each identified causal link, A non-temporary, computer-readable storage medium that causes the multimodal discourse tree to be created by linking each node generated for each entity match and each causal link to each node of the extended discourse tree, at least partially based on the determined corresponding rhetorical relationships.

16. Performing the operation to generate the extended discourse tree further involves the computing device, The process involves generating a first discourse tree from a first text in the corpus of the aforementioned text, the first discourse tree corresponding to a first portion of the first text, and further performing the operation to generate the extended discourse tree. The process involves generating a second discourse tree from a second text in the corpus of the aforementioned text, the second discourse tree corresponding to a second portion of the second text, and further performing the operation to generate the extended discourse tree. A non-temporary computer-readable storage medium according to claim 15, which, in response to determining specific rhetorical relationships between the respective basic discourse units of the first and second discourse trees, causes the system to generate links between the first and second discourse trees using the rhetorical relationships.

17. The non-temporary computer-readable storage medium according to claim 16, wherein the first discourse tree and the second discourse tree are communication discourse trees, each containing a verb signature generated for each basic discourse unit of the first discourse tree and the second discourse tree.

18. Identifying the aforementioned entity match further involves the computing device identifying the second entity from the data record and from the second data record. A non-temporary computer-readable storage medium according to claim 15, which is made to be compared with an identified third entity, wherein the entity refers to one of (i) a person, (ii) a company, (iii) a place, (iv) the name of a document, or (v) a date or time.

19. The computing device classifies subsequent inputs based at least partially on the multimodal discourse tree, and classifying the subsequent inputs is performed by the computing device, The system generates a training dataset containing multiple multimodal discourse trees, each multimodal discourse tree corresponding to a text corpus and a set of data records, and each multimodal discourse tree is associated with a label corresponding to a classification, and the subsequent input is further classified by the computing device. Training a machine learning model to classify inputs based at least partially on the aforementioned training dataset and supervised learning algorithm, The computing device is made to generate a corresponding multimodal discourse tree from the subsequent input, the subsequent input including a corresponding text corpus and a corresponding set of data records, and classifying the subsequent input is performed by the computing device. A non-temporary computer-readable storage medium according to claim 15, wherein the corresponding multimodal discourse tree generated from the subsequent input is provided as input to the machine learning mode, causing the subsequent input to be classified, at least in part, on the basis of receiving an output from the machine learning model indicating a classification label for the subsequent input.

20. The computing device uses the multimodal discourse tree to navigate the corpus of text, and navigating the corpus of text further involves Accessing the aforementioned multimodal discourse tree, This includes determining a first basic discourse unit from the multimodal discourse tree to respond to a query from a user device, wherein the first basic discourse unit corresponds to a first node of the first discourse tree in the multimodal discourse tree, and further includes navigating the text corpus. Navigating the text corpus further includes determining a set of navigation options from the multimodal discourse tree, the set of navigation options including at least two of (i) a first rhetorical relationship between the first node of the first discourse tree and the second node of the first discourse tree, (ii) a second rhetorical relationship between the first node and the third node of the second discourse tree of the multimodal discourse tree, or iii) a third rhetorical relationship between the first node of the first discourse tree and the fourth node of the multimodal discourse tree associated with the corresponding data record, and further includes navigating the text corpus. Presenting at least two of the first, second, or third rhetorical relationships to the user device, In response to receiving further user input, including the selection of the first rhetorical relationship, the second rhetorical relationship, or the third rhetorical relationship, To present a second basic discourse unit corresponding to the second node, at least in part, based on the selection decision corresponding to the first rhetorical relationship, Presenting a third basic discourse unit corresponding to the third node, at least in part, based on the selection decision corresponding to the first rhetorical relationship, or A non-temporary computer-readable storage medium according to claim 15, comprising presenting at least a portion of the corresponding data records based at least in part on the selection decision corresponding to the third rhetorical relationship.