Explainable Natural Language Understanding Platform

JP2025515316A5Pending Publication Date: 2026-05-07GYAN INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GYAN INC
Filing Date
2023-04-24
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing natural language processing technologies are difficult to understand natural language content in a complete structure, lack global context understanding and interpretability, and rely on large-scale tasks to specific data, so they cannot work effectively in the case of scarce data.

Method used

Using knowledge-based computational linguistics approaches, combining discourse models, constitutive structures and rhetorical structures, a real-time human-like natural language understanding engine (NLU platform) is developed, which is able to understand natural language content without or with a small amount of training data and generate interpretable mechanical representations.

Benefits of technology

It realizes the ability to understand natural language content in a complete structure, provides global context understanding and interpretability, can work effectively in the case of scarce data, and is suitable for various natural language understanding tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000024_0000
    Figure 00000024_0000
  • Figure 00000024_0001
    Figure 00000024_0001
  • Figure 00000025_0000
    Figure 00000025_0000
Patent Text Reader

Abstract

A system and method for a real-time human-like natural language understanding engine (NLU platform). The platform understands natural language content in its full compositional context, is fully explainable, and is effective with little or no training data. The invention utilizes machine representation of natural language to enable human-like understanding by machines. The invention combines knowledge-based linguistics, discourse models, compositionality, and rhetorical structures in language. It can incorporate global knowledge about concepts to improve their representation and understanding. The platform does not rely on statistically derived distributional semantics and can be used for any natural language understanding task in the world. In one exemplary application, the NLU platform can be used for automated knowledge acquisition or as an automated research engine for any topic or set of topics from one or more document repositories such as the Internet.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to natural language understanding, computational linguistics and machine representation of meaning, and in particular to methods, computer programs and systems for human-like understanding of natural language data in the full range of automated natural language understanding tasks. [Background technology]

[0002] Natural language content represents a large proportion of content both on the public Internet and within private content repositories. Not surprisingly, therefore, there has been a considerable growth in the need for and interest in automated machine processing of natural language content. Nearly all currently available techniques for machine processing of natural language rely on large neural language models of word patterns, including patterns based on real-valued vector representations of each word. Such approaches are not designed to understand text in its full compositional context. Humans understand text in its full compositional context.

[0003] All current language models learn to predict the probability of a sequence of words. The generated sequence of words is associated with a query by the user. The next word with the highest probability is selected as the output by the model. These models are trained on a very large amount of raw text. Two categories of statistical methods have been used to build language models. One uses hidden Markov models and the other uses neural networks. Both are statistical approaches that rely on a large corpus of natural language text (the input data) for prediction. With regard to the input data, the bag of words approach is very popular. However, although the bag of words model is preferred due to its simplicity, it does not preserve any context and different words have different representations every time, regardless of how they are used. This has led to the development of a new technique called word embedding, which tries to create similar representations for similar words. The famous adage "A word is known by the company it keeps" is the motivation for word embedding. Proponents of word embedding believe that this is similar to capturing the meaning of a word and its usage.

[0004] Word embeddings have some significant limitations, mainly related to the corpus-dependent nature of such embeddings and their inability to handle the different senses in which a word is used: word representations learned based on a corpus to determine similar words are not equivalent to a human-like machine representation of the meaning of the entire text.

[0005] Current language models have significant limitations, including an inability to preserve configuration context, a lack of explainability, and the need for large-scale task-specific data to be combined with neural language models. Additionally, current language models are far from universal in their ability to "learn" or "transfer" knowledge from one domain to the next.

[0006] More recently, there has been a trend to integrate external knowledge into these models to address some of their limitations, especially in the case of context preservation and context-specific global knowledge. Although these enhancements improve the performance of language models, they are still based on word-level patterns. They do not address fundamental issues such as full contextual understanding of constructions, explainability and the need for large-scale task-specific data. In the real world, large amounts of data are often lacking for most natural language processing or understanding tasks.

[0007] As an example of the need for machine natural language understanding, consider "research," a common task most knowledge workers perform on a daily basis. It is estimated that there are roughly one billion knowledge workers in the world today, and the number is growing. Examples of knowledge workers include students, teachers, academic researchers, analysts, consultants, scientists, lawyers, marketers, etc. The research or knowledge acquisition tasks they perform today involve the use of search engines on the Internet, and the use of search techniques against in-house or private document collections. Search engines have provided great value by providing access to previously inaccessible information, but they are only the starting point of research.

[0008] Today's search engines still cannot fully understand the text within a document, so there are both false positives and false negatives in the result set. In a setting that shows 10 items per page based on a query, irrelevant results may appear as early as page 1 and certainly on subsequent pages. Similarly, relevant results may be listed much later, such as pages 10-20.

[0009] More importantly, today's technology does not provide any assistance with the rest of the process: determining relevance, organizing, reading, understanding, and synthesizing knowledge from each document, and developing a collective understanding across all documents. Current search technologies or language models cannot understand the full text within these documents, and therefore cannot automatically assist in the complete research or knowledge acquisition process.

[0010] Other exemplary use cases that require a complete understanding of unstructured content include tagging unstructured content, comparing multiple documents to each other to detect plagiarism or duplication, automating the generation of intelligent natural language output, etc. For the avoidance of doubt, unstructured content could be video or audio data that can be transcribed into text.

[0011] It is therefore clear that there is a need for methods and systems that automate human-like understanding of textual discourse in its full composed context in an explainable and tractable manner, and even in situations where data is sparse. Summary of the Invention

[0012] A system and method for a real-time human-like natural language understanding engine (NLU platform) is disclosed. The NLU platform understands natural language content in its full compositional context, is fully explainable, and is effective with little or no training data. The present invention discloses a novel method of machine representation of natural language that enables human-like machine understanding. The present invention combines knowledge-based linguistics, discourse models, compositionality, and rhetorical structures in the language. It improves its representation and understanding by incorporating global knowledge about the concepts. The present invention does not utilize statistically derived distributional semantics. The disclosed invention can be used for a full range of natural language understanding tasks performed universally.

[0013] In one exemplary application, the NLU platform is used for automated knowledge acquisition or as an automated exploration engine for any topic or set of topics from one or more document repositories such as the Internet. This includes determining the relevance of each item, organizing the relevant items into a table of contents or automatically deriving a table of contents from its human-like understanding of each item, creating a multi-level summary of each relevant item, creating a multi-level semantic graph for each item, and identifying the main topic of each item. Content from multiple content items is integrated to create an integrated cross-document multi-level semantic representation and a multi-level integrated summary across multiple documents. The NLU platform automatically discovers new knowledge on an ongoing basis about all aspects of the initially discovered knowledge, such as the main topics.

[0014] One or more aspects of the present invention are particularly pointed out and distinctly claimed by way of example in the claims at the conclusion of the specification. These and other objects, features, and advantages of the present invention will become apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief description of the drawings]

[0015] [Figure 1] The following shows the results of a relevance analysis of 10 search queries using popular Internet search engines. [Diagram 2] 1 illustrates an NLU platform high-level architecture in accordance with one or more embodiments described herein. [Diagram 3] 1 illustrates an NLU platform pipeline according to one or more embodiments described herein. [Figure 4] 1 illustrates an example NLU platform semantic representation graph main layer, according to one or more embodiments described herein. [Diagram 5] 1 illustrates an NLU platform semantic representation graph sentence level, according to one or more embodiments described herein. [Figure 6] 1 illustrates an NLU platform semantic representation graph concept level layer, according to one or more embodiments described herein. [Figure 7] 1 illustrates several discourse models according to one or more embodiments described herein. [Figure 8A] Here is an example of a news article: [Figure 8B] Here is an example of a news article: [Figure 8C] Here is an example of a news article: [Figure 8D] Here is an example of a news article: [Figure 9A] 1 illustrates NLU platform 200 determining the relevance of a document to a topic of interest. [Figure 9B] 1 illustrates NLU platform 200 determining the relevance of a document to a topic of interest. [Figure 10] We present an analysis of results from NLU platform 200 for the same 10 queries reported in Figure 2. Note: After concept detection (occurrence of concepts in the document's SKG), for complex concepts (multiple topics), besides sentence importance (main idea, main cluster, etc.) and concept importance (theme, etc.), NLU platform 200 also determines the context strength. Context strength determines how well a document references multiple concepts within a complex concept in the context of each other. [Figure 11] 4 illustrates a partial NLU platform semantic representation graph concept layer for the example news article of FIG. 3, in accordance with one or more embodiments described herein. [Figure 12] 10 illustrates an example outline for the document of FIG. 9 created by the NLU platform 200, according to one or more embodiments described herein. [Figure 13] 2 illustrates an exemplary knowledge collection on the topic of electric vehicles using NLU platform 200. [Figure 14]1 illustrates an exemplary real-time knowledge discovery process in which an NLU platform proactively discovers new knowledge on any topic and integrates it with an existing semantic representation graph, according to one or more embodiments described herein. [Figure 15] 1 illustrates the creation of top topics across a collection of documents, in which the NLU platform 200 combines top topics across multiple documents, in accordance with one or more embodiments described herein. [Figure 16] 2 illustrates an example of an exemplary aggregated semantic representation in which the NLU platform 200, in accordance with one or more embodiments described herein, consolidates the semantic representation structures of individual documents by eliminating redundant semantics. [Figure 17] 1 illustrates how a user can input global knowledge about a topic that the NLU platform 200 considers when determining relevance or synopsis, according to one or more embodiments described herein. [Figure 18] 2 illustrates how a user can input symbol rules to identify specific senses of different words for consideration by the NLU platform 200, according to one or more embodiments described herein. [Figure 19] Demonstrates the ACL Platform I Learning Content Authoring / Assembly component, which enables any institution to rapidly assemble learning content on emerging topics or topics related to specific skills. [Figure 20] The results of the ACL Platform I pilot are shown below. As you can see, all 10 students received perfect peer grades. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016] Background technology There are countless use cases for automatic machine processing of natural language content around the world. Nearly all currently available technologies for machine processing of natural language rely on neural language models of word patterns, which contain patterns based on the semantic representation of each word. This approach fails to understand text in its full composed context. Humans understand text in its full composed context. For the avoidance of doubt, unstructured content refers to video or audio data that has been transcribed into text for processing.

[0017] Traditionally, the "bag of words" method has been a common approach for natural language processing (NLP) tasks. Bag of words models are often used in NLP problems that require classification. The majority of NLP tasks, such as internet search, are classification problems. In the bag of words (BoW) model, text (such as a sentence or document) is represented as a collection of words, ignoring grammar and word order. The frequency of occurrence of each word across a collection of documents is used as a feature to train a classifier. The classifier is then applied to new documents to classify them into a set of categories, identify topics, or determine a match to a search query (e.g., a search query for a search engine). Because this method focuses only on the frequency of words occurring in a document, the BoW model does not preserve the context of where and how the words were used and does not understand the meaning of the documents being processed.

[0018] To overcome the limitations of the popular "bag of words" method, several large pre-trained language models have been introduced. Two categories of statistical methods have been used for building language models. One uses probabilistic approaches such as hidden Markov models, and the other uses neural networks. Both rely on large corpora of natural language texts for prediction. In bag of words models, different words have different representations, regardless of how they are used. This has led to the development of a new technique called word embedding, which attempts to create similar representations for similar words. The famous adage "A word is known by the company it keeps" motivates word embedding. Proponents of word embedding believe that this is similar to capturing the meaning of a word and its usage.

[0019] Word embeddings have several significant limitations, mainly related to the corpus-dependent nature of such embeddings and their inability to handle the different senses in which a word is used. Moreover, word representations learned based on corpora to determine similar words bear little resemblance to human-like semantic representations of the full text.

[0020] Language models learn to predict the probability of a sequence of words. The next word with the highest probability is selected by the model as the output. These models are trained on very large amounts of raw text. The probability of a sequence of words given preceding words or other word characteristics obviously does not represent the meaning of the text. These models preserve all the biases of the data used to train the model. For example, Marcus and David (2021) found that GPT3, one of the more powerful language models, outputs the following bold italicized text as a continuation of a provided sentence:

[0021] "You pour yourself a glass of cranberry juice, but then you absentmindedly pour about a teaspoon of grape juice into it. It looks okay. You try sniffing it, but you have a bad cold, so you can't smell anything. You are very thirsty. So you drink it. You are now dead."

[0022] It is also a well-known fact that NLP applications such as chatbots that rely on language models can spit out racist or sexist comments, as Microsoft discovered with the release of its chatbot Tay. It has also been noted that typing "three Muslims" into GPT2 often results in text portraying them as terrorists or criminals. These examples demonstrate the biases introduced by training datasets.

[0023] Current language models are far from universal in their ability to "learn" or "transfer" knowledge from one application to the next. These models suffer from the same significant limitations, including the inability to preserve full context, lack of explainability, the need for large task-specific data to combine with neural language models, and biases from the data.

[0024] Recently, there has been a trend to integrate external knowledge into the models to address some of the limitations of these models. This work commonly advocates integrating some form of external knowledge to pre-train large-scale knowledge-augmented LMs. These augmentations improve the performance of language models, but do not address at all the fundamental problem of explainable semantic representation of natural language texts. Bender and Koller (2017) argue that language models do not in principle lead to learning meaning. They properly define "meaning" as human-like natural language understanding, and argue that it is disingenuous to claim that these language models come close to understanding meaning.

[0025] Natural language discourse involves structuring the author's ideas and knowledge, using rhetorical devices when necessary. Good discourse is coherent and follows widely accepted principles and rules regarding discourse structure and content. At a higher level, there are many different types of discourse models: news articles, blogs, research papers, legal documents, emails, social media posts, etc. Each of these has a well-defined discourse model that is understandable by experts in the field, e.g. lawyers understand the discourse model underlying legal contracts, most of us understand the model for emails, and news articles follow a discourse model. A human-like semantic representation model needs to be aware of the discourse model underlying the construction. Moreover, such representations need to be explainable, which is not possible with any current method.

[0026] As an example of the need for machine natural language understanding, consider research, a common task that most knowledge workers perform on a daily basis. Today, it is estimated that there are roughly one billion knowledge workers in the world, and that number is growing. Examples of knowledge workers include students, teachers, academic researchers, analysts, consultants, scientists, lawyers, marketers, etc. Knowledge workers exist in all fields. The research or knowledge acquisition tasks they perform today involve the use of search engines on the Internet, as well as the use of search techniques against in-house or private document collections. Search engines have provided great value by providing access to information that was previously inaccessible. However, they only provide access, which is only the starting point for research. However, current search engines, even when they provide access, whether on the Internet or on private document repositories, generate significant false positives due to the limitations outlined above.

[0027] When a user performs a query using a search engine such as Google, the results typically display 1-10 items from a large number of hits. Because today's search engines still cannot fully understand the text within documents, there are both false positives and false negatives in the result set. In a setting that displays 10 items per page based on the query, irrelevant results may appear as early as page 1 and will certainly appear on subsequent pages regardless of the nature of the query. Similarly, highly relevant results may appear much further down the list, such as pages 10-20, relative to the query. When dealing with results from today's search technologies, users must first open each item to determine which search results are truly relevant to their query.

[0028] In the experiment, a search for “data science and ethics” was performed in MIT’s popular open courseware content collection. Over 80% of the results were irrelevant, including 5 of the 10 results on the first page. Similarly, Figure 1 shows the results of an analysis of 10 random queries using popular search engines. The selected queries had one, two, and three words, some of which had very ambiguous words. For each query, human experts reviewed all 1,785 articles in the entire result set from the search engine and classified the articles as relevant or irrelevant. In the analysis, the home page, index page, audio and video pages, books, and shopping pages were considered irrelevant. Figure 1 shows the relevant and irrelevant results according to the browser’s default page view. Across the 10 queries, Google returned only 59% of the relevant results. As can be seen, later pages are full of relevant results and vice versa. Users must first read all documents to determine relevance.

[0029] The level of inaccuracy in search engine results is only the first problem. This step is the "access" step of the research task. Current technology does not assist with the rest of the steps. Once a knowledge worker has identified and acquired the relevant documents, the next step is to organize them for effective manual processing. This can take the form of organizing into subtopics. A human must then read each document, understand and integrate the knowledge within each document to develop a holistic understanding across all documents. Furthermore, knowledge acquisition is not a one-time event, so research efforts must be continually updated with the latest knowledge on the topic. Current search technology cannot automatically assist with the complete research or knowledge acquisition process from unstructured content.

[0030] Text summarization is another NLU task that requires a full understanding of the document's content, making it an integrated task in the research workflow described above.

[0031] An increasing amount of text accumulates on the Internet and in private corpora, and to be useful it needs to be summarized. Naturally, the need is to summarize text automatically without losing meaning and coherence. Automatic text summarization is the task of producing a concise and fluent summary without human assistance, while preserving the meaning of the original text document.

[0032] There are two types of text summaries: extractive and abstract. In extractive text summaries, the summary preserves the original text of the content. Abstract summaries attempt to replicate how humans summarize documents using their global knowledge. With current technology, extractive summaries are only partially successful and abstract summaries are very inadequate. Effective extractive summaries can be very helpful to knowledge workers.

[0033] Topic modeling is another very common NLU task. Current techniques fall short in identifying top topics within documents because they are not designed to understand the content of the document in its full compositional context. A typical digital publishing site wanted to make it easier for users or subscribers to find articles and content that interest them. The organization defined a topic ontology to classify the content documents. A typical language model achieved only 20% correct classification. The model is unexplainable and trained on a very large corpus. Clearly, the training data may not be sufficient to classify the site's content, and the model does not focus on understanding the full content of each document before identifying the top topics.

[0034] Other example use cases that require a complete understanding of unstructured content include tagging unstructured content, comparing multiple documents to each other to detect plagiarism or duplication, and automating the generation of intelligent natural language output.

[0035] Therefore, there is a clear need for a platform that automatically interprets natural language content in its full compositional context, works with little or no data, and is explainable and easy to handle. The method and system should also be able to integrate relevant knowledge beyond the discourse of the text. In the research or knowledge acquisition example above, the platform should be able to automate all research tasks or significantly assist knowledge workers. And in the topic modeling example, the platform should be able to improve the identification of top topics for each document with a full understanding of the content, including its compositional nature.

[0036] SUMMARY OF THE PRESENT APPLICATION In this section, we describe a Natural Language Understanding Platform [NLU Platform] based on deep learning using knowledge-based computational linguistics [CL]. This platform is also called the constituent language model. The platform retains full context and traceability and does not necessarily require data to operate. Thus, although the invention relies on deep language learning, it does not require learning from data, as in traditional machine learning, which must be trained with data to operate. If there is data available, the platform can leverage it, and the platform only requires a small fraction of the data required for modern machine learning techniques. This is observed from several result sets provided later in this disclosure.

[0037] The NLU platform generates a semantic representation of natural language. This representation enables zero-shot or few-shot learning. The use of the NLU platform is illustrated for an example NLU task for automated research that requires human-like understanding. In this disclosure, the word "document" is used frequently, but the definition also includes video or audio content that has been transcribed into text for processing. In this disclosure, the words "document" and "content item" are used interchangeably.

[0038] In general, the longer the text, the more complex the task of comprehension, and therefore representation. The longer the text, the more complex it is for humans to understand, and any machine representation will inherit that complexity. Comprehension is conceptualized in terms of "layers", starting with a high-level understanding of the main ideas of a document and broadening to understand additional details or layers. A user may want to understand the essence of an article, the explicit theme of an article, an article from a particular perspective, the personal viewpoint expressed in the article, and so on. This suggests that users need flexible, layered representations to "understand" discourse.

[0039] A universal semantic representation of natural language involves breaking down a text back into the concepts and discourse relations that authors used to express their ideas. This requires reverse engineering the text back into its basic ideas, as well as the phrases and building blocks that combine to form sentences and paragraphs that lead to a complete document. Then, to fully understand it in the same way a human would, it is necessary to integrate prior / global knowledge.

[0040] More specifically, the machine representation needs to be able to answer or support the following needs:

[0041] What is the main idea of ​​the document?

[0042] · What are the other ideas in the document?

[0043] · How is the idea supported by other ideas and reasoning?

[0044] Can it be used to create an overview of the contents of a document without having to read the entire document?

[0045] · Could it be used to create idea-specific summaries of documents based on interest in particular ideas within the document?

[0046] Can it be used to create a layered representation of comprehension, allowing readers to peel back layers based on their own interests and pace?

[0047] Can it be used to reduce cognitive load through multi-layered presentations and visually effective displays such as graphs?

[0048] Can representations be enriched with global knowledge external to the document, just as humans naturally do?

[0049] · Can it be used to dynamically create knowledge collections on any topic?

[0050] Can it be used to aggregate / synthesize understanding across multiple documents or knowledge collections?

[0051] Can it be used to identify major topics across a document collection or knowledge collection?

[0052] Can it be used to generate natural language responses?

[0053] FIG. 2 illustrates the NLU platform 200 at a high level. FIG. 2 provides an overview of the major components of the NLU platform and their relationships, according to a preferred embodiment of the present invention. A user 205 interacts with an NLU platform application 210 to conduct ongoing research or discover knowledge. The application 210 forwards one or more topics of interest to the user 205 to the NLU platform 215. The NLU platform 200 encapsulates all the document level 250 and document collection level components shown in FIG. 2 on the right. The NLU platform 200 looks for content for the topic(s) using any of the common search engines on the Internet 220. The NLU platform 200 can also aggregate content from the public Internet 220, private or internal document repositories 225, and public document repositories 228 such as PubMed, USPTO, etc. The NLU platform 200 then processes the results to generate output that is displayed as document level 250 and document collection level 280 components. The stored knowledge 240 contains knowledge previously created by the user at the topic or concept level. Global knowledge 260 contains knowledge available to all users about concepts and relationships. Discourse models 290 contain discourse structures for different discourse types such as emails, contracts, research documents, among others. Both stored knowledge and global knowledge are not explicitly present in the document being processed. For example, "Can I Xerox this document?" implies copying the document.

[0054] FIG. 3 illustrates the process flow of the NLU platform. When the NLU platform 200 receives a user's topic(s) or query from the NLU platform application 210, the query is normalized by the query normalizer 302. The query normalizer 302 examines the query to extract concepts and, if present in the user's query, also extracts the user's intent expressed in rhetorical expressions in the query. For example, if the query is "Does coffee increase blood pressure?", the query normalizer 302 extracts "coffee", "blood pressure", and "increase" as specific rhetorical relations between "coffee" and "blood pressure". Consider another example where the query is "coffee and cancer". In this case, the query normalizer 302 normalizes the query as a compound concept of the user's interest and finds all kinds of rhetorical relations between two simple concepts ("coffee" and "cancer") by default.

[0055] The document aggregator 305 performs queries on the accessible sources for the user 205. In an exemplary embodiment, the document aggregator 305 has access to the Internet 220, the public document repository 228, and the private document repository 225. On the Internet 220, the document aggregator uses a search engine such as Google. The document aggregator 305 passes the results to the document preprocessor 307, which performs a series of preprocessing steps on each document. Examples of such preprocessing actions include removing or cleaning up the document content of any html tags that are not part of the document content, any extraneous information. The document preprocessor 307 provides the clean document to the discourse model determination 300, which examines the document structure and determines its discourse model. The discourse model determination 300 loads the available discourse model structures from the discourse models 290. Common examples of discourse models are emails, articles, blogs, and research reports. The administrator can add new discourse models using the abstract discourse model structures without programming. If the new discourse model requires changes to the abstract discourse model structure, then the Discourse Model Determination 300 component is enhanced. Discourse model structure is described in more detail later in this section.

[0056] Once the NLU platform 200 has determined the discourse model, the document is parsed using the document parser 310. Document parsing takes the discourse model as input and performs format-specific native document parsing to convert the data into a format-independent representation. Example formats include HTML, Docx, PDF, Pptx and eBook. Parsing includes explicit parsing of native elements including headers, footers, footnotes, paragraphs, bulleted lists, navigation links and graphics. After parsing, the document parser 310 organizes the content into a hierarchical graph with sufficient attributes to maintain absolute traceability of all elements.

[0057] Once basic processing is complete, the document is passed through the iterative surface language parser 320, which performs surface NLP tasks to evaluate all machine and linguistic attributes of the bulk text from a hierarchical graph representation of the document: Identify paragraphs, sentences, and clauses, and tokenize sentences by words; Determine sections (subsections), paragraphs (child paragraphs), sentences, words, part-of-speech (POS) tags, topics, phrases, and linguistic attributes (see Wikipedia's Grammar Category page for an expanded taxonomy). Every attribute has a role to play in subsequent steps in the pipeline; these attributes form the building blocks of linguistic rules in subsequent steps. The iterative surface language parser 320 is iterative due to common errors produced by out-of-the-box surface NLP parsers. This is done to prevent some common errors in part-of-speech tags from popular part-of-speech taggers. These rules correlate with the word itself, the words before and after the word, clauses and their attributes, sentence types and sentence attributes. Such rules are stored as metadata and can be enforced by users or administrators without programming.

[0058] The language simplification tool 330 performs further simplification using linguistic and structural rules (at various levels) to improve the accuracy of the parsed output. The language simplification tool 430 also incorporates external knowledge to detect actual concepts (through minor sense disambiguation) and aggregate words to overall simplify sentence length. Mechanical subsentence detection (and separate processing) to further normalize sentence length while preserving information and relationships is also performed by this component. The simplification tool 330 performs mechanical substitutions to present familiar words to the parser and semantic substitutions to present 1-2 words of entire subordinate clauses or phrases to the parser to prevent common errors. Substitutions are identified using KnowledgeNet and linguistic attributes (from surface analysis), parallel boundaries are expanded (e.g., law vs. constitutional rights = legal rights vs. constitutional rights), and subordinate clauses are simplified (e.g., single-family households that own homes = single-family homeowners).

[0059] The iterative deep parser 340 performs deeper parsing of the sentence. The iterative deep parser 340 is iterative to correct certain errors in the parse tree caused by misidentification by the POS tagger. After correction, the process is restarted. The deep parser 440 also determines the concepts and their coreferences. The coreferences may be pronominal (e.g., the use of a pronoun to refer to a noun used before or after in the document) or may be pragmatic, such as using "xerox" to refer to a copy. To allow for deeper coreferences and other rhetorical relationships between concepts, the deep parser 340 also integrates real-world knowledge of additional contexts, e.g., Civil Rights Movement also means "period" when accompanied by a preposition such as "during" or "after". According to Wikipedia, this period is 1954-1968. References (direct or implied) throughout this document also inform us of the fact that Rights is here a noun, rather than a verb or adjective. This component also determines the structure of the sentence, if any, including the modifiers and constraints (if any) within the structure (e.g., Civil Rights Movement during Kennedy Administration).

[0060] The Recursive Semantic Encoder 350 uses all the outputs of the preceding components to create a human-like semantic representation layer. This includes idea detection, rhetorical relations within and between ideas. The Recursive Semantic Encoder 350 also determines important linguistic metrics such as importance of ideas / rhetorical expressions, central themes, and major topics. The semantic encoding process performed by the Recursive Semantic Encoder 350 outputs a combined semantic representation hypergraph (MR) that combines and preserves hierarchies, ideas, arguments / rhetorical expressions, sentences, clauses, and concepts. The MR is fully lossless, and entire documents can be generated from the MR and vice versa. The Semantic Encoder 450 also derives attributes (importance, layer, summary, etc.) to accommodate the needs of more selective analysis and specialized understanding.

[0061] Graphs are widely recognized as an effective way to mirror the way humans organize and view information. A graph-based representation of discourse that preserves compositionality can effectively provide a computer surrogate for the above levels of text understanding. A Semantic Representation Graph [MR] provides a semantic representation of any document or collection of documents that matches the above definition of multi-layered meaning.

[0062] MR graphs should be distinguished from other graph forms in the literature, such as knowledge graphs, concept graphs, and AMR graphs. While there is overlap between all these graph representations, MR graphs represent complete documents in a surface form in terms of compositionality [sentences, phrases, clauses], concepts, and rhetorical connections between concepts. MR maps surface rhetorical phrases to abstract rhetorical relation types, although the reverse translation is possible.

[0063] MR is a multi-layered semantic representation of any document. Semantic is defined as understanding the concepts and relationships between concepts described in a discourse. Naturally, this is supplemented by existing knowledge of concepts and interrelationships in the human mind. A layered representation of semantics in the NLU platform 200 is presented first and then extended to include knowledge outside the document that is relevant to understanding the document.

[0064] Conceptually, a document can be defined as a set of ideas on a subject that an author is trying to communicate to a reader. The ideas are then expressed in the form of sentences, with each sentence expressing the relationship between certain concepts using rhetorical connections. Thus, a simple sentence can be thought of as a subject and an object linked by a verb.

[0065] This idea can be expressed mathematically as follows: (Number 1) Document D={I}, where I is a set of ideas or themes; Where: I={S}, where S is a set of sentences, S={C,R}, where C is a set of concepts and R is a set of rhetorical expressions.

[0066] The NLU platform 200 computationally enumerates the above definitions to generate an MR graph. As previously mentioned, the generated MR graph is multi-layered and can gradually reveal meaning. "Idea," "theme," and "topic" are used interchangeably in the remainder of the disclosure.

[0067] Consider a hypothetical document D that contains many sentences {S1,...Sn}. At the highest level, the NLU platform 200MR is a graph of ideas in the document. Figure 4 illustrates this conceptually. For short documents, the document may reflect only one idea. The longer the document, the higher the number of ideas. However, most documents are likely to have a main theme and several subthemes. In some cases, such as analyst reports, the subthemes may be independent of each other.

[0068] Conceptually, sentences that reflect an idea or theme are grouped together, and each emerging theme should ideally appear as a new set of sentences. Remember that all sentences are related because they come from the same document. However, it is interesting to discover major sub-themes expressed within the documents.

[0069] The NLU platform 200 discovers ideas by delving into the linguistic structure / attributes of a document at the word, clause, sentence, paragraph, section and entire document levels. It analyzes all parts of a document to determine how they relate rhetorically and linguistically to other parts of the document.

[0070] Each idea takes up multiple sentences in the discourse, and the main sentence that introduces the idea becomes the head sentence for that idea. An idea can be expressed through a sentence or the main concept of the main sentence.

[0071] Figure 5 shows the semantic representation of Figure 4 along with an expanded representation of IDEA 1. The expanded representation shows that IDEA 1 includes multiple sentences from the first and second paragraphs in the document. NLU platform 200 uses a comprehensive set of linguistic attributes at the word, clause, sentence, paragraph, section, and whole document levels, as well as discourse model elements to detect the beginning of an idea or theme. An example of a discourse model element is the "new paragraph" marker.

[0072] FIG. 5 shows that idea 1 includes S1, S2, S3 in paragraph P1, S1 in paragraph 2, etc. FIG. 5 also shows rhetorical relations between sentences (e.g., R12). R12 shows the relationship between sentences S1 (P1) and S2 (P1). Examples of such relations are elaboration, temporal, etc. In simple terms, if R12 is identified as "elaboration," then S2 is determined to be a elaboration of S1 (P1). NLU platform 200 has the capability to discover the complete set of such rhetorical relations.

[0073] The MR is then extended to a concept-level representation of the document. The NLU platform 200 concept graph is a complete cognitive representation of the document at the concept level. The NLU platform 200 extends the sentences to reveal their constituent structure in the form of concepts and their relationships. The relationships between sentences are effectively also the relationships between the subjects of the sentences. For example, a sentence beginning with a pronoun refers to a person in the associated sentence. A concept can be a noun or a phrase, and sometimes it is a normalized form of a noun or verb phrase. It is a directed graph of nodes and edges, where the nodes are concepts and the edges are phrases of text that connect one concept to another.

[0074] A Concept Graph (GG) of NLU platform 200 may be formally defined as follows: (Number 2) GG={V,E} Where GG = concept graph, V = the set of vertices in GG {A,B,C,D,E,...}, and E = the set of edges containing surface-level rhetorical phrases; E ij = the edge between the i-th and j-th vertices R = Abstract relationship type of surface-level rhetorical phrase (E)

[0075] Figure 6 shows an example of a directed graph with concept nodes {C1,...C8} and edges E. Each edge in E is ij where i is the source and j is the target. ij is the type of abstraction relation (R ij ) and R ij belongs to a set of abstract relationship types R. Edges E may have weights associated with them, which may indicate their semantic importance in interpreting the overall meaning of the document. The NLU platform 200 uses default weights based on the abstract relationship type R to which the edge is mapped. For clarity, FIG. 6 does not show edge weights (RW).

[0076] FIG. 7 shows four typical discourse models: a news article 710, a research paper 720, an email 730, and a legal contract 740. As shown, the discourse models have a hierarchical structure. Any node in the hierarchy of the discourse models is optional, the [...] node is optional, and the other nodes are typically always present. Structurally, however, an administrator can create any number of discourse models in the discourse model 290, using unique labels to name the discourse models. Similarly, an administrator can create any depth level in the hierarchy. For example, the legal contract DM 740 has a node titled "Clause 1 Test" and a child node titled "[Sub-Clause]". An administrator can create child nodes for sub-clause as needed. The NLU platform 200 attempts to find the closest discourse model to the document.

[0077] Figure 8 shows a sample news article related to immigration enforcement. For ease of explanation and reference, each sentence is labeled S1, S2, S3...Sn.

[0078] Figures 9a-b show the NLU platform's analysis of the results of the 10 queries reported in Figure 1. The search engine results were processed using the NLU platform 200. The output of the NLU platform is shown in Figures 9a-b. Using the NLU platform 200, the detection of related documents is significantly improved. Figure 9a shows the accuracy metrics of the NLU platform 200 versus the search engine. The accuracy of the detection of related items has improved by a significant 29%, from 59% to 88%. The F1 score has also improved from 74% to 89%. The mean squared error [MSE] statistics are also shown based on the ranking of items based on relevance by the NLU platform 200 and the order of items in Google's search results. Figure 9b shows this graphically. It shows the frequency distribution of page differences calculated as [Google page - NLU platform page]. The results are surprising, the NLU platform 200 is much more accurate at detecting relevance.

[0079] FIG. 10 illustrates the process of determining the relevance of a document to a query topic by the NLU platform 200. A concept 1010 and a document 1015 are input to the NLU platform 200, which goes through a three-step process to determine the relevance of the document to a query. After concept detection (occurrence of the concept in the MR graph of the document), for complex concepts (multiple topics), in addition to sentence importance (main idea, set of sentences of the main idea) and concept importance (subject), the NLU platform 200 also determines the context strength. The context strength determines how essential the concepts are to each other in the context of the document. For example, do the words of a complex concept appear in various places in the document without any direct contextual relationship to each other, or do they appear in a direct contextual relationship to each other?

[0080] Figure 11 shows a concept-level representation of the first sentence of the document in Figure 8. The nodes in the diagram represent concepts in the sentence, and the edges reflect the rhetorical relations expressed.

[0081] FIG. 12 shows a summary of the document shown in FIG. 8. The NLU platform 200 found that 7 of the 27 sentences reflect the essence of the document. If the user wants to expand on any of the 7 separate ideas / themes from the document included in the summary of FIG. 12, the NLU platform 200 will expand on that idea using sentences that are part of that idea. The expansion can be filtered by a particular set of rhetorical relations that the user is interested in (e.g., “causal relations”). The summary will be at the level of granularity that the user is interested in. Similarly, the summary can reflect only the types of rhetorical relations that the user is interested in. For example, the user may only be interested in the main idea of ​​the document and not the “details”. The NLU platform 200 also generates summaries for any topics of the document that the user is interested in. For example, the user may be interested in the main topic of the document or any of the other topics in the document. The summaries generated by the NLU platform 200 will vary depending on the user's settings.

[0082] FIG. 13 illustrates an electric vehicle knowledge collection generated by NLU platform 200. The tables of contents (COP26, government, enterprise, SME, individual, and judiciary) are generated by NLU platform 200 based on a thorough understanding of the content of all documents it determines to be relevant. Each table of contents topic has a number of content items that NLU platform 200 has categorized into that topic. A content item may be categorized into multiple table of contents topics. In such cases, it should be noted that the summaries for different topics are different. Upon selecting or clicking on any table of contents topic, NLU platform application 310 displays articles related to that topic. Upon selecting the root topic “Climate Change Action”, NLU platform application 210 displays all articles related to “Climate Change Action”.

[0083] FIG. 14 shows an example of continuous knowledge discovery and updating. It shows that auto-discovery is turned on for a particular node. NLU platform 200 then automatically searches and finds new content items related to that node and all of its child nodes. The newly found content items are processed by NLU platform 200 and integrated into the current "knowledge collection" using normal processing steps. In one embodiment of the present invention, the newly found knowledge can be staged for a super-user or author to review and decide which of the newly found items should be integrated.

[0084] FIG. 15 shows the top topics or tags generated by the NLU platform 200 for an example collection of seven documents. Top topics can be created for the entire content item or individually for different sections of the content item. FIG. 16 is an illustration of the zero-shot learning capabilities of the NLU platform 200. The NLU platform 200 was applied to these seven documents to determine whether it could detect topics within the documents that correctly matched the manually generated tags. The use case was to maximize the discoverability of such documents when customers / prospects searched by keywords. FIG. 16 also shows that the NLU platform found 80% matches without any training data, outperforming contemporary language models by an order of magnitude. Leading language models with no task-specific training could only find 20% of the tags on the same data.

[0085] Figure 16 shows the unified semantic representation graph generated by the NLU platform in more detail. It shows the semantic representations from three content items that are intelligently aggregated by the NLU platform to generate the unified semantic representation graph shown in Figure 17. For simplicity, the figure does not show each of the three semantic representation graphs. Figure 17 also shows the use of global knowledge from the global knowledge store. This is evident when looking at how "Biden's Immigration Policy" relates to "Biden's Executive Order on Public Assistance Rules."

[0086] FIG. 17 illustrates the structure of a global knowledge input that the NLU platform 200 can use for its processing. The format of the input is flexible to reflect relationships between any pair of concepts. The relationship column reflects a rhetorical relationship type from an extensible set of relationship types. Changing the relationship type does not require any programming and is governed by data consistency principles such as isolation control. A directional cue phrase (before or after) can be associated with the relationship. The sense column reflects the sense in which the concept is referenced in the specified relationship. Sense is an extensible attribute. The sense list for any concept (or word) in the NLU platform 200 can be extended without programming. Global knowledge can be specified to be applicable to all knowledge collections, specific knowledge collections or domains, or only to content items. If necessary, the NLU platform 200 can also invoke an automatic knowledge acquisition system or access a global knowledge store that contains global knowledge about the concept.

[0087] As with FIG. 17, FIG. 18 allows a user of the NLU platform application 310 to input sense disambiguation rules. These rules can be a combination of linguistic characteristics of a word, the set of words that appear before and after the word in the sentence that contains it, part-of-speech tags for those words, cue phrases, and content words such as proper nouns and distinct nouns. The NLU platform 200 can also invoke automatic classifiers, if present, to disambiguate the sense. Such automatic classifiers automatically identify how words are used in a particular context. As with global knowledge, the scope indicates how broadly the rule applies (all knowledge collections, a specific knowledge collection or domain, or just the document).

[0088] In another embodiment, an exemplary method and use of automated natural language understanding tasks performed by the NLU platform of the present invention in an agile continuous learning process is described below.

[0089] Agile Continuous Learning Platform (ACL Platform): An effective framework for an agile continuous learning platform includes a process that provides users with the following core element requirements:

[0090] 1. The ability for learners to self-study from a continuously updated knowledge portal, with or without assessments and micro-credentials

[0091] 2. The ability to rapidly assemble, organize, integrate, and keep up to date new learning content from within and outside an institution's content library (rapid learning content assembly).

[0092] 3. Automated, explainable evaluation of free responses scored against a rubric (Scalable open-text response evaluation.

[0093] 4. Personalization of skill-based, career and demand-driven learning journeys.

[0094] The ACL Platform application of the automated NLU platform of the present invention automatically creates a continuous knowledge collection on any topic within its continuous knowledge portal, as also shown in Figures 13-14.

[0095] The ACL platform application of the present invention includes one or more portals that facilitate self-study on a collection of topics. For example, a third party institution can provide a portal to lifelong learners by combining it with assessments and credentials. Instructors or course designers can control the content that learners can access. Instructors can annotate ACL content as they see fit to convey their knowledge and teaching methods.

[0096] The ACL Platform Portal can also be integrated with an institution's learning delivery infrastructure and learning management system.

[0097] Rapid learning content collection (course creation) The ACL Platform I Learning Content Authoring / Assembly component (Figure 19) enables any institution to rapidly assemble learning content around new topics or topics related to specific skills. The ACL Platform uses its natural language understanding engine to break down content down to the concept level. It can incorporate existing learning content in any format. It can also augment the content with new content from external repositories such as other private learning content collections or the public internet.

[0098] A large global educational institution is using Gyan to rapidly break down vast amounts of existing content into granular learning units and supplement them with external content to enable the creation of new learning modules on demand. The institution can combine granular learning units to meet new learning needs. Content can be personalized for individual learners as needed. Personalization can be based on individual learner learning preferences.

[0099] The platform also has the intelligence to build learning overviews, create assessments based on desired coverage, and integrate the final course or learning content with popular learning management solutions (LMS).

[0100] Scalable free-text response evaluation Grading of essays and free text answers [ORA] by humans is time consuming and prone to inconsistencies and biases. Existing research on automated grading of essays [AES] has been criticized for lack of context awareness, need for large amounts of training data, data-driven bias, and their black box nature. Furthermore, peer grading is a common practice in online courses such as MOOCs. Peer grading has significant drawbacks as a grading method, the biggest of which is that it naturally leads to grade inflation. The ACL platform of the present invention effectively addresses these challenges. The ACL platform evaluates essays or ORAs in their context, gleaning intelligence from a sparse dataset of labeled essays, and can be improved in a way that is easy to understand because its reasoning or logic is transparent and easy to handle. Furthermore, the ACL platform can be rapidly configured to reflect course-specific requirements.

[0101] A prominent, globally recognized higher education institution currently uses peer grading for the evaluation of free text responses in online courses. Assignments from already graded courses were selected for evaluation of the ACL platform of the present invention. For the selected assignments, peer graders gave each other perfect marks.

[0102] The ACL platform was configured using the rubrics specified by the professor. After the ACL platform was configured, all the assignments submitted by the 10 students were evaluated through the ACL platform.

[0103] To obtain the output grading provided by the ACL platform and compare it to human standards, we asked three Teaching Assistants (TAs) to manually grade the same essays. Figure 20 shows the results of the pilot. As can be seen, the peer grades of all 10 students were full marks. The TAs had a large variability for the same essays. For each evaluation of the ACL platform, a detailed reasoning report was provided to the professor / institution explaining how the ACL platform arrived at the final score / grade. In all cases, the professor and TAs agreed with the ACL platform's evaluation.

[0104] Personalized Learning Pathways Employers are increasingly eager to hire people who have the skills they need for the job, regardless of whether they have a traditional four-year college degree. Large employers are partnering with educational institutions to create learning programs to develop the talent and workforce they need.

[0105] The ACL platform enables the automatic creation of personalized learning pathways for individuals and is used by global higher education institutions and startups focused on career guidance for unemployed adults.

[0106] The process of determining a personalized learning pathway begins with an individual user uploading and completing their resume or profile. The ACL platform extracts the individual's skills from the resume. The ACL platform processes the different sections of the resume (education, experience, objectives, and skills if explicitly stated). In addition to the skills, the ACL platform also determines the proficiency level of each skill. The ACL platform may be integrated with any skills taxonomy, such as ONET, EMSI Burning Glass, or a proprietary skills taxonomy specific to the hiring company.

[0107] The ACL platform identifies skill gaps between an individual's skills listed on their resume and the desired occupation. The skills extracted from the resume are normalized by the ACL platform using a selected skills taxonomy, such as ONET. The normalized skills are then compared to the desired occupation to identify the individual's skill gaps in that particular occupation.

[0108] Document discoverability tags Improving the discoverability of an institution's content typically requires manually creating tags and / or manually creating a content creation structure (e.g., a hierarchical set of topics). Keeping this tagging process ahead of the curve as new content arrives is equally challenging. Automatically generating high-quality tags would not only make this process significantly more efficient, but would also enable tagging of content at a much larger scale.

[0109] The ACL platform of the present invention was evaluated in a pilot project that required the platform to generate tags that best describe documents in order to maximize discoverability in searches by customers / prospects.

[0110] As shown in Figure 15, for a sample size of 7 articles, the ACL platform matched 80% of the tags without any training data, which was an order of magnitude better than state-of-the-art AI-based machine learning systems (LMs). The leading LMs currently in use without task-specific training only accurately detected 20% of the tags with the same data compared to our ACL platform application.

[0111] In summary, the present invention is directed to a system and method for natural language understanding, the system including a processor configured to automatically perform the task of human-like natural language understanding of natural language content from one or more text documents, the system having a natural language understanding engine that creates a machine representation of the meaning of the document using a constituent language model, the engine creates the machine semantic representation of the document by breaking down the entire document into its constituent structures to reflect multiple layers of meaning, and generates the machine semantic representation in a reversible manner by parsing the document according to a discourse model identified for the document, identifying main ideas of the document, identifying beginnings of new ideas in the document, breaking down the document into sub-documents by ideas, and breaking down the sub-documents into constituent parts to create a semantic representation of the entire document, the computing system may include one or more virtual or dedicated servers, or similar computing devices, and may be programmed with executable instructions.

[0112] The method of the present invention automatically performs human-like natural language understanding of natural language content from one or more text documents, the method includes generating a machine representation of the meaning of the document using a constituent language model, then generating a machine semantic representation of the document by decomposing the entire document into its constituent structures to reflect multiple layers of meaning, and generating the machine semantic representation in a reversible manner by parsing the document according to a discourse model identified for the document, identifying main ideas of the document, identifying beginnings of new ideas in the document, decomposing the document into sub-documents by ideas, and creating a semantic representation of the entire document by decomposing the sub-documents into their constituent structures.

[0113] A compositional language model does not utilize statistical machine learning or statistically derived distributional semantics such as word embeddings to create its semantic representation. A natural language understanding engine utilizes computational linguistics, the compositionality and rhetorical structure of language, and can process one or more documents without any training or training data. A natural language understanding engine works entirely with natural language and does not convert any part of a document into a real-valued vector encoding.

Claims

1. A natural language understanding system comprising at least one processor configured to automatically perform a human-like natural language understanding task of natural language content from one or more text documents, wherein the system This includes a computer system that hosts a natural language understanding engine that uses a constructive language model to create a machine representation of the meaning of a document, The engine decomposes the entire document into its constituent structure and creates a machine semantic representation of the document by reflecting multiple layers of meaning, and creates the machine semantic representation in a reversible manner. The engine analyzes the document according to a discourse model identified for the document, identifies the main ideas of the document, identifies the beginnings of new ideas within the document, decomposes the document into subdocuments by ideas, and decomposes the subdocuments into their constituent elements to create a semantic representation of the entire document, thereby decomposing the document into its constituent structure. The computing system includes one or more virtual servers or dedicated servers, or similar computing devices, and is programmed with executable instructions.

2. The system according to claim 1, wherein the constituent language model does not utilize statistical machine learning or statistically derived distributional semantics such as word embeddings to create its semantic representation graph.

3. The system according to claim 1, wherein the natural language understanding engine utilizes computational linguistics, linguistic constructivity, and rhetorical structure to process one or more documents without any training or training data.

4. The system according to claim 1, wherein the natural language understanding engine operates entirely in natural language and does not convert any part of the document into a real-valued vector encoding.

5. The system according to claim 1, wherein a new discourse model is created and integrated without programming.

6. The system according to claim 1, wherein the main idea of ​​the document is determined on the importance of the concept using a combination of linguistic and rhetorical structures, attributes, and discourse models.

7. The system according to claim 1, wherein the origin of the aforementioned new idea can be determined by determining the nature of the relationships between each consecutive constituent element, including paragraphs, sentences, and clauses, based on the nature of linguistic attributes and rhetorical expressions that link constituent elements, including multiple linguistic rules, rhetorical relationships, and coreference relationships.

8. The system according to claim 1, wherein the decomposition of the subdocument includes, at a first level, identifying sentence-level components constituting the subdocument and the relationships between them using linguistic attributes and the properties of linguistic relationships between the components and the main sentences of the subdocument and between the components themselves.

9. The system according to claim 8, wherein breaking down the subdocument at the next level includes using linguistic attributes and rules and rhetorical attributes and rules to identify the nature of the rhetorical and linguistic relationships within and between sentence-level components.

10. The system according to claim 9, wherein the decomposition is continued until each constituent unit, including word-level units, becomes part of the semantic representation graph.

11. The system according to claim 8, wherein new rules can be added for detecting relationships between components without any programming.

12. The system according to claim 1, wherein the semantic expression takes the form of a multilayer graph comprising the decomposed components which can be queried by a concept, the surface form of the rhetorical expression, or a type of abstract rhetorical relationship.

13. The system according to claim 12, wherein the semantic representation graph can be reversibly returned to a complete document.

14. The system according to claim 1, wherein the decomposition of the constituent structure includes establishing coreference relationships between words, and the detection of such coreference relationships is improved by external knowledge stored in a global knowledge store.

15. The system according to claim 14, wherein the global knowledge store stores knowledge at a sense level for each concept and can be continuously improved without any programming.

16. The system according to claim 14, wherein the global knowledge store can store knowledge at a sense level for each concept.

17. The system according to claim 1, wherein the document may be a text document of any length in any format, including text, HTML, PDF, etc., based on any discourse model, and a video transcribed to text.

18. The system according to claim 1, wherein the document includes grammatically incomplete and incomprehensible social media messages.

19. The system according to claim 1, wherein the engine identifies rhetorical phrases between sentences and between constituent units, including between subjects and objects, within each sentence, and maps such identified rhetorical phrases to a normalized set of rhetorical relationships.

20. The system according to claim 12, wherein the mapping between rhetorical expressions and abstract rhetorical relationships is configurable and can function as cue words, cue phrases, and linguistic attributes of the discourse such as parts of speech, sentence patterns, and coreference relationships.

21. The system according to claim 1, wherein the engine creates an extensible multilevel outline of the document corresponding to the multilevel semantic representation graph.

22. The system according to claim 1, wherein the engine creates an extensible multi-level summary of a document according to user selections expressed as rules that operate based on the rhetorical relationships identified within the document.

23. The system according to claim 1, wherein the engine can process one or more documents without requiring any training or set of training documents.

24. A system capable of creating a knowledge collection of related documents using the semantic representation of the system of claim 1.

25. The system according to claim 24, wherein the knowledge collection is obtained based on queries using an internet search engine, a private document collection, and sources, or a combination thereof.

26. The engine organizes the collection of related documents into a predetermined set of subtopics by determining the most relevant match for each document. If a predetermined set of subtopics is unavailable, the engine uses the semantic representation of the system of claim 1 to surface a unique set of topics from the collection of related documents for each of the documents in the collection, according to claim 25.

27. The system according to claim 26, wherein the specific set of subtopics is determined to be the most representative idea in the entire collection of documents by aggregating the ideas in each of the documents.

28. The system according to claim 27, wherein the process for determining the most representative idea across the entire collection of documents involves integrating knowledge from the global knowledge store so that the engine can recognize relationships between ideas that are not explicitly present in the documents, and using linguistic characteristics that include their role in the documents in which they are found.

29. The system according to claim 26, wherein the degree of relevance between each article in the document collection and each predefined subtopic is determined using a combination of conceptual importance, sentence importance, and contextual strength.

30. The system according to claim 24, wherein the engine removes redundant concepts from the entire semantic representation graph at the document level, then aggregates the semantic representation graphs for each of the documents to create a combined semantic representation graph for the entire collection of documents.

31. The system according to claim 24, wherein the engine aggregates summaries for each of the documents and removes redundant content to create an aggregated summary of the entire collection of documents.

32. The system according to claim 24, wherein the engine creates an aggregated summary for each of the subtopics by aggregating summaries of the documents classified into those subtopics.

33. The system according to claim 1, wherein the aforementioned coreference relationship between concepts in two sentences may be the aforementioned use of the exact same concept, or a pronoun or practical reference to the concept contained in an external global knowledge store.

34. A method for automatically performing human-like natural language understanding of natural language content from one or more text documents, wherein the method is Using a constructive language model to generate a machine representation of the meaning of a document, The entire document is broken down into its constituent structure, and by reflecting multiple layers of meaning, a machine semantic representation for the document is generated, and the machine semantic representation is generated in a reversible manner. A method comprising: analyzing the document according to a discourse model identified for the document; identifying the main ideas of the document; identifying the beginnings of new ideas within the document; decomposing the document into its constituent structure by decomposing the document into subdocuments by ideas; and creating a semantic representation of the entire document by decomposing the subdocuments into their constituent elements.

35. The method according to claim 34, wherein the construct language model does not utilize statistical machine learning or statistically derived distributional semantics such as word embeddings to construct its semantic representation graph.

36. The method according to claim 34, wherein the constructive language model can process one or more documents without any training or training data, utilizing computational linguistics and the rhetorical structures of language.

37. The method according to claim 34, wherein the method operates entirely in natural language and does not convert any part of the document into a real-valued vector encoding.

38. The method according to claim 34, wherein a new discourse model is created and integrated without programming.

39. The method according to claim 34, wherein the main idea of ​​the document is determined on the basis of the importance of the concept using a combination of linguistic and rhetorical structures, attributes, and discourse models.

40. The method according to claim 34, wherein the origin of the aforementioned new idea may be determined by determining the nature of the relationships between each consecutive constituent element, including paragraphs, sentences, and clauses, based on the nature of linguistic attributes and rhetorical expressions that link constituent elements, including multiple linguistic rules, rhetorical relationships, and coreference relationships.

41. The method according to claim 34, wherein the decomposition of the subdocument includes, at a first level, identifying sentence-level components constituting the subdocument and the relationships between them using linguistic attributes and the properties of linguistic relationships between the components and the main sentences of the subdocument and between the components themselves.

42. The method according to claim 41, wherein breaking down the subdocument at the next level includes using linguistic attributes and rules and rhetorical attributes and rules to identify the nature of the rhetorical and linguistic relationships within and between sentence-level components.

43. The method according to claim 42, wherein the decomposition is continued until each constituent unit, including word-level units, becomes part of the semantic representation graph.

44. The method according to claim 41, wherein new rules can be added for detecting relationships between components without any programming.

45. The method according to claim 34, wherein the semantic expression takes the form of a multilayer graph comprising the decomposed components, which can be queried by a concept, the surface form of the rhetorical expression, or a type of abstract rhetorical relationship.

46. The method according to claim 34, wherein the decomposition of the constituent structure includes establishing coreference relationships between words, and the detection of such coreference relationships is improved by external knowledge stored in the global knowledge store.

47. The method according to claim 46, wherein the global knowledge store stores knowledge at a sense level for each concept and can be continuously improved without any programming.

48. The method according to claim 34, wherein the document contains grammatically incomplete and incomprehensible social media messages.

49. The method according to claim 45, wherein the mapping between rhetorical expressions and abstract rhetorical relationships is configurable and can function as cue words, cue phrases, and linguistic attributes of the discourse such as parts of speech, sentence patterns, and coreference relationships.

50. The method according to claim 34, wherein the method creates an extensible multilevel outline of the document corresponding to the multilevel semantic representation graph.

51. The method according to claim 1, wherein it creates an extensible multi-level overview of the document according to user selections expressed as rules that operate based on the rhetorical relationships identified within the document.

52. The system according to claim 1, wherein the engine can process one or more documents without requiring any training or set of training documents.

53. A method for creating a knowledge collection of related documents using the semantic representation of the method of claim 34.

54. The method according to claim 53, wherein the knowledge collection is obtained based on queries using an internet search engine, a private document collection, and sources, or a combination thereof.

55. The method organizes the collection of related documents into a predetermined set of subtopics by determining the most relevant match for each document. The method of claim 54, wherein, if a predetermined set of subtopics is unavailable, the method uses the semantic representation of the method of claim 34 to surface a unique set of topics from the collection of related documents for each of the documents in the collection.

56. The method according to claim 55, wherein the specific set of subtopics is determined to be the most representative idea in the entire collection of documents by aggregating the ideas in each of the documents.

57. The method according to claim 56, wherein the process for determining the most representative idea in the entire collection of documents also integrates knowledge from the global knowledge store so that the method can recognize relationships between ideas that are not explicitly present in the documents.

58. The method according to claim 57, wherein the degree of relevance between each article in the document collection and each predefined subtopic is determined using a combination of conceptual importance, sentence importance and contextual strength.

59. The method according to claim 34, wherein the method removes redundant concepts from the entire semantic representation graph at the document level, and then aggregates the semantic representation graphs for each of the documents to create a combined semantic representation graph for the entire collection of documents.

60. The method according to claim 34, wherein the method creates an aggregated summary of the entire collection of documents by aggregating summaries for each of the documents and removing redundant content.