An automated platform for explainable semantic knowledge graphs for a concept and explainable sense disambiguation
The AI-based platform with KN and SenseNet systems addresses the limitations of current language models by providing explainable, human-like natural language understanding through semantic knowledge graphs and sense disambiguation, enhancing the accuracy and relevance of natural language processing tasks.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GYAN INC
- Filing Date
- 2025-11-01
- Publication Date
- 2026-05-07
AI Technical Summary
Current language models lack full contextual understanding, explainability, and require large task-specific data, leading to inefficiencies in natural language processing tasks, particularly in real-world applications.
An AI-based platform comprising a Knowledge Net (KN) and SenseNet system that uses deep linguistic learning to create semantic knowledge graphs and provide explainable, human-analogous understanding of natural language data, leveraging a two-level rhetorical relationship taxonomy and tree-based classifiers for sense disambiguation.
Enables machines to understand natural language in its complete compositional context with explainability, improving the accuracy and relevance of natural language processing tasks, such as query expansion and document analysis.
Smart Images

Figure US2025053666_07052026_PF_FP_ABST
Abstract
Description
[0001] TITLE OF INVENTION
[0002] AN AUTOMATED PLATFORM FOR EXPLAINABLE SEMANTIC KNOWLEDGE GRAPHS FOR A CONCEPT AND EXPLAINABLE SENSE DISAMBIGUATION
[0003] RELATED APPLICATIONS
[0004] This application claims priority to US provisional applications S.N. 63 / 715597 and S.N. 63 / 715620 both filed on November 3, 2024, the disclosures of which are hereby incorporated by reference as if fully set forth herein.
[0005] TECHNICAL FIELD
[0006] The present invention relates generally to an artificially intelligent (Al) platform for natural language learning, and more specifically, to an Al platform comprising computer programs and systems comprising a knowledge discovery engine that creates a machine representation of the semantic knowledge graph of a concept for explainable, human- analogous understanding of natural language data for the full range of automated natural language understanding tasks wherein the knowledge discovery engine discovering the related concepts uses a variety of linguistic learning methods. The present invention further relates to an Al platform for natural language understanding, computational linguistics, machine representation of word meaning that provides explainable, human-analogous understanding of natural language data for the full range of automated natural language understanding tasks.
[0007] BACKGROUND OF THE INVENTION
[0008] Natural language content makes up a large portion of content both on the public internet and in private content repositories. Understandably, therefore, there is considerable need and interest in the automated machine processing of natural language content. Almost all the technologies available today for machine processing of natural language rely on large neural language models of word patterns, including patterns based on a real-value vector representation of each word. Such approaches are, however, not designed to understand the text in its complete compositional context. Conversely, humans understand the text in its full compositional context.
[0009] Current language models suffer from significant limitations - the tendency to hallucinate, the lack of intellectual property (IP) protection, the co-mingling of private data, the inability to retain compositional context, lack of explainability, and the need for large task-specific data to be combined with the neural language model. Furthermore, the current crop of language models is far from being universal in their ability to ‘learn’ or ‘transfer’ knowledge from one field to the next. These limitations are a major hurdle to the adoption of such models in real world production applications.
[0010] More recently, here has been a trend towards integration of external knowledge into these models to address some of their limitations especially context retention and contextspecific global knowledge. While these enhancements improve the performance of language models, they are still based on word level patterns. They have not addressed the fundamental issues of full contextual understanding of the composition, explainability and the need for large scale task-specific data. In the real world, for most natural language processing or understanding tasks, there is often a paucity of large amounts of data.
[0011] As an exemplary illustration of the need for natural language understanding by machines, consider a common task most knowledge workers perform routinely, namely research. It is estimated that there are roughly one billion knowledge workers in the world today and growing. Examples of knowledge workers include students, teachers, academic researchers, analysts, consultants, scientists, lawyers, marketers and so on. The research or knowledge acquisition work they perform today involves the use of search engines on the internet and the use of search technologies on internal or private document collections.
[0012] While search engines have provided immense value in providing access to information where there was no access before, that is just the starting point for research.
[0013] Because today’s search engines are still unable to fully understand the text inside documents, there are both false positive and false negatives in the result set. Based on the query, irrelevant results can be seen as early as the first page of the displayed search results in a display setting of 10 items on a page and most certainly in subsequent pages. Similarly, highly relevant results can be seen listed in much later like on pages 10-20 of the displayed search results.
[0014] In this common use case, there is a need to expand the query context and potentially the result context to find more relevant results in an explainable way without any of the limitations of neural language models. Word or other forms of embeddings are inadequate. Humans, on the other hand, effortlessly expand the query and result context in their minds to find matches, interpret or analyze.
[0015] Other exemplary use cases which require full understanding of unstructured content include, for example, tagging unstructured content, comparing multiple documents to each other to detect plagiarism or duplicates, and automating the generation of intelligent natural language output. For avoidance of doubt, unstructured content may be in the form of videos or audio data that can be transcribed to text.
[0016] One specific exemplary use case is the need automated machine processing of natural language content. Natural language content makes up a large portion of content both on the public internet and in private content repositories. Understandably therefore, there is considerable need and interest in the automated machine processing of natural language content. Almost all the technologies available today for machine processing of natural language rely on large neural language models of word patterns including patterns based on a real-value vector representation of each word. Such approaches are, however, not designed to understand the text in its complete and full compositional context like humans.
[0017] Current language models learn to predict the probability of a sequence of words. The generated sequence of words is related to a query by the user. The next word with the highest probability is selected as the output by the model. These models are trained on very large amounts of raw text. The two categories of statistical methods have been used for building language models - one using hidden Markov models and the other using neural networks. Both are statistical approaches that rely on large corpus of natural language text (input data) for their prediction. In terms of the input data, the ‘bag of words’ approach has been very popular. However, a bag of words model while appealing because of its simplicity, does not retain any context and different words have different representations each time regardless of how they are used. This has led to the development of a novel technique called word embeddings in an attempt to create similar representations for similar words. The popular saying, “A word is known by the company it keeps’ is the motivation for word embeddings. The proponents of word embeddings believe this is analogous to capturing the semantic meaning of those words and their usage.
[0018] Word embeddings suffer from several significant limitations which primarily have to do with the corpus dependent nature of such embeddings and the inability to handle the different senses in which words are used. Learned word representations based on a corpus to determine similar words, are not equivalent to human analogous machine representation of the meaning of the text. There is, therefore, a need to accurately and understandably determine the sense in which a polysemic word is used in any context, and a need for a method and system to provide a comprehensive, explainable semantic context for various words. SUMMARY OF THE INVENTION
[0019] The present invention provides a system and method for an artificially intelligent (Al) based Knowledge Net (KN) Platform for automatically creating accessible, transparent knowledge stores which reflect a complete spectrum of semantic relationships between concepts, a Concept Semantic Knowledge Graph (CSKG). The Knowledge Net (KN) component of the Knowledge Net Platform extracts the relationships for any concept from various private and public corpora that are made available to it, including both generic corpora such as, for example, Wikipedia, or the Internet and domain specific corpora, such as, for example, PubMed or Arxiv. Relationships are extracted by the Knowledge Net Platform automatically by analyzing the text related to the concept of interest using deep linguistic learning. Extracted information is stored in a transparent, accessible form in the Knowledge Net Store. Applications can obtain a rich set of semantic information on a concept for purposes of their Al models by querying the Knowledge Net Store.
[0020] The present invention further provides a computer system and method comprising an Al based platform for natural language understanding, computational linguistics, machine representation of word meaning that provides explainable, human-analogous understanding of natural language data for the full range of automated natural language understanding tasks (SenseNet Platform). The SenseNet Platform of the invention comprises one or more processors and storage devices that are configured to automatically discover and identify the sense in which a word is used in a given context with full explainability and transparency.
[0021] One or more aspects of the present invention are described herein and in the examples below. The foregoing and other objects, features, and advantages of the invention will be apparent from the following detailed description in conjunction with the accompanying figures describing the invention. Embodiments of the invention are described, by way of examples only, with reference to the attached figures herein.
[0022] BRIEF DESCRIPTION OF THE FIGURES
[0023] Figure 1 shows the high level architecture and components of the Al based GY AN Knowledge Net (GY AN KN) Platform.
[0024] Figure 2 shows a Concept Semantic Knowledge Graph (CSKG) created by the Knowledge Net (KN) Platform.
[0025] Figure 3 shows an illustrative example of the Concept Semantic Knowledge Graph (GSKG). Figure 4 shows the high level architecture and components Al based SenseNet Platform.
[0026] DETAILED DESCRIPTION OF INVENTION
[0027] The present invention provides an artificially intelligent (Al) platform comprising a computer system that provides a novel method of machine representation of natural language to enable human-analogous understanding by machines. In one embodiment, the computer system of the invention comprises one or more processors and storage devices that are configured to provide an artificially intelligent (Al) based Knowledge Net (KN) Platform. The storage devices may be physical servers or a cloud system that access the internet and other locally saved or remotely located information repositories that store and process relevant information identified by the KN Platform to provide the desired output. The present invention also comprises a to an artificially intelligent (Al) based component comprising a fully explainable, Knowledge Net Store of semantic information on concepts created automatically without utilizing statistical machine learning or statistically derived distributional semantics like embeddings. In one aspect, the KN platform can provide a CSKG for a concept of interest, for example, Climate Change. The CSKG contains all concepts related to ‘Climate Change’ reflecting a complete set of rhetorical relations optionally along with the literal phrases that might characterize the relationship. Illustratively, CSKG contains concepts that are synonyms including hypernymy and hyponymy relationships, concepts that cause ‘Climate Change’, concepts that are affected by ‘Climate Change’ and so on. Where specified, the platform also provides extended CSKG expanding it to include CSKGs of related concepts. For example, for Climate Change, the KN Platform can return a CSKG containing Global Warming and also include a CSKG for Global Warming.
[0028] In an exemplary application, the KN Platform of the invention is used for expanding a natural language query in a natural language processing (NLP) application to find a more complete set of relevant documents pertaining to the query. In another exemplary application, the KN platform is used to assess responses to an essay to determine the relevance of the response to the essay context or to detect overlap across two documents. In a yet another application, CSKG can also be used to provide interactive writing assistance and to develop critical thinking skills.
[0029] Given a concept, the KN Platform can access any corpora and discover related concepts and classify the relationship type into one of a set of predefined relationship types. The scope of discovery can be defined and can be just the immediate context like the sentence or broader like the paragraph or the entire document. The KN Platform uses a variety of methods for its discovery, including deep linguistic learning, cue phrases that denote a type of relationship, and also leverages existing knowledge in the Knowledge Net Store. Deep linguistic learning refers to the full compositional analysis of the text related to the scope of discovery, including part of speech tagging, parse trees, dependency trees, semantic roles, etc. The platform can also reflect different senses of word usage for words or concepts which are polysemic.
[0030] In another embodiment of the invention, the KN platform of the invention is GY AN Knowledge Net Platform (GY AN KN), which uses two-level rhetorical relationship taxonomy to analyze, detect expressed rhetorical expressions in language data. Given a set of relations and their definitions, GY AN KN identifies all its instances from the unstructured text. Relationships are defined between two concepts or two events or a concept and an event or vice-versa. GY AN KN also implements a compound concept identifier and concept normalizer that intelligently reduce complex noun-phrases into specific normalized concepts, so that different relations about the same event or concept can be perceived. GY AN KN can discover the relationship types that are described in Example 1 below. Some of the relations GY AN KN is capable of detecting are listed below in Example 1.
[0031] Example 1.
[0032] 1) Attribution: Named Entity A is saying something about Concept B (Example: France stated that it will back Palestine on its non-member observer entity status)
[0033] 2) Causal: Event A causes Event B (Example: the stagnant housing industry got a rare boost last month, as more people bought new homes after the worst winter for sales in almost 50 years)
[0034] 3) Comparison: Event A is compared to Event B (Example: the housing sector continues to lag, whereas other sectors have begun a rebound in earnest)
[0035] 4) Conclusion: Event A is a conclusion of Event B (Example: the inflation rate over the longer run is primarily determined monetary policy and hence the committee has the ability to specify a longer run goal for inflation) 5) Conditional: If Event A occurs, then Event B may occur (Example: if home prices dip again, then consumers may curb their spending)
[0036] 6) Contrast: Event A and Event B have contrasting behaviors.
[0037] 7) Contra-Expectation: Event A occurs even when Event B occurred, which was opposite to the expectations. (Example: the housing market continues to remain low, even though it did get a significant boost in March)
[0038] 8) Elaboration: Event A is an elaboration of Event B (Example: Economists forecast that incomes may also rise)
[0039] 9) Hypernym: Event A is a hypernym of Event B (Example: retailers such as Home Depot, Inc. describe the broader category of home building materials)
[0040] 10) Justification: Concept B is used to justify the event on Concept A.
[0041] 11) Reason: Event A is a reason of Event B (Example: pending home sales are considered a leading indicator because they track contract signings)
[0042] 12) Result: Event A is a result of Event B (Example: this raises incomes in the respective foreign countries thus supporting increased sales)
[0043] 13) Temporal Simultaneous: Event A occurred simultaneously with Event B (E.g.: in Bristol, sales dropped 43.8 percent in April compared with the same month last year, while the median sales price fell 3 percent to $ 225,000)
[0044] 14) Temporal Succession: Event A is succeeded by Event B (E.g. : many markets began a decline, once those tax credits expired in April)
[0045] Figure 1 describes the GY AN Knowledge Net Platform (GY AN KN) engine functional architecture at a high level, showing the major platform components and their configuration and relationships according to a preferred embodiment of the invention. Users 105 interact with Knowledge Net Platform (KN Platform) 100 by inputting text related to the concept of interest at a concept input interface 110. Relationships to the concept are extracted automatically by KN Platform 100 by analyzing the text related to the concept of interest from various repositories such as the Internet 120, Private Document Repository 125 and Public Document Repositories 130 using deep linguistic learning. The extracted information from these repositories is stored in a transparent, accessible form in the Knowledge Net Store component 140 as GY AN Concept Semantic Knowledge Graphs (GYAN SKG).
[0046] Applications can obtain a rich set of semantic information on a concept for purposes of their Al models by querying the Knowledge Net Store.
[0047] Figure 2 provides a schematic illustration of the GYAN Concept Semantic Knowledge Graphs (GYAN CSKG).
[0048] The GYAN CSKG can be formally defined as follows:
[0049] CSKG = {V, E}
[0050] Where GG = Concept Graph;
[0051] V = set of vertices {A, B, C, D, E, . . . } in GG; and
[0052] E = set of edges with surface level rhetorical phrases
[0053] Eij = edges between ithand jthvertices
[0054] R = Abstract relationship types for the surface level rhetorical phrases E.
[0055] Figure 2 shows a GYAN CSKG directed graph 200 with concept nodes {Ci,.. .C7} 210, 220, 230, 240, 250, 260 and 270, respectively, and Edges E. Each Edge in E is denoted by Eij where i is the source and j is the target. Eij is mapped to an abstract relationship type, Rij where Ry belongs to a set of abstract relationship types R. Edges E can have a weight associated with them which can indicate their semantic importance in interpreting the overall meaning of the document. The GYAN CSKG directed graph 200 uses default weights based on the abstract relationship type R to which the edge is mapped to. For clarity, the edge weights RW is not shown in Figure 2. Figure 3 describes an illustrative example of a GY AN Concept Semantic Knowledge
[0056] Graph (GY AN CSKG). Specifically, Figure 3 shows an Electric Vehicles Knowledge Collection produced by the GY AN CSKG Platform 300. The content generated by the Platform 300 is based on its complete understanding of the content of all the documents and information it found relevant, and shows meaning representation from several content items 310 - 400 that are intelligently aggregated by the GY AN CSKG Platform to generate the consolidated GY AN CSKG.
[0057] In another embodiment of the invention, the Al based platform of the invention comprises a computer system and method comprising an artificially intelligent automated platform for explainable sense disambiguation for automatically discovering and identifying the sense in which a word is used in a given context with full explainability and transparency (SenseNet Platform). The SenseNet Platform comprises a sense disambiguation engine (SenseNet Engine) that provides an artificially intelligent, fully explainable system and method for automatically determining word sense for polysemic words.
[0058] SenseNet identifies word sense in two ways - real time using linguistic rules and / or using pre-built classifiers for that word if it exists. Real time linguistic rules relate to the linguistic attributes of the words before and after the word of interest and their semantic roles. These linguistic rules are obviously fully explainable. On the other hand, word sense classifiers are created from a word specific glossary data set created by SenseNet for that word. Given a word, SenseNet identifies a large number of word usage instances and captures the semantic composition surrounding each instance, e.g., the sentence or the paragraph or the document in which the word occurs. SenseNet then performs a linguistic characterization of that context including the use of part of speech tags, parse trees, dependency trees, semantic roles, and uses those attributes along with the actual words and sequences of words in the context to develop tree-based classifiers to identify word sense. SenseNet relies on tree-based classifiers, as opposed to developing neural networks, in order to maintain explainability. In an exemplary application, the SenseNet Platform is used for determining the sense of polysemic words in queries and documents in a natural language processing (NLP) application. Additionally, SenseNet includes a computer user interface that enables the senses for a word to be edited by an expert without any programming and stored in the system.
[0059] Figure 4 describes the SenseNet Platform high level functional architecture and components of the present invention. The SenseNet Platform has two loosely coupled components: SenseNet Engine 500 comprising a natural language understanding component, which determines the sense in which a word is used in a given content, and SenseNet Classifier Builder 540, where the SenseNet engine develops classifiers for specific words using word usage data.
[0060] To invoke the SenseNet Engine at run time, Users 505 provide (i.e. input) the SenseNet Engine with a word of interest 510 along with the context [sentence, paragraph, document]. The SenseNet Engine 500 determines the sense of the word in that context using a combination of linguistic rules including disambiguation rules and word specific classifiers stored in repository 530. The sense determination by SenseNet Engine 500 has significant implications for all downstream processing in Natural Language Processing (NLP) applications.
[0061] Another component of the SenseNet Platform, SenseNet Classifier Builder 540, develops an explainable model of word sense disambiguation for specific words and gather a collection of word usage data from external sources. SenseNet Classifier Builder 540 creates a collection of word usage examples by querying the internet 520 and other private and public corpora 525, analyzes each of these word usage instances and identifies their linguistic properties using disambiguation rules and classifiers generated by the platform that are then stored in the system in Storage component 550. Finally, a classifier model for input word is developed using linguistic and other properties which uniquely classifies each usage instance into a specific word classification. Figure 4 further provides a computer User Interface inquiry using which Users 505 who are experts can modify any of the rules generated by the machine that are stored as Expert Rules component 560.
[0062] As an example, the SenseNet Platform of the invention can be used to determine the word sense for the word ‘rose’ in the sentences - ‘The rose on that plant is so red’; ‘Rose came home to visit’, ‘The markets rose for the third straight day’. The word ‘rose’ is used differently in these three sentences. It is a noun in the first two sentences, and a verb in the third. Further, it is used to denote a flower in the first sentence and denotes a person’s name in the second sentence and an occurrence in the third. SenseNet can classify the word sense for the word ‘rose’ either using its linguistic rules to analyze the words and sentences complemented or preceded by the use of a sense classifier for the word ‘rose’ if it is available.
[0063] The functioning of the SenseNet Platform shown in Figure 4 may be understood by using an exemplary word ‘rose’ as the user input. SenseNet Classifier Builder has a set of senses for the word ‘rose’. This set may be either input manually or rely on external taxonomies, dictionaries and frameworks, such as for example, WordNet. According to WordNet, the word ‘rose’ is reported to have 21 different senses in which it can be used. If an application encountered the word ‘rose’ in a sentence and needed to know the precise sense in which it is being used in that sentence, it can send the word to the SenseNet Engine along with the sentence, paragraph of the entire related text. SenseNet Classifier returns a two part response - first is the role ‘rose’ is playing in that sentence and the second, the sense which it is playing. For example, in the sentence, ‘The markets rose appreciably today, SenseNet Engine will return ‘verb’ as the role and ‘sense #2 as’ the sense. Example 2 below shows the WordNet search input by the user for the word of interest (i.e.
[0064] Word to search for) ‘rose’ and the sentence or paragraph in which it appears. As shown in
[0065] Example 2, the SenseNet Engine of the invention correctly identifies and returns the correct sense for the word based on the context in which it was used.
[0066] Example 1:
[0067] WordNet Search - 3.1
[0068] - WordNet home page - Glossary - Help
[0069] Word to search for:
[0070] Display Options: (Select option to change) Hide Example Sentences Hide Glosses Show Frequency Counts Show Database Locations Show Lexical File Info Show Lexical File Numbers Show Sense Keys Show Sense Numbers Show all Hide all
[0071] Key: "S:" = Show Synset (semantic) relations, "W:" = Show Word (lexical) relations
[0072] Display options for sense: (gloss) "an example sentence"
[0073] Noun
[0074] • S: (n) rose, rosebush (any of many shrubs of the genus Rosa that bear roses)
[0075] • S: (n) blush wine, pink wine, rose, rose wine (pinkish table wine from red grapes whose skins were removed after fermentation began)
[0076] • S: (n) rose, rosiness (a dusty pink color)
[0077] Verb
[0078] • S: (v) rise, lift, arise, move up, go up, come up, uprise (move upward) "The fog lifted"; "The smoke arose from the forest fire"; "The mist uprose from the meadows"
[0079] • S: (v) rise, go up, climb (increase in value or to a higher point) "prices climbed steeply"; "the value of our house rose sharply last year"
[0080] • S: (v) arise, rise, uprise, get up, stand up (rise to one's feet) "The audience got up and applauded"
[0081] • S: (v) rise, lift, rear (rise up) "The building rose before them"
[0082] • S: (v) surface, come up, rise up, rise (come to the surface)
[0083] • S: (v) originate, arise, rise, develop, uprise, spring up, grow (come into existence; take on form or shape) "A new religious movement originated in that country"; "a love that sprang up from friendship" ; "the idea for the book grew out of a short story"; "An interesting phenomenon uprose "
[0084] • S: (v) ascend, move up, rise (move to a better position in life or to a better job) "She ascended from a life of poverty to one of great renown"
[0085] • S: (v) wax, mount, climb, rise (go up or advance) "Sales were climbing after prices were lowered"
[0086] • S: (v) heighten, rise (become more extreme) "The tension heightened"
[0087] • S: (v) get up, turn out, arise, uprise, rise (get up and out of bed) "I get up at 7 A.M. every day"; "They rose early"; "He uprose at night"
[0088] • S: (v) rise, jump, climb up (rise in rank or status) "Her new novel jumped high on the bestseller list"
[0089] • S: (v) rise (become heartened or elated) "Her spirits rose when she heard the good news"
[0090] • S: (v) rise (exert oneself to meet a challenge) "rise to a challenge"; "rise to the occasion" • S: (v) rebel, arise, rise, rise up (take part in a rebellion; renounce a former allegiance)
[0091] • S: (v) rise, prove (increase in volume) "the dough rose slowly in the warm room "
[0092] • S: (v) rise, come up, uprise, ascend (come up, of celestial bodies) "The sun also rises"; "The sun uprising sees the dusk night fled... "; "Jupiter ascends"
[0093] • S: (v) resurrect, rise, uprise (return from the dead) "Christ is risen!"; "The dead are to uprise "
[0094] Adjective
[0095] • S: (adj) rose, roseate, rosaceous (of something having a dusty purplish pink color) "the roseate glow of dawn"
[0096] Although specific embodiments of the invention have been described herein in detail, it should be understood that variations, substitutions and alterations can be made herein, and in some instances, some features of the embodiments may be employed without a corresponding use of other features. Moreover, the scope of the present invention is not intended to be limited to the particular embodiments of the system, method and steps described in the specification. As can be understood, the examples described above and illustrated are intended to be exemplary only.
Claims
WHAT IS CLAIMED IS1. A system comprising at least one processor configured to automatically carry out a task of discovering concept semantic knowledge graphs and creating explainable concept semantic knowledge stores from one or more text documents or compositions, the system comprising: a computer system that hosts a knowledge discovery engine which creates a machine representation of the semantic knowledge graph of a concept using linguistics; a knowledge discovery engine creating its machine meaning representation for a concept semantic knowledge graph by linking the concept with the discovered related concepts using the rhetorical expressions which were found to relate them; and at least one virtual or dedicated server and one or more computing devices that are programmed with executable instructions; wherein the knowledge discovery engine discovering the related concepts uses a variety of linguistic learning methods including relying on part of speech tags, parse trees, dependency trees, semantic roles, cue phrases for various rhetorical relations.
2. The system of claim 1, where the engine does not utilize statistical machine learning or statistically derived distributional semantics like word embeddings, for discovering related concepts.
3. The system of claim 1 where the knowledge net engine utilizes computational linguistics, compositionality and rhetorical structure in language and can process one or more text passages or compositions without any training or training data.
4. The system of claim 1 where the natural language understanding engine works entirely in natural language and does not convert any part of the document into real valued vector encodings.
5. The system of claim 1 where external vocabulary, dictionaries, taxonomies and similar documents can be ingested into the knowledge store to assist the discovery engine and also be part of the concept semantic knowledge graph.
6. The system of claim 1 where a plurality of linguistic rules or cue phrases can be added to detect relationships without any programming.
7. The system of claim 1 where the relationship detected is represented both in its surface form or abstract rhetorical relationship types.
8. The system of claim 1 where the decomposing the compositional structure includes establishing co-referential relationships between words and the detection of such co- referential relationships are improved by external knowledge stored in a Knowledge Net Store.
9. The system of claim 8, wherein the Knowledge Net Store stores knowledge at a sense level for each concept and continuously improves said concept without any programming.
10. The system of claim 8, wherein the Knowledge Net Store stores knowledge integrates external knowledge with context retention and context-specific global knowledge.
11. The system of claim 7, wherein the map between rhetorical expressions and abstract rhetorical relations are configurable and can be a function of cue words, cue phrases and linguistic attributes of the discourse including parts of speech, sentence type, and co- referential relationships.
12. A method for automatically carrying out a task of discovering concept semantic knowledge graphs and creating explainable concept semantic knowledge stores from one or more text documents or compositions, the method comprising: a computer system hosting a knowledge discovery engine that creates a machine representation of the semantic knowledge graph of a concept using linguistics; a knowledge discovery engine creating its machine meaning representation for a concept semantic knowledge graph by linking the concept with the discovered related concepts using the rhetorical expressions which were found to relate them; and at least one virtual or dedicated server and one or more computing devices that are programmed with executable instructions; wherein the knowledge discovery engine discovers the related concepts using linguistic learning methods comprising reliance on part of speech tags, parse trees, dependency trees, semantic roles and cue phrases for various rhetorical relations.
13. A system comprising at least one processor configured to automatically carry out a task of identifying the word sense for a word in one or more text documents or compositions, the system comprising: a computer system comprising a natural language understanding engine , a sense disambiguation engine and a sense classifier builder that determines the sense in which a word used in a provided context using linguistic rules wherein the sensedisambiguation engine supplements its use of linguistic rules using explainable classification models using stored linguistic rules based on a combination of linguistic learning methods and metrics including part of speech tags, parse trees, dependency trees, semantic roles, word sequences and cue phrases for various rhetorical relations; wherein the system comprises one or more virtual servers, dedicated servers and computing devices that are programmed with executable instructions.
14. The system of claim 13 where the sense disambiguation engine does not utilize statistical machine learning or statistically derived distributional semantics or word embeddings for discovering related concepts.
15. The system of claim 13 where the sense disambiguation engine utilizes computational linguistics, compositionality and rhetorical structure in language.
16. The system of claim 13 where the natural language understanding engine works entirely in natural language and does not convert any part of the context into real valued vector encodings.
17. The system of claim 13 where the sense disambiguation engine is assisted by storage of natural language from external sources comprising vocabulary dictionaries and taxonomies.
18. The system of claim 13 where a plurality of linguistic rules or cue phrases can be added to detect word sense without any programming.
19. The system of claim 13 where the senses for a word are ingested from external taxonomies, dictionaries and frameworks.
20. A system comprising at least one processor configured to develop an explainable classifier for word sense for a word, the system comprising: a computer system comprising a computing device or at least one virtual or dedicated server that is programmed with executable instructions wherein the computer system hosting a sense disambiguation classification model engine which creates an optimized model to identify the sense in which a word is used in a provided context; wherein the disambiguation classification model engine relies on explainable classification models including tree-based classifiers or Markov chains to develop an optimized classifier for a word.
21. The system of claim 20 where the system is capable of discovering a target number of instances of usage of a word from searching for it in any available corpora including the internet.
22. The system of claim 20 where the identified word usage instances are stored along with a defined context for that usage including the sentence, paragraph or the entire document.
23. The system of claims 20 where the senses for a word are ingested from external taxonomies, dictionaries and frameworks.
24. A method for automatically carrying out a task of identifying the word sense for a word in one or more text documents or compositions, the method comprising: a computer system comprising one or more virtual or dedicated servers or computing devices programmed with executable instructions; and a sense disambiguation engine hosted by the computer system that determines the sense in which a word used in a provided context using linguistic rules;wherein the sense disambiguation engine uses linguistic rules based on a combination of linguistic learning methods and metrics including part of speech tags, parse trees, dependency trees, semantic roles, word sequences and cue phrases for various rhetorical relations and supplements its use of linguistic rules with the use of explainable classification models.
Citation Information
Patent Citations
Mapping natural language utterances to operations over a knowledge graph
US20210326531A1
Knowledge graph-based case retrieval method, device and equipment, and storage medium
US20220121695A1
Generation of optimized knowledge-based language model through knowledge graph multi-alignment
US20220230625A1
Dynamic question generation for information-gathering
US20230061906A1