Machine-Learning Knowledge Graph Generation for Cross-Domain Inquiry
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The fragmentation of scientific knowledge into siloed domains hinders interdisciplinary research, making it difficult for researchers to find innovative solutions to complex problems due to the lack of broad expertise across multiple fields.
Innovation Solution
A machine learning-based tool, HypoFinder, automates the initial phases of scientific inquiry by mining semantic metainformation from papers, selecting appropriate scientific contexts and formalisms, and generating hypotheses using large language models to facilitate interdisciplinary research.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If researchers specialize deeply in their domains to achieve proficiency, then their expertise and reliability in their specific field improve, but their ability to traverse and understand multiple domains deteriorates
Solution Approach 1:
The patent introduces an AI assistant as an intermediary tool that bridges the gap between specialized researchers and cross-domain knowledge. The assistant automatically retrieves and synthesizes information from multiple scientific domains, allowing researchers to maintain their specialized expertise while accessing broad interdisciplinary context without needing to personally traverse multiple domains.
Solution Approach 2:
The system creates a virtual copy of broad scientific knowledge through the AI assistant, which encapsulates information from numerous domains. This digital replica provides researchers with access to cross-domain insights without requiring them to personally acquire such extensive knowledge, effectively copying the function of a polymath for the specialized researcher.
2Ease of manufacture
If scientific knowledge is organized into siloed domains for clarity and structure, then the organization and accessibility within domains improve, but the potential for innovative discoveries through cross-domain connections deteriorates
Solution Approach 1:
The AI assistant serves as a universal tool that operates across multiple scientific domains simultaneously. It can retrieve, synthesize, and connect information from different fields (e.g., biology, chemistry, physics) in response to research questions, enabling innovative cross-domain discoveries while preserving the structured organization of individual domains.
Solution Approach 2:
The system introduces an intermediary layer between the siloed domain structures and the researcher. This AI-mediated layer automatically bridges the silos by retrieving and synthesizing information across domains, allowing the structured organization of individual fields to be maintained while enabling innovative cross-domain connections through automated synthesis.
3Measurement precision
If researchers manually search across multiple domains for interdisciplinary connections, then the depth and quality of their analysis improve, but the time required for literature review and hypothesis generation increases
Solution Approach 1:
The AI assistant performs self-service by automatically retrieving, filtering, and synthesizing information from multiple domains in response to research questions. This eliminates the need for researchers to manually search across domains, significantly reducing literature review time while maintaining high analysis quality through the assistant's sophisticated information processing capabilities.
Solution Approach 2:
The system performs preliminary actions by pre-processing and organizing vast amounts of scientific literature into structured formats that the AI assistant can efficiently query. This preliminary organization of knowledge enables rapid, high-quality analysis during actual research without requiring time-consuming manual searches, as the groundwork is already laid by automated systems.
Data Source
AI summary
A system receives a plurality of scientific documents for inclusion in an extended knowledge graph. For each respective scientific document of the plurality of scientific documents, the system classifies the respective scientific document with a theoretical framework of a plurality of theoretical frameworks each comprising terms and principles associated with a particular scientific topic, extracts metainformation of the respective scientific document, structures the metainformation in a document-specific ontology model that further comprises an indication of the theoretical framework, generates a plurality of text chunks from the respective scientific document of a given size, and generates, using a first ML model, one or more concepts from each of the plurality of text chunks. The system generates, using a second ML model, the extended knowledge graph using each of the plurality of text chunks, each concept, and the metainformation, and stores the extended knowledge graph in a graph document database.


