Scientific Concept Framework for Provenance-Rich Research Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional approaches to organizing and disseminating scientific research, particularly in life sciences, rely on semantic triples that oversimplify complex scientific concepts, reducing informational value and eroding trust, and fail to adequately document provenance, hindering further investigation and reproduction of research findings.
Innovation Solution
A framework that defines concepts as building blocks, incorporating multiple entities, relationships, and qualifications, encapsulating underlying research data and contextual information, and implementing independent compute environments for each concept to enhance data analysis and dissemination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If semantic triples are used to organize scientific research data, then data organization simplicity is improved, but data fidelity and information completeness deteriorate
Solution Approach 1:
The patent segments scientific knowledge into distinct components: concepts (representing scientific assertions), evidence (supporting data), and provenance (source information). This segmentation allows each component to be independently captured and managed, preserving data fidelity while maintaining organizational simplicity through modular structure.
Solution Approach 2:
The patent implements a nested structure where concepts contain evidence, which in turn contain provenance information. This nested doll approach allows complex scientific information to be organized hierarchically, with simple access points at the concept level while preserving detailed information at deeper levels, thus balancing simplicity and completeness.
2Reliability
If complex scientific concepts are fully captured with all nuances and qualifications, then data fidelity is improved, but data organization complexity increases
Solution Approach 1:
The patent extracts complex qualifications and nuances from scientific statements and encapsulates them within structured concept objects. By taking out detailed information and placing it in standardized fields within the concept structure, the system maintains data fidelity while presenting a simplified interface for common operations.
Solution Approach 2:
The patent transforms unstructured scientific text into structured parameters with defined types and validation rules. By changing the representation from free-text to parameterized data structures, the system preserves complex information while enabling efficient querying and manipulation through standardized interfaces.
3Loss of information
If research data is made more accessible through streamlined querying, then information accessibility is improved, but research reproduction capability deteriorates
Solution Approach 1:
The patent introduces concept objects as intermediaries between raw research data and query interfaces. These concept objects serve as mediators that maintain complete provenance information and evidence links while providing simplified access points for querying, thus enabling both accessibility and reproducibility.
Solution Approach 2:
The patent performs preliminary organization of research data into structured concepts with embedded provenance and evidence information before querying occurs. This preliminary structuring enables efficient access while preserving all necessary information for reproduction, as the complete data lineage is established in advance.
Data Source
AI summary
Computing systems methods, and non-transitory storage media are provided for ingesting data, which includes entities, within a data platform, formulating concepts associated with a subset of the entities, defining the concepts as building blocks within a framework of the data platform, categorizing the data within the concepts, and linking the concepts with one another and with the subset of the entities. The concepts include relationships among the subset of the entities.


