Visual Query Builder for Genomic Data Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional genomic data analysis models require downloading large datasets, which is impractical due to the enormous growth in biomedical data sizes, and querying such data using existing visual query browsers is not intuitive for users lacking expertise in RDF formats and SPARQL queries.
Innovation Solution
A visual query builder generates RDF queries by receiving entity identifications from a graphical user interface, allowing users to select entities and relationships, and automatically constructs SPARQL queries to retrieve genomic data sets, simplifying the querying process through a user-friendly interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If researchers download large genomic data sets for local analysis, then they can perform computational analyses using local hardware, but the download time and storage requirements become impractical given the enormous growth in data size
Solution Approach 1:
The patent introduces a cloud-based intermediary system that acts as a mediator between researchers and genomic data stores. Instead of downloading data locally, researchers interact with a web-based interface that translates their queries into SPARQL queries executed against cloud-hosted RDF data stores. This intermediary cloud infrastructure eliminates the need for local data downloads while maintaining research capabilities.
Solution Approach 2:
The patent replaces the mechanical process of downloading and storing physical data files with a digital query-based retrieval system. Instead of transferring terabytes of data over networks to local storage, the system uses web-based interfaces and SPARQL queries to retrieve only the specific data needed, substituting physical data movement with intelligent data access.
2Ease of operation
If visual query browsers are used to query genomic data, then data selection becomes more intuitive, but users still require expertise in RDF formats and SPARQL queries
Solution Approach 1:
The patent introduces a web-based interface as an intermediary layer between the user and the complex RDF/SPARQL query system. This interface allows users to construct queries using familiar web technologies and visual elements, while automatically translating user intent into proper SPARQL queries against the RDF data store, shielding users from underlying complexity.
Solution Approach 2:
The patent creates a simplified copy or abstraction of the complex query interface. Instead of requiring users to directly interact with RDF triples and SPARQL syntax, the system provides a web-based representation that mirrors the data structure in a more accessible format, allowing users to work with familiar web interface patterns rather than specialized query languages.
3Adaptability or versatility
If cloud computing resources are implemented to provide shared data access, then data accessibility improves, but new issues arise regarding data access, computing capacity, interoperability, training, usability, and governance
Solution Approach 1:
The patent introduces a web-based interface as an intermediary layer between the user and the complex RDF/SPARQL query system. This interface allows users to construct queries using familiar web technologies and visual elements, while automatically translating user intent into proper SPARQL queries against the RDF data store, shielding users from underlying complexity.
Solution Approach 2:
The patent creates a universal access layer that works across different platforms and devices through standard web technologies. The system provides multiple functions including data querying, visualization, and analysis through a single web interface, eliminating the need for specialized software installations and ensuring broad compatibility across different user environments.
Data Source
AI summary
A method for generating a query of a genomic data store includes receiving, by a query generator executing on a computing device, from a graphical user interface, an identification of a first entity of a first entity class for inclusion in a resource description framework (RDF) query. The method includes receiving from the graphical user interface, an identification of a second entity of the first entity class, the second entity having a bi-directional relationship with the first entity. The method includes automatically generating an RDF query based upon the received identification of the first entity and the received identification of the second entity. The method includes executing the RDF query to select, from a plurality of genomic data sets, at least one genomic data set for at least one patient cohort. The method includes providing a listing of genomic data sets resulting from executing the RDF query.


