Graph-Based Formulation Data Extraction System
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The chemical industry faces inefficiencies in formulating new products due to reliance on manual data entry and file searches within vast textual datasets, making it time-consuming and costly for experts to find and recreate similar formulations, as existing techniques lack automated methods for extracting and querying domain-specific information from textual data.
Innovation Solution
A processor-implemented method and system that extracts information from chemical formulation texts, classifies sentences into subject-verb-object triples, and represents recipes as graphs, enabling structured querying and storage of formulations in a graph database, allowing for efficient retrieval and analysis of ingredients, actions, and conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If manual data entry and file search methods are used to extract formulation information from textual data, then experts can find and analyze formulation recipes, but the process becomes time-consuming and costly
Solution Approach 1:
The patent replaces manual mechanical processes (data entry, file searching, expert analysis) with an automated computer-implemented system that uses natural language processing, machine learning, and text mining algorithms to extract formulation information from unstructured textual data sources, thereby eliminating time-consuming manual operations
Solution Approach 2:
The system enables self-service by automatically extracting formulation information from various text sources without requiring expert intervention for data collection and initial analysis, allowing the system to autonomously process and structure formulation data from unstructured sources
2Reliability
If experts manually search through vast textual datasets to find similar formulations, then they can make rational judgments on ingredients and procedures, but the process extends time frames and increases costs
Solution Approach 1:
The patent replaces expert manual searching and analysis with an automated computer-implemented system that uses natural language processing and machine learning to rapidly search through vast textual datasets, extract relevant formulation information, and present structured data for analysis, thereby maintaining reliability while dramatically improving productivity
Solution Approach 2:
The system introduces an intermediary layer between the vast textual datasets and the expert analysis process, using automated information extraction and text mining technologies to bridge the gap between unstructured text and structured formulation data, enabling faster access to relevant information without losing analytical quality
3Ease of operation
If existing techniques rely on file search or crude data search in manually created databases, then formulation data can be accessed, but the techniques lack automated methods for extracting and querying domain-specific information
Solution Approach 1:
The patent replaces crude manual data search methods with automated natural language processing and text mining systems that can extract domain-specific formulation information from unstructured textual data, automatically identifying ingredients, quantities, procedures, and properties without requiring manual database creation or simple file searching
Data Source
AI summary
This disclosure relates to method of extracting an information associated with design of formulated products and representing as a graph. A graph domain model of a plurality of vertices, and at least one formulation text as text file are received as an input. The information extraction is applied to identify at least one sentence and extract at least one subject-verb-object triple from every sentence of the at least one formulation text. A sentence including an ingredient listing and associated weights indicated by presence of weight numerals, and a sentence including at least one verb from the at least one subject-verb-object based on the graph domain model are classified. A representation of the recipe text is generated in terms of at least one action, ingredients on which the at least one action is performed, and condition. An insert query string is generated and executed to store the formulations as the graph.


