Chemical Formula Extrapolation for Novel Compound Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems are inadequate in identifying overlapping subject matter and generating new chemical entities within a chemical database, as they fail to visualize relationships between chemical structures and do not predict additional compounds in a low-dimensional space effectively.
Innovation Solution
A computer-implemented method that generates queries representing generic chemical formulas, converts chemical identifiers into coded forms, and uses a virtual n-dimensional manifold to identify and predict new chemical entities by analyzing relationships between coded forms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If generic chemical formulas are used in text documents to represent multiple actual chemical formulas, then the document can cover a broader range of chemical compounds, but it becomes difficult to identify specific overlapping subject matter between documents
Solution Approach 1:
The patent introduces an intermediary system that includes: (1) extracting generic chemical formulas from source documents, (2) converting them to coded forms (e.g., SMILES, InChI), (3) generating specific chemical formulas by substituting variables with values from a database, and (4) comparing these specific formulas across documents to identify overlaps. This intermediary process bridges the gap between generic representation and specific identification.
Solution Approach 2:
The patent segments the generic chemical formula into variable components and invariant components. By extracting variables (e.g., R groups, substituents) and their possible values, the system can generate multiple specific formulas from a single generic formula, enabling detailed comparison while preserving the original broad coverage.
2Measurement precision
If manual techniques are used to identify structural formulas from source documents, then specific chemical compounds can be identified, but the process becomes time-consuming and inefficient
Solution Approach 1:
The patent replaces manual mechanical processes (hand-drawing and comparing structural formulas) with automated computational processes. The system uses software to extract generic formulas, convert them to coded representations, generate specific formulas through variable substitution, and automatically compare them across documents, dramatically improving productivity while maintaining precision.
Solution Approach 2:
The patent changes the representation parameter of chemical formulas from visual structural drawings to coded forms (SMILES, InChI, IUPAC names). This parameter change enables automated processing and comparison while preserving the ability to identify specific compounds, resolving the contradiction between precision and productivity.
3Reliability
If existing analytical techniques are used to search for specific structural formulas in databases, then matching compounds can be identified, but the system cannot predict new chemical entities or visualize relationships in a comprehensive manner
Solution Approach 1:
The patent adds a new dimension to chemical formula analysis by organizing formulas in an n-dimensional space where each dimension represents a chemical feature or property. This enables the system to not only identify existing matching compounds but also predict new chemical entities by exploring unoccupied regions in this multidimensional space and visualizing relationships between compounds in a comprehensive manner.
Solution Approach 2:
The patent performs preliminary actions by generating all possible specific formulas from generic formulas before comparison, and by pre-organizing chemical data in coded forms and n-dimensional space. This preliminary preparation enables both reliable identification of matches and prediction of new entities, as the system has already explored the chemical space and identified patterns.
Data Source
AI summary
A system and method for extrapolating a set of specific representational identifiers that are represented or covered by a generic representational identifier found in a target document. Queries are constructed and performed on a corpus of source documents in which members of the extrapolated set of specific representational identifiers are compared to a database of representational data. By matching representational data in this way, any overlap between the generic representational data and specific instances of the generic representational identifier within the source documents is determined. In a more specific implementation, the system and method reduces the scope of the generic representational identifier such that the reduced scope generic representational identifier encompasses only novel specific representational identifiers.


