Word Embedding Model for Chemical Substance Representation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current natural language processing technologies face challenges in effectively extracting structured knowledge from texts related to chemical substances, particularly in representing and searching for substances based on their structural, compositional, and physical properties.
Innovation Solution
A word embedding method and apparatus that trains a word embedding model using characteristic information such as structure, composition, and physical properties of chemical substances, enabling the prediction of context words and retrieval of substances with similar characteristics through a word embedding matrix.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If natural language processing technology is used to extract structured knowledge from chemical substance texts, then knowledge extraction capability is improved, but the ability to represent and search substances based on their structural, compositional, and physical properties deteriorates
Solution Approach 1:
The patent segments the substance representation into multiple characteristic dimensions including structure information (molecular graphs, SMILES), composition information (chemical formulas, elemental composition), and physical property information (melting point, boiling point, density). This segmentation allows each dimension to be processed and embedded independently, preserving the specific characteristics of chemical substances while enabling comprehensive knowledge extraction.
Solution Approach 2:
The patent transforms chemical substance data from traditional tabular or text formats into multi-dimensional embedding vectors that capture structural, compositional, and physical properties simultaneously. By mapping substances into a high-dimensional vector space where similar substances are positioned closer together, the system enables efficient similarity search and knowledge extraction while maintaining accurate substance representation.
2Speed
If traditional word embedding models are used for chemical substances, then general text processing speed is improved, but the accuracy of predicting context words specific to chemical domains deteriorates
Solution Approach 1:
The patent performs preliminary action by pre-processing chemical substance data into structured formats (molecular graphs, SMILES strings, chemical formulas) before embedding. This pre-processing step organizes the complex chemical information into standardized representations that the embedding model can efficiently process, maintaining high processing speed while ensuring domain-specific accuracy in context word prediction.
Solution Approach 2:
The patent changes the input parameters of the word embedding model to accommodate chemical substance characteristics. Instead of using only text sequences, the model accepts multiple parameter types including structural parameters (molecular connectivity), compositional parameters (elemental ratios), and physical property parameters. This parameter expansion enables accurate context word prediction for chemical domains while maintaining computational efficiency.
Data Source
AI summary
A word embedding method and apparatus and a word search method are provided, wherein the word embedding method includes training a word embedding model based on characteristic information of a chemical substance, and acquiring an embedding vector of a word representing the chemical substance from the word embedding model, wherein the word embedding model is configured to predict a context word of an input word.


