Word Embedding Model for Chemical Substance Representation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current natural language processing technologies face challenges in effectively extracting structured knowledge from texts related to chemical substances, particularly in representing and searching for substances based on their structural, compositional, and physical properties.

Innovation Solution

A word embedding method and apparatus that trains a word embedding model using characteristic information such as structure, composition, and physical properties of chemical substances, enabling the prediction of context words and retrieval of substances with similar characteristics through a word embedding matrix.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If natural language processing technology is used to extract structured knowledge from chemical substance texts, then knowledge extraction capability is improved, but the ability to represent and search substances based on their structural, compositional, and physical properties deteriorates

Engineering Contradiction:
Improveknowledge extraction capabilityVSAvoidsubstance representation accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the substance representation into multiple characteristic dimensions including structure information (molecular graphs, SMILES), composition information (chemical formulas, elemental composition), and physical property information (melting point, boiling point, density). This segmentation allows each dimension to be processed and embedded independently, preserving the specific characteristics of chemical substances while enabling comprehensive knowledge extraction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms chemical substance data from traditional tabular or text formats into multi-dimensional embedding vectors that capture structural, compositional, and physical properties simultaneously. By mapping substances into a high-dimensional vector space where similar substances are positioned closer together, the system enables efficient similarity search and knowledge extraction while maintaining accurate substance representation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Speed

If traditional word embedding models are used for chemical substances, then general text processing speed is improved, but the accuracy of predicting context words specific to chemical domains deteriorates

Engineering Contradiction:
Improvetext processing speedVSAvoidcontext word prediction accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by pre-processing chemical substance data into structured formats (molecular graphs, SMILES strings, chemical formulas) before embedding. This pre-processing step organizes the complex chemical information into standardized representations that the embedding model can efficiently process, maintaining high processing speed while ensuring domain-specific accuracy in context word prediction.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the input parameters of the word embedding model to accommodate chemical substance characteristics. Instead of using only text sequences, the model accepts multiple parameter types including structural parameters (molecular connectivity), compositional parameters (elemental ratios), and physical property parameters. This parameter expansion enables accurate context word prediction for chemical domains while maintaining computational efficiency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11443118B2Word embedding method and apparatus, and word search method
Publication Date: 2022.09.13 SAMSUNG ELECTRONICS CO LTD
  • US11443118B2 patent drawing
  • US11443118B2 patent drawing
  • US11443118B2 patent drawing

AI summary

A word embedding method and apparatus and a word search method are provided, wherein the word embedding method includes training a word embedding model based on characteristic information of a chemical substance, and acquiring an embedding vector of a word representing the chemical substance from the word embedding model, wherein the word embedding model is configured to predict a context word of an input word.