SMILES Text Preprocessing for Small Molecule Data Standardization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional methods for cleaning and standardizing small molecule compounds using chemical informatics are inefficient and lack unified standards, particularly in handling SMILES compound information from various sources, leading to issues with data duplication and non-standard structures, which hinder practical applications in machine learning and deep learning.

Innovation Solution

A data preprocessing method that includes text preprocessing and chemical graph formatting steps, where SMILES text is normalized by removing heavy metal components, multimers, and charges, and split into digitized graph structures for standardized chemical information, facilitating the construction of artificial intelligence models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional chemical informatics methods are used for standardizing small molecule compounds, then data cleaning can be performed, but the processing efficiency is low and computing speed is slow

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidcomputing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent changes the fundamental parameters of data representation by converting SMILES strings into graph structure representations with node and edge features. This parameter transformation enables the use of graph neural networks which process chemical structures more efficiently than traditional chemical informatics methods, directly resolving the contradiction between processing efficiency and computing time.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical text-based chemical informatics processing with a graph-based computational system. By substituting the mechanical string manipulation approach with graph structure analysis and neural network processing, the system achieves significantly improved processing efficiency and reduced computing time while maintaining accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If traditional cleaning methods are used, then some data standardization is achieved, but unified standards are not established and data duplication cannot be effectively distinguished

Engineering Contradiction:
Improvedata standardization precisionVSAvoiddata consistency
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent creates a universal graph-based representation system that can handle diverse chemical structures from multiple sources (SMILES, molecular graphs, 3D structures) through a unified conversion process. This multi-functional approach establishes consistent standards across different data formats and sources, enabling reliable duplication detection and ensuring data consistency throughout the processing pipeline.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent segments chemical structures into discrete graph components (nodes representing atoms, edges representing bonds) with standardized features. This segmentation approach creates a uniform granular representation that enables precise comparison and standardization across different chemical compounds, resolving the issue of inconsistent data standards while maintaining high precision in data cleaning.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If SMILES text is directly used for machine learning applications, then the process is simple, but non-standard or repetitive structures cannot be effectively handled

Engineering Contradiction:
Improveprocessing simplicityVSAvoidstructure standardization accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent introduces graph structure representation as an intermediary between raw SMILES text and machine learning models. This intermediary transformation layer converts variable-length SMILES strings into standardized graph formats with consistent node and edge features, maintaining processing simplicity while dramatically improving structure standardization accuracy and enabling effective handling of non-standard and repetitive structures.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240021276A1Data preprocessing system for cleaning small molecule compound and method thereof
Publication Date: 2024.01.18 AINNOCENCE LLC
  • US20240021276A1 patent drawing
  • US20240021276A1 patent drawing
  • US20240021276A1 patent drawing

AI summary

The present invention provides a data preprocessing method for cleaning a small molecule compound, the data preprocessing method comprising: an S1 text preprocessing step including: preprocessing an original SMILES text of a small molecule compound into a standardized SMILES text of the small molecule compound; and an S2 chemical graph formatting step including: splitting in a format each text element of the standardized SMILES text of the small molecule compound of S1 to obtain chemical graph information of the small molecule compound. The present invention also provides a data preprocessing system for cleaning a small molecule compound. The present invention enables the cleaning, deduplication, and standardization of global datasets, providing an efficient, fast, accurate integration method for the cleaning of end-to-end small molecule compounds.