SMILES Text Preprocessing for Small Molecule Data Standardization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional methods for cleaning and standardizing small molecule compounds using chemical informatics are inefficient and lack unified standards, particularly in handling SMILES compound information from various sources, leading to issues with data duplication and non-standard structures, which hinder practical applications in machine learning and deep learning.
Innovation Solution
A data preprocessing method that includes text preprocessing and chemical graph formatting steps, where SMILES text is normalized by removing heavy metal components, multimers, and charges, and split into digitized graph structures for standardized chemical information, facilitating the construction of artificial intelligence models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional chemical informatics methods are used for standardizing small molecule compounds, then data cleaning can be performed, but the processing efficiency is low and computing speed is slow
Solution Approach 1:
The patent changes the fundamental parameters of data representation by converting SMILES strings into graph structure representations with node and edge features. This parameter transformation enables the use of graph neural networks which process chemical structures more efficiently than traditional chemical informatics methods, directly resolving the contradiction between processing efficiency and computing time.
Solution Approach 2:
The patent replaces traditional mechanical text-based chemical informatics processing with a graph-based computational system. By substituting the mechanical string manipulation approach with graph structure analysis and neural network processing, the system achieves significantly improved processing efficiency and reduced computing time while maintaining accuracy.
2Manufacturing precision
If traditional cleaning methods are used, then some data standardization is achieved, but unified standards are not established and data duplication cannot be effectively distinguished
Solution Approach 1:
The patent creates a universal graph-based representation system that can handle diverse chemical structures from multiple sources (SMILES, molecular graphs, 3D structures) through a unified conversion process. This multi-functional approach establishes consistent standards across different data formats and sources, enabling reliable duplication detection and ensuring data consistency throughout the processing pipeline.
Solution Approach 2:
The patent segments chemical structures into discrete graph components (nodes representing atoms, edges representing bonds) with standardized features. This segmentation approach creates a uniform granular representation that enables precise comparison and standardization across different chemical compounds, resolving the issue of inconsistent data standards while maintaining high precision in data cleaning.
3Ease of operation
If SMILES text is directly used for machine learning applications, then the process is simple, but non-standard or repetitive structures cannot be effectively handled
Solution Approach 1:
The patent introduces graph structure representation as an intermediary between raw SMILES text and machine learning models. This intermediary transformation layer converts variable-length SMILES strings into standardized graph formats with consistent node and edge features, maintaining processing simplicity while dramatically improving structure standardization accuracy and enabling effective handling of non-standard and repetitive structures.
Data Source
AI summary
The present invention provides a data preprocessing method for cleaning a small molecule compound, the data preprocessing method comprising: an S1 text preprocessing step including: preprocessing an original SMILES text of a small molecule compound into a standardized SMILES text of the small molecule compound; and an S2 chemical graph formatting step including: splitting in a format each text element of the standardized SMILES text of the small molecule compound of S1 to obtain chemical graph information of the small molecule compound. The present invention also provides a data preprocessing system for cleaning a small molecule compound. The present invention enables the cleaning, deduplication, and standardization of global datasets, providing an efficient, fast, accurate integration method for the cleaning of end-to-end small molecule compounds.


