Structural Formula Image Analysis Across Drawing Formats

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image analysis techniques struggle to accurately identify structural formulas of compounds from images due to variations in drawing formats, making it difficult to search for and manage such data effectively.

Innovation Solution

An image analysis apparatus and method using machine learning to generate symbol information from structural formula images, employing a convolutional neural network for feature extraction and a recurrent neural network for symbol recognition, capable of handling different drawing styles and formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If pattern recognition and predetermined algorithms are used to recognize text information and line diagram information separately, then the structural formula can be identified according to established rules, but the system cannot cope with new or varied drawing formats requiring a large number of predefined rules

Engineering Contradiction:
Improveidentification accuracyVSAvoidadaptability to different drawing formats
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent replaces the mechanical rule-based recognition system with a deep learning-based image processing system. Instead of using predetermined algorithms to recognize text and line diagrams separately according to fixed rules, the system uses a deep learning model trained on diverse structural formula images to automatically extract and identify structural elements, enabling the system to adapt to various drawing formats without requiring manual rule updates

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters of the recognition approach by transitioning from discrete rule-based processing to continuous deep learning-based processing. The system learns optimal recognition parameters automatically from training data, allowing it to handle variations in drawing styles, bond line thicknesses, orientations, and other format differences through learned feature representations rather than fixed parameter thresholds

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If a large number of rules are established to cope with different ways of drawing structural formulas, then various drawing formats can be handled, but the complexity of the identification system increases significantly

Engineering Contradiction:
Improvecoverage of drawing formatsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal deep learning model that performs multiple functions simultaneously: feature extraction, structural element recognition, and formula identification. This single multi-functional system replaces numerous specialized rules that would otherwise be needed to handle different drawing formats, bond representations, and atomic symbol variations, thereby reducing overall system complexity while maintaining broad format coverage

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent extracts the essential structural features from images using deep learning, separating the core structural information from the variable formatting elements. By extracting only the relevant structural features (atomic symbols, bond connections, spatial relationships) and discarding format-specific variations, the system achieves format coverage without requiring complex rules for each variation

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If identification rules are not established for new drawing formats, then the system remains simple, but identification of structural formulas in new formats becomes impossible

Engineering Contradiction:
Improvesystem simplicityVSAvoididentification capability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent performs preliminary training of the deep learning model on a diverse dataset of structural formula images representing various current and potential drawing formats. This preliminary action equips the system with pre-learned knowledge of structural patterns across different formats, enabling it to reliably identify structural formulas in new or unseen formats without requiring subsequent rule updates or system modifications

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12417648B2Image analysis apparatus, image analysis method, and program
Publication Date: 2025.09.16 FUJIFILM CORP
  • US12417648B2 patent drawing
  • US12417648B2 patent drawing
  • US12417648B2 patent drawing

AI summary

There are provided an image analysis apparatus, an image analysis method, and a program for implementing an image analysis method that can, when text information about a structural formula of a compound is generated from an image showing the structural formula, cope with a change in the way of drawing of the structural formula.An image analysis apparatus according to one embodiment of the present invention includes a processor, and the processor is configured to generate, on the basis of a feature value of a subject image showing a structural formula of a subject compound, symbol information representing the structural formula of the subject compound with a line notation, by using an analysis model. The analysis model is a model created through machine learning using a learning image and symbol information representing a structural formula of a compound shown by the learning image with a line notation.