Scientific Figure Segmentation for Experiment Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Natural language processing (NLP) technologies struggle to effectively analyze scientific documents due to the presence of complex images that contain heterogeneous elements and represent multiple experiments in different sub-images, which are not adequately captured by text-based analysis.

Innovation Solution

A computer-implemented method that segments images into sub-images using machine learning models, classifies experimental techniques, determines macromolecules and contexts, and maps sub-images to experiments, integrating this information into a knowledge base for querying and correlation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If NLP technology is used to analyze scientific documents, then text-based analysis can be performed, but information contained in images is not captured

Engineering Contradiction:
Improveinformation extraction completenessVSAvoidmulti-modal analysis capability
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The patent segments images into multiple sub-images, where each sub-image represents a distinct experiment. This segmentation allows the system to process and analyze individual experimental units separately, enabling comprehensive extraction of experimental information from complex multi-panel scientific figures while maintaining the ability to handle diverse image structures

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mapping mechanism that connects sub-images to experiment descriptions. This intermediary layer translates visual information from sub-images into structured experimental data, bridging the gap between image analysis and text-based NLP processing, thereby enabling complete information extraction from both modalities

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If machine learning models are trained on complex images with heterogeneous elements, then image analysis capability improves, but training difficulty increases

Engineering Contradiction:
Improveimage analysis accuracyVSAvoidmodel training complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides complex heterogeneous images into simpler sub-images, each representing a single experiment. This segmentation reduces the complexity of individual training samples, making it easier to train machine learning models while maintaining high analysis accuracy for each experimental unit

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing and analysis approaches to different sub-images based on their specific characteristics. Each sub-image can be analyzed with appropriate methods tailored to its content, improving overall accuracy without requiring a single complex model to handle all variations

Inventive Principle:
Principle #3Local quality

3Productivity

If images are treated as single units, then processing is simpler, but multiple experiments within sub-images cannot be distinguished

Engineering Contradiction:
Improveexperiment identification accuracyVSAvoidimage processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent automatically segments single images into multiple sub-images, where each sub-image corresponds to a distinct experiment. This segmentation enables the system to identify and process multiple experiments within what would otherwise be treated as a single image unit, significantly improving experiment identification accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of analysis by creating sub-image level granularity within images. This dimensional transformation allows the system to navigate between image-level and sub-image-level processing, enabling distinction of multiple experiments while maintaining an organized hierarchical structure

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentEP3757877B1Determining experiments represented by images in documents
Publication Date: 2026.04.08 SCINAPSIS ANALYTICS INC D B A BENCHSCI
  • EP3757877B1 patent drawingFigure 1
  • EP3757877B1 patent drawingFigure 2
  • EP3757877B1 patent drawingFigure 3

AI summary

A method may include acquiring one or more image texts from an image of a document, segmenting the image into one or more sub-images using the one or more image texts, determining, by applying a machine learning model, one or more experimental techniques of one or more experiments for the one or more sub-images, and adding, to a knowledge base, one or more mappings of the one or more sub-images to the one or more experiments.