Molecule Metadata Extraction for Search Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current document search systems are limited by available metadata, making it difficult to search for complex representations of document content, which increases time and computational resources needed to find relevant information.

Innovation Solution

A metadata database is used to store molecule representations extracted from documents, allowing for searches across multiple representations such as text-based, image-based, and graph-based, improving search accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text-based search metadata is used, then search simplicity is maintained, but search accuracy for complex molecular representations deteriorates

Engineering Contradiction:
Improvesearch accuracyVSAvoidmetadata complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments molecular information into multiple representation types (text-based names, SMILES strings, InChI identifiers, molecular images, and graph structures) and stores them as separate metadata fields. This segmentation allows the system to maintain simple text-based search capabilities while adding specialized search paths for complex molecular representations, thereby improving search accuracy without overwhelming system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds multiple dimensions to molecular search by transforming molecular data into different representation formats. Instead of relying solely on text-based search, the system creates parallel search dimensions including chemical structure graphs, SMILES strings, and molecular images, each enabling searches from different perspectives and improving overall search accuracy for complex molecular representations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If multiple molecule representations are indexed, then search thoroughness is improved, but computational resources required increase

Engineering Contradiction:
Improvesearch thoroughnessVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-processing molecular information during document ingestion and converting it into multiple standardized representations (SMILES, InChI, graph structures) that are stored in the metadata database. This preliminary transformation eliminates the need for complex real-time conversions during search operations, thereby improving search thoroughness while minimizing computational resource consumption during actual search execution.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If comprehensive metadata extraction is performed, then information completeness is improved, but processing time increases

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent uses copying by creating multiple representation copies of the same molecular information in different formats (text-based names, SMILES strings, InChI identifiers, graph structures) during the initial metadata extraction phase. These copies enable comprehensive information retrieval without requiring repeated processing of the original molecular data, thereby improving information completeness while reducing processing time for subsequent search operations.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250139154A1Enhancing document metadata with contextual molecular intelligence
Publication Date: 2025.05.01 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250139154A1 patent drawing
  • US20250139154A1 patent drawing
  • US20250139154A1 patent drawing

AI summary

A molecule representation is extracted from a document and associated with the document in a metadata database. For example, an image of a molecular structure may be extracted from a document and stored in the metadata database in a text-based representation such as SMILES. The metadata database may be searched to identify documents that mention a particular molecule. Continuing the example, the metadata database may be searched with a SMILES representation to identify the document and other documents that refer to the same molecule. The metadata database may index documents based on different types of molecule representations, including text-based, image-based, graph-based, name, abbreviation, etc. This allows search over multiple representations of a molecule, improving accuracy and thoroughness. These improvements reduce the time and computational resources needed to search for documents that refer to a particular molecule.