Chemical Compound Retrieval via Substructure Vector Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for retrieving documents in the chemical field face challenges due to the numerous names of compounds and the difficulty in collecting large amounts of text data required for accurate distributed expression vectors, leading to low retrieval accuracy.

Innovation Solution

A retrieval device that calculates substructure vectors for chemical compounds, allowing for high-accuracy document retrieval by specifying chemical structures and generating vectors based on substructure counts, and combines these with document vectors for semantic comparison.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If distributed expression vectors are used for document retrieval in the chemical field, then retrieval accuracy may be improved, but large amounts of text data are required which are difficult to collect

Engineering Contradiction:
Improveretrieval accuracyVSAvoidamount of text data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments compound names into substructure units (e.g., functional groups, molecular fragments) to create substructure vectors. This segmentation allows the system to represent chemical compounds using structured components that can be compared without requiring large amounts of text data, thus resolving the contradiction between retrieval accuracy and data quantity requirements

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces substructure vectors as an intermediary representation between compound names and document retrieval. Instead of directly comparing distributed expression vectors of compound names (which require extensive data), the system uses substructure vectors that capture chemical structural information, enabling accurate retrieval with limited data

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If compound names with multiple names are not uniquely identified, then retrieval process is simpler, but retrieval accuracy decreases

Engineering Contradiction:
Improveretrieval process simplicityVSAvoidretrieval accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent changes the parameter used for identification from compound names (which have multiple names) to substructure vectors based on chemical structures. This parameter change allows unique identification of compounds regardless of naming variations, improving retrieval accuracy while maintaining process simplicity through automated structure-based comparison

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20220215907A1Retrieval method, computer-readable recording medium, and retrieval device
Publication Date: 2022.07.07 FUJITSU LTD
  • US20220215907A1 patent drawing
  • US20220215907A1 patent drawing
  • US20220215907A1 patent drawing

AI summary

A retrieval device specifies the chemical structure of a compound indicated by a compound name included in an input document. The retrieval device totalizes, for each substructure of the chemical structure, the number of substructures included in the input document. The retrieval device generates a substructure vector of the input document based on the substructure and the number. The retrieval device outputs one or more documents similar to the input document from a plurality of documents including a stored compound name based on comparison between the substructure vector of the input document and each substructure vector of the documents.