Chemical Compound Retrieval via Substructure Vector Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for retrieving documents in the chemical field face challenges due to the numerous names of compounds and the difficulty in collecting large amounts of text data required for accurate distributed expression vectors, leading to low retrieval accuracy.
Innovation Solution
A retrieval device that calculates substructure vectors for chemical compounds, allowing for high-accuracy document retrieval by specifying chemical structures and generating vectors based on substructure counts, and combines these with document vectors for semantic comparison.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If distributed expression vectors are used for document retrieval in the chemical field, then retrieval accuracy may be improved, but large amounts of text data are required which are difficult to collect
Solution Approach 1:
The patent segments compound names into substructure units (e.g., functional groups, molecular fragments) to create substructure vectors. This segmentation allows the system to represent chemical compounds using structured components that can be compared without requiring large amounts of text data, thus resolving the contradiction between retrieval accuracy and data quantity requirements
Solution Approach 2:
The patent introduces substructure vectors as an intermediary representation between compound names and document retrieval. Instead of directly comparing distributed expression vectors of compound names (which require extensive data), the system uses substructure vectors that capture chemical structural information, enabling accurate retrieval with limited data
2Ease of operation
If compound names with multiple names are not uniquely identified, then retrieval process is simpler, but retrieval accuracy decreases
Solution Approach 1:
The patent changes the parameter used for identification from compound names (which have multiple names) to substructure vectors based on chemical structures. This parameter change allows unique identification of compounds regardless of naming variations, improving retrieval accuracy while maintaining process simplicity through automated structure-based comparison
Data Source
AI summary
A retrieval device specifies the chemical structure of a compound indicated by a compound name included in an input document. The retrieval device totalizes, for each substructure of the chemical structure, the number of substructures included in the input document. The retrieval device generates a substructure vector of the input document based on the substructure and the number. The retrieval device outputs one or more documents similar to the input document from a plurality of documents including a stored compound name based on comparison between the substructure vector of the input document and each substructure vector of the documents.


