Document Search Support Device for Metabolomics Data Interpretation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In metabolomics and related fields, interpreting analysis data results is challenging due to the need for appropriate keyword selection from a large number of literature sources, often leading to excessive narrowing or omission of relevant information, even with the help of information analyzers.
Innovation Solution
A document search support device that acquires information from analysis data, extracts relevant terms, calculates relevance scores, and provides an index value of statistical likelihood to help users efficiently select keywords for literature search.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all relevant terms are presented as keywords for literature search, then comprehensive search coverage is achieved, but excessive narrowing occurs causing omission of relevant information
Solution Approach 1:
The patent changes the parameter of keyword selection by introducing statistical likelihood indices calculated from co-occurrence frequencies. Instead of using all relevant terms or relying on interpreter knowledge, the system quantifies the relevance of each term through statistical parameters (co-occurrence frequency with analyte and with MeSH terms), allowing objective selection of optimal keywords that balance comprehensiveness and precision.
2Reliability
If each relevant term is searched individually to avoid omission, then search completeness is maintained, but the number of extracted literatures becomes excessively large
Solution Approach 1:
The system transforms the search strategy by introducing statistical parameters (co-occurrence frequencies and calculated likelihood indices) to evaluate and rank relevant terms. This allows the interpreter to select a limited number of high-priority keywords with the highest statistical likelihood of yielding useful literature, thereby maintaining search completeness while drastically reducing the volume of extracted literature to a manageable size.
3Measurement precision
If appropriate search keywords are selected based on interpreter knowledge, then accurate literature extraction is achieved, but the process becomes highly dependent on interpreter expertise
Solution Approach 1:
The system enables self-service by automatically calculating statistical likelihood indices for relevant terms based on co-occurrence data from the database. The interpreter no longer needs deep domain knowledge to select appropriate keywords; instead, the system provides objectively ranked term suggestions with statistical evidence, making the process accessible to interpreters regardless of their expertise level while maintaining high extraction accuracy.
4Loss of information
If comprehensive literature search is performed without keyword optimization, then all potential information is captured, but the search process becomes inefficient and time-consuming
Solution Approach 1:
The system performs preliminary action by pre-calculating co-occurrence frequencies between analytes, relevant terms, and MeSH terms from the database. This statistical information is prepared in advance and used to generate ranked keyword suggestions before the actual literature search begins. This preliminary statistical analysis enables efficient keyword selection that captures comprehensive information while significantly reducing search time and improving productivity.
Data Source
AI summary
A device to support work of searching document data for interpreting an information analysis result of analysis data obtained by analyzing a sample containing an analyte, includes: an acquisition unit to acquire first information for identifying the analyte from the analysis data; a reception unit to receive input of second information for searching data of a document for interpreting the information analysis result of the analysis data; an extraction unit to extract, based on the first and second information, terms relevant to the information analysis result, from among terms in data of documents in a database; a calculation unit to calculate, for each relevant term, relevance scores indicating a relevance degree between the relevant term and the first information, and a relevance degree between the relevant term and the second information; and a processing unit to obtain an index value of statistical likelihood from the relevance scores.


