Automated Product Identification in Scientific Documents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Researchers face challenges in accessing and identifying commercial products, such as antibodies and chemicals, mentioned in scientific or research-related documents, due to the complexity of extracting relevant information from text.
Innovation Solution
A system comprising a server that receives documents, divides them into tokens, computes scores for each token to identify commercial products based on features like names, catalog numbers, and manufacturers, and provides outputs indicating product mentions, with a user interface for displaying product information and links.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review of documents is used to identify products, then accuracy of product identification can be maintained, but time consumption and labor requirements increase significantly
Solution Approach 1:
The patent replaces manual mechanical review with an automated computer-based system that uses natural language processing and machine learning algorithms to identify products in scientific documents. The system processes text digitally, extracting product mentions, catalog numbers, and manufacturer information automatically, thereby eliminating the need for manual reading and analysis while maintaining high accuracy through trained models.
Solution Approach 2:
The patent introduces an intermediary computational layer between the raw document text and the final product identification. This intermediary system uses trained algorithms to bridge the gap between unstructured text and structured product information, enabling automated extraction without direct human intervention while preserving identification accuracy through sophisticated pattern recognition.
2Productivity
If automated text processing is used to identify products, then processing speed increases, but accuracy of product identification may decrease
Solution Approach 1:
The patent performs preliminary training of machine learning models using labeled datasets before actual product identification. This preliminary action prepares the system in advance with learned patterns and features, enabling it to accurately process documents at high speed without sacrificing precision during the actual extraction process.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously improves its identification accuracy based on processed data and outcomes. The machine learning models are refined using feedback from training data and performance metrics, allowing the system to maintain high accuracy while operating at automated processing speeds.
3Loss of information
If comprehensive product information is extracted from documents, then research data completeness improves, but system complexity increases
Solution Approach 1:
The patent segments the product information extraction into distinct modular components: product mention detection, catalog number extraction, manufacturer identification, and data validation. Each module handles a specific aspect of information extraction independently, then integrates results to provide comprehensive product data without requiring a monolithic complex system.
Solution Approach 2:
The patent creates a universal extraction framework that handles multiple types of product information (product names, catalog numbers, manufacturers, URLs) using a single integrated system. This multi-functional approach extracts diverse data elements through unified processing logic, reducing overall system complexity compared to having separate specialized systems for each information type.
Data Source
AI summary
Aspects of the present disclosure relate to identifying a product in a document. A computing machine accesses a product mention in a scientific or research-related text, the product mention including one or more attribute values for a plurality of attributes, each attribute being associated with either a single attribute value or no attribute value. The computing machine determines that the attribute values of the product mention correspond to two or more candidate product matches in a product directory. The computing machine identifies, based at least in part on stored data related to the scientific or research-related text, a product match from among the candidate product matches, the product match corresponding to the product mention in the scientific or research-related text. The computing machine provides an output of the product match for storage in conjunction with the product mention in the scientific or research-related text.


