Automated Product Identification in Scientific Documents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Researchers face challenges in accessing and identifying commercial products, such as antibodies and chemicals, mentioned in scientific or research-related documents, due to the complexity of extracting relevant information from text.

Innovation Solution

A system comprising a server that receives documents, divides them into tokens, computes scores for each token to identify commercial products based on features like names, catalog numbers, and manufacturers, and provides outputs indicating product mentions, with a user interface for displaying product information and links.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual review of documents is used to identify products, then accuracy of product identification can be maintained, but time consumption and labor requirements increase significantly

Engineering Contradiction:
Improveproduct identification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical review with an automated computer-based system that uses natural language processing and machine learning algorithms to identify products in scientific documents. The system processes text digitally, extracting product mentions, catalog numbers, and manufacturer information automatically, thereby eliminating the need for manual reading and analysis while maintaining high accuracy through trained models.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary computational layer between the raw document text and the final product identification. This intermediary system uses trained algorithms to bridge the gap between unstructured text and structured product information, enabling automated extraction without direct human intervention while preserving identification accuracy through sophisticated pattern recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated text processing is used to identify products, then processing speed increases, but accuracy of product identification may decrease

Engineering Contradiction:
Improveprocessing speedVSAvoidproduct identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary training of machine learning models using labeled datasets before actual product identification. This preliminary action prepares the system in advance with learned patterns and features, enabling it to accurately process documents at high speed without sacrificing precision during the actual extraction process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where the system continuously improves its identification accuracy based on processed data and outcomes. The machine learning models are refined using feedback from training data and performance metrics, allowing the system to maintain high accuracy while operating at automated processing speeds.

Inventive Principle:
Principle #23Feedback

3Loss of information

If comprehensive product information is extracted from documents, then research data completeness improves, but system complexity increases

Engineering Contradiction:
Improveresearch data completenessVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the product information extraction into distinct modular components: product mention detection, catalog number extraction, manufacturer identification, and data validation. Each module handles a specific aspect of information extraction independently, then integrates results to provide comprehensive product data without requiring a monolithic complex system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal extraction framework that handles multiple types of product information (product names, catalog numbers, manufacturers, URLs) using a single integrated system. This multi-functional approach extracts diverse data elements through unified processing logic, reducing overall system complexity compared to having separate specialized systems for each information type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11120362B2Identifying a product in a document
Publication Date: 2021.09.14 RESGATE
  • US11120362B2 patent drawing
  • US11120362B2 patent drawing
  • US11120362B2 patent drawing

AI summary

Aspects of the present disclosure relate to identifying a product in a document. A computing machine accesses a product mention in a scientific or research-related text, the product mention including one or more attribute values for a plurality of attributes, each attribute being associated with either a single attribute value or no attribute value. The computing machine determines that the attribute values of the product mention correspond to two or more candidate product matches in a product directory. The computing machine identifies, based at least in part on stored data related to the scientific or research-related text, a product match from among the candidate product matches, the product match corresponding to the product mention in the scientific or research-related text. The computing machine provides an output of the product match for storage in conjunction with the product mention in the scientific or research-related text.