Chemical Entity Recognition for Patent Relevance Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for extracting chemical compounds from patent documents are inefficient, as they often require manual processes that are costly and time-consuming, and automatic approaches struggle with interpreting chemical structures and relevance, leading to incomplete and inaccurate data.

Innovation Solution

A system and method for training a chemical entity recognition system to normalize patent documents, extract chemical entities, and determine their relevance, using a chemical patent corpus with annotated relevancy, enabling automatic identification of relevant compounds within patent documents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual processes are used to extract chemical compounds from patent documents, then extraction accuracy is improved, but processing time and cost increase

Engineering Contradiction:
Improveextraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary normalization of patent documents into a unified format, extracting and annotating chemical entities with relevance labels before the main extraction process. This preprocessing step improves the accuracy of subsequent automated extraction while reducing the time required for manual review.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an automated chemical entity recognition system as an intermediary between manual extraction processes. The system uses trained models to identify and classify chemical compounds, providing high-accuracy extraction results that reduce reliance on time-consuming manual processes while maintaining precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated approaches are used to extract chemical compounds, then processing speed is improved, but extraction accuracy and completeness deteriorate

Engineering Contradiction:
Improveprocessing speedVSAvoidextraction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary training using a curated chemical patent corpus with annotated chemical entities and relevance labels. This pre-training phase enables the automated system to achieve high extraction accuracy before deployment, ensuring that speed improvements do not compromise precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates feedback mechanisms where extraction results are evaluated against the trained corpus, and the model is continuously refined. This feedback loop maintains high extraction accuracy while processing large volumes of patent documents at automated speeds.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If all chemical compounds are cataloged from patent documents, then completeness of data is improved, but data volume and maintenance cost increase

Engineering Contradiction:
Improvedata completenessVSAvoiddata management complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system extracts only the most relevant chemical entities from patent documents based on trained relevance criteria, rather than cataloging all mentioned compounds. This selective extraction approach maintains data completeness for important compounds while significantly reducing the overall data volume and associated maintenance complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different quality standards to different chemical entities based on their relevance to the patent document. High-relevance compounds receive detailed annotation and cataloging, while less relevant compounds are either excluded or given minimal processing, optimizing the balance between completeness and manageability.

Inventive Principle:
Principle #3Local quality

4Measurement precision

If manual annotation of chemical entities is performed, then annotation accuracy is improved, but time and resource consumption increase

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system uses automatically annotated chemical patent corpus to train the entity recognition model, enabling the system to perform self-annotation of new patent documents. This self-service capability maintains high annotation accuracy while eliminating the need for continuous manual annotation, significantly improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary annotation using the trained model before any manual review process. This preliminary automated annotation covers the majority of cases accurately, requiring manual intervention only for edge cases, thereby maintaining high accuracy while dramatically improving overall annotation efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11537788B2Methods, systems, and storage media for automatically identifying relevant chemical compounds in patent documents
Publication Date: 2022.12.27 ONTOCHEM
  • US11537788B2 patent drawing
  • US11537788B2 patent drawing
  • US11537788B2 patent drawing

AI summary

Methods, systems, and non-transitory media for training a chemical entity recognition system to extract chemical compounds from a patent document and determine a relevance of the chemical compounds to the patent document are disclosed. A method includes obtaining patent documents from patent databases, normalizing each patent document into a unified format, and generating a chemical patent corpus. The chemical patent corpus includes chemical entities, each having relevancy annotations that indicate a relevance to the patent document from which the chemical entity is extracted. The method further includes providing the chemical patent corpus to the chemical entity recognition system, which tags the one or more chemical entities in a corresponding normalized patent document, extracts additional chemical entities, assigns a confidence score to each additional chemical entity, and labels each additional chemical entity as relevant or irrelevant to an associated patent document based on information contained in the chemical patent corpus.