Chemical Entity Recognition for Patent Relevance Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for extracting chemical compounds from patent documents are inefficient, as they often require manual processes that are costly and time-consuming, and automatic approaches struggle with interpreting chemical structures and relevance, leading to incomplete and inaccurate data.
Innovation Solution
A system and method for training a chemical entity recognition system to normalize patent documents, extract chemical entities, and determine their relevance, using a chemical patent corpus with annotated relevancy, enabling automatic identification of relevant compounds within patent documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual processes are used to extract chemical compounds from patent documents, then extraction accuracy is improved, but processing time and cost increase
Solution Approach 1:
The system performs preliminary normalization of patent documents into a unified format, extracting and annotating chemical entities with relevance labels before the main extraction process. This preprocessing step improves the accuracy of subsequent automated extraction while reducing the time required for manual review.
Solution Approach 2:
The patent introduces an automated chemical entity recognition system as an intermediary between manual extraction processes. The system uses trained models to identify and classify chemical compounds, providing high-accuracy extraction results that reduce reliance on time-consuming manual processes while maintaining precision.
2Productivity
If automated approaches are used to extract chemical compounds, then processing speed is improved, but extraction accuracy and completeness deteriorate
Solution Approach 1:
The system performs preliminary training using a curated chemical patent corpus with annotated chemical entities and relevance labels. This pre-training phase enables the automated system to achieve high extraction accuracy before deployment, ensuring that speed improvements do not compromise precision.
Solution Approach 2:
The system incorporates feedback mechanisms where extraction results are evaluated against the trained corpus, and the model is continuously refined. This feedback loop maintains high extraction accuracy while processing large volumes of patent documents at automated speeds.
3Quantity of substance
If all chemical compounds are cataloged from patent documents, then completeness of data is improved, but data volume and maintenance cost increase
Solution Approach 1:
The system extracts only the most relevant chemical entities from patent documents based on trained relevance criteria, rather than cataloging all mentioned compounds. This selective extraction approach maintains data completeness for important compounds while significantly reducing the overall data volume and associated maintenance complexity.
Solution Approach 2:
The system applies different quality standards to different chemical entities based on their relevance to the patent document. High-relevance compounds receive detailed annotation and cataloging, while less relevant compounds are either excluded or given minimal processing, optimizing the balance between completeness and manageability.
4Measurement precision
If manual annotation of chemical entities is performed, then annotation accuracy is improved, but time and resource consumption increase
Solution Approach 1:
The system uses automatically annotated chemical patent corpus to train the entity recognition model, enabling the system to perform self-annotation of new patent documents. This self-service capability maintains high annotation accuracy while eliminating the need for continuous manual annotation, significantly improving productivity.
Solution Approach 2:
The system performs preliminary annotation using the trained model before any manual review process. This preliminary automated annotation covers the majority of cases accurately, requiring manual intervention only for edge cases, thereby maintaining high accuracy while dramatically improving overall annotation efficiency.
Data Source
AI summary
Methods, systems, and non-transitory media for training a chemical entity recognition system to extract chemical compounds from a patent document and determine a relevance of the chemical compounds to the patent document are disclosed. A method includes obtaining patent documents from patent databases, normalizing each patent document into a unified format, and generating a chemical patent corpus. The chemical patent corpus includes chemical entities, each having relevancy annotations that indicate a relevance to the patent document from which the chemical entity is extracted. The method further includes providing the chemical patent corpus to the chemical entity recognition system, which tags the one or more chemical entities in a corresponding normalized patent document, extracts additional chemical entities, assigns a confidence score to each additional chemical entity, and labels each additional chemical entity as relevant or irrelevant to an associated patent document based on information contained in the chemical patent corpus.


