Concept-Based Patent Search Engine for Scientific Relatedness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current prior-art search engines rely on semantic similarity, which poorly performs in identifying scientific relatedness due to the assumption that semantic similarity reflects conceptual relatedness, leading to inadequate identification of relevant patent documents, especially in fields like software where technical phrases differ significantly.
Innovation Solution
A search engine trained using patent examination reports to learn relationships between technical phrases, forming concepts from grouped phrases and ranking documents based on conceptual relatedness, rather than textual overlap, allowing for more accurate identification of scientifically related documents.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If semantic similarity algorithms are used to determine relatedness of patent documents, then the search process is simple and fast, but the accuracy of identifying scientifically related documents deteriorates
Solution Approach 1:
The patent introduces an intermediary layer of scientific concepts as mediators between patent documents. Instead of directly comparing textual similarity of documents, the system maps documents to scientific concepts and compares concepts to determine relatedness. This intermediary concept layer resolves the contradiction by enabling accurate scientific relatedness detection while maintaining search efficiency through pre-computed concept mappings.
Solution Approach 2:
The patent replaces the mechanical textual overlap calculation with a conceptual mapping system. Rather than mechanically counting word overlaps or computing semantic similarity scores, the system substitutes this with a knowledge-based concept mapping approach that uses scientific taxonomy and ontology to determine relatedness, achieving both accuracy and efficiency.
2Extent of automation
If automated prior-art search is performed using existing search engines, then the search process is automated, but the identification of relevant prior-art deteriorates due to reliance on semantic similarity
Solution Approach 1:
The patent incorporates feedback mechanisms where patent examiners review and validate automated prior-art search results. The system uses examiner feedback to refine and improve the automated search algorithms over time, creating a feedback loop that enhances reliability while maintaining automation. Examiner feedback is used to train and retrain the automated search system.
Solution Approach 2:
The patent uses scientific concepts as intermediaries to bridge automated search processes with expert knowledge. The concept mapping system serves as an intermediary that translates automated textual analysis into scientifically meaningful relatedness assessments, improving reliability while preserving automation through systematic concept-based processing.
3Measurement precision
If patent examination reports are used to train search engines, then the accuracy of conceptual relatedness identification improves, but the complexity of the search system increases
Solution Approach 1:
The patent applies preliminary action by pre-processing and storing patent examination reports and scientific concept mappings in a structured database before actual search operations. The system performs preliminary extraction, classification, and indexing of scientific concepts from examination reports, creating ready-to-query concept structures that simplify the actual search process while maintaining high accuracy in conceptual relatedness identification.
Data Source
AI summary
A search engine for searching based on related scientific or technological concepts, comprises: a learning module for learning about relationships between technical phrases based on their rates of occurrence in related documents, therefrom to form concepts from groupings of related phrases, and a search module for searching for related documents to a query document based on occurrence in said related documents of concepts present in said query document, the learning module carrying out said learning based on a training set of documents and inter-document relations.


