Text Term Extraction Using Multiple Ranking Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search technologies are inefficient due to the need for iterative searches, reliance on manual review of dense documents, and the challenge of identifying relevant terms amidst variants in patent searching, where different patents use different terms to describe similar concepts.
Innovation Solution
A method and system that analyze text using multiple rankings based on various metrics to extract key terms, allowing for the selection of top-ranked terms for display and use in searches, thereby improving search efficiency and effectiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review of dense patent documents is performed to evaluate search results, then search accuracy can be improved, but time consumption and operational complexity increase significantly
Solution Approach 1:
The system performs preliminary extraction of key terms and generation of document summaries before the search evaluation phase. By pre-processing documents to extract essential terms and create condensed representations, the system prepares information in advance so that evaluators do not need to read entire dense patent documents, thus maintaining accuracy while reducing time consumption.
Solution Approach 2:
The system introduces an intermediary layer between the search query and full document review by generating term-based summaries and relevance indicators. These intermediaries (extracted terms, key phrases, and summary statistics) serve as proxies that capture the essence of dense patent documents, allowing accurate evaluation without direct manual review of complete texts.
2Reliability
If multiple search terms and variants are used to improve search completeness, then more relevant results are found, but the complexity of search operations increases
Solution Approach 1:
The system segments the search process into distinct phases: automatic term extraction from the corpus, variant identification, ranking generation based on multiple metrics, and selective term selection. By breaking down the complex task of comprehensive searching into manageable segments, the system achieves search completeness while simplifying user interaction, as users only need to initiate the search rather than manually construct multiple term variations.
Solution Approach 2:
The system performs self-service by automatically extracting terms, identifying variants, and generating ranked lists of search terms without requiring user intervention. The system autonomously analyzes the patent corpus, extracts relevant terminology, and presents optimized search terms, thereby achieving comprehensive search coverage while eliminating the complexity of manual search term construction.
3Reliability
If iterative searching is performed to discover helpful search terms, then search effectiveness improves, but the number of search operations and time required increase
Solution Approach 1:
The system performs preliminary term extraction and ranking operations on the entire patent corpus before the actual search is executed. By pre-computing term frequencies, variants, and relevance metrics across the corpus, the system prepares a ready-to-use ranked list of search terms, eliminating the need for iterative searching and significantly improving search efficiency while maintaining effectiveness.
Solution Approach 2:
The system incorporates feedback mechanisms by analyzing search results and user interactions to refine term rankings. The system monitors which extracted terms lead to successful searches and uses this feedback to adjust the ranking algorithms, continuously improving search effectiveness without requiring manual iterative searching, thus maintaining high productivity.
Data Source
AI summary
A method of analyzing text is performed at an electronic system that includes one or more processors and memory storing instructions for execution by the one or more processors. The method includes extracting terms from a corpus of text and generating a plurality of rankings of the extracted terms. Each ranking of the plurality of rankings is based on a respective metric of a plurality of metrics. The method further includes defining a set of terms, which includes selecting a number of top-ranked terms from each ranking of the plurality of rankings. The set of terms may be provided for display and/or used as search terms to perform a search.


