Patent Search Support Using Language Model Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Even when a user finds appropriate search terms, a large number of patent publication hits requires significant time to confirm the contents, increasing the user's burden and risking the elimination of important patent publications when adjusting search terms.
Innovation Solution
A search support method utilizing a language model that involves receiving a patent document and a reference group, extracting overviews, identifying similar points between the patent document and reference group documents, clustering these points to divide the reference group, and deleting unnecessary clusters based on labels and different points.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the user performs a broad patent search to ensure comprehensive coverage, then the search recall is improved, but the user burden and time required to confirm contents increases significantly
Solution Approach 1:
The patent replaces the mechanical manual review process with an automated semantic analysis system. The system extracts overviews from patent documents and reference groups, uses language models to identify similar points, performs clustering to organize results, and automatically deletes redundant groups. This substitution of automated intelligent systems for manual mechanical review resolves the contradiction by maintaining comprehensive search recall while eliminating the time-consuming manual confirmation burden.
2Loss of time
If the user narrows down search terms to reduce the number of hits, then the time to confirm contents is reduced, but important patent publications may be eliminated from the search results
Solution Approach 1:
The system replaces manual search term adjustment with automated semantic analysis and clustering. By extracting overviews and using language models to identify similar points across all search results, the system automatically clusters patent documents and removes only truly redundant groups. This ensures that important patent publications are retained while still reducing the number of hits, resolving the contradiction between reducing confirmation time and maintaining search recall.
Solution Approach 2:
The system employs feedback mechanisms where the language model continuously analyzes the relationship between patent documents and reference groups, adjusting clustering and deletion decisions based on semantic similarity assessments. This feedback loop ensures that narrowing down results does not eliminate important publications, as the system learns from the semantic content to preserve relevant hits while reducing redundant ones.
3Measurement precision
If the system performs detailed analysis of each patent document to ensure accurate filtering, then the filtering precision is improved, but the processing time and computational resources increase
Solution Approach 1:
The patent extracts only the essential overview information from each patent document and reference group, rather than analyzing the complete content. This extraction of key semantic elements enables accurate filtering and clustering while significantly reducing processing time and computational resources, resolving the contradiction between filtering precision and processing efficiency.
Solution Approach 2:
The system segments the patent documents into manageable overviews and processes them through modular steps: extraction, language model analysis, clustering, and deletion. This segmentation allows for detailed semantic analysis of each segment while maintaining overall processing efficiency, as the system can handle multiple segments in parallel and avoid the computational burden of analyzing complete documents.
Data Source
AI summary
A search support method using a language model is provided. The search support method includes a step of receiving a patent document and a first reference group; a step of obtaining a first overview extracted from the patent document and a plurality of second overviews each of which is extracted from any one of documents belonging to the first reference group; a step of inputting a first instruction sentence for outputting a similar point between the first overview and each of the plurality of second overviews to a language model; a step of dividing the first reference group into two or more second reference groups by performing clustering on the similar points and obtaining a label for each of the second reference group; and a step of deleting at least one of the two or more second reference groups based on the label.


