Document Retrieval System Using Classification Code Definition Sentences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document retrieval methods face challenges in accurately retrieving classification codes due to hierarchical and minute classification codes, leading to retrieval noises and missing documents, especially when dealing with large numbers of assigned codes and hierarchical structures.
Innovation Solution
A document retrieval program and apparatus that acquires key information, extracts candidate classification codes, and evaluates them based on comparison with related key information, allowing for the retrieval of classification codes of an appropriate grain size by utilizing a multi-step process involving keyword and classification code interactions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If classification codes are used as key for document retrieval, then retrieval accuracy is improved, but device complexity increases due to hierarchical and minute classification codes
Solution Approach 1:
The patent introduces classification code definition sentences as intermediary elements between user queries and the hierarchical classification code system. These sentences serve as mediators that bridge the gap between simple keyword searches and the complex hierarchical classification structure, allowing users to retrieve documents using natural language descriptions rather than navigating complex code hierarchies directly
Solution Approach 2:
The patent segments the complex hierarchical classification code system into manageable components by extracting and utilizing individual definition sentences for each classification code. This segmentation allows the system to handle the complexity by processing discrete, understandable units (definition sentences) rather than presenting the entire hierarchical structure to users
2Ease of operation
If definition sentences of classification codes are used for retrieval, then ease of operation is improved, but retrieval noises increase due to keyword-based retrieval limitations
Solution Approach 1:
The patent implements a feedback mechanism by retrieving documents that contain both the classification code definition sentence and the search keyword, then using the classification codes associated with these retrieved documents to refine and validate the search results. This feedback loop helps eliminate retrieval noises by verifying results against the hierarchical classification structure
Solution Approach 2:
The patent merges two retrieval approaches: keyword-based retrieval using definition sentences and classification code-based retrieval. By combining these methods and cross-validating results, the system maintains the ease of operation provided by natural language queries while reducing retrieval noises through the structured validation of classification codes
3Reliability
If co-occurrence relations of classification codes and keywords are utilized, then missing documents are reduced, but retrieval noises remain when large numbers of classification codes are assigned
Solution Approach 1:
The patent applies local quality by evaluating and weighting individual classification codes based on their specific relevance to the query context, rather than treating all co-occurring classification codes equally. This allows the system to highlight locally relevant classification codes while filtering out less relevant ones, reducing retrieval noises even when many codes are assigned to documents
4Reliability
If classification codes of multiple hierarchies are summed up, then document completeness is improved, but difficulty of detecting and measuring increases due to hierarchical relationships
Solution Approach 1:
The patent performs preliminary action by pre-processing and organizing classification codes according to their hierarchical relationships before the actual retrieval process. This preliminary organization allows the system to efficiently handle multiple hierarchies without increasing detection difficulty during query execution, as the hierarchical structure has already been prepared and indexed
Data Source
AI summary
A medium storing a document retrieval program, a document retrieval apparatus and a document retrieval method that can retrieve a classification code of an appropriate grain size are provided. A medium computer-readably storing a document retrieval program causes a computer to execute a process comprising an acquisition step that acquires the first key information, or the key for retrieving a document, a first retrieval step that causes documents containing the first key information to be retrieved from a predetermined range in the database, a first extraction step that extracts the first classification codes from the documents retrieved by the first retrieval step as candidate classification codes, the first classification code being related to the documents, a second retrieval step that causes documents containing candidate classification codes to be retrieved from the predetermined range, a second extraction step that extracts the second key information contained in the documents retrieved by the second retrieval step, the second key information being of the type same as the first key information and an evaluation step that evaluates the candidate classification codes on the basis of comparison of the first key information and the second key information.


