Document Priority Scoring for Gene and Drug Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large content repositories like PubMed face challenges in providing intelligent access to vast amounts of documents, with limited accessibility and no intelligent way to filter or rank relevant information, making it difficult for users to find specific combinations of information such as gene, drug, and cancer-type related data.
Innovation Solution
A system that preprocesses documents to make them machine-readable, classifies them into categories based on specific sections, and assigns a priority score based on the frequency of search terms, allowing users to search and rank documents by relevance using a user interface that differentially weights sections for accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If documents are stored in large content repositories without intelligent filtering, then the quantity of stored information increases, but the accessibility and ease of finding specific information deteriorates
Solution Approach 1:
The patent segments documents into structured sections (title, abstract, introduction, methods, results, discussion, conclusion) and applies different weighting factors to each section. This segmentation allows the system to efficiently search and filter specific portions of documents rather than processing entire documents, resolving the contradiction between storing large quantities of documents and enabling easy access to specific information.
Solution Approach 2:
The patent introduces an intermediary priority scoring system that mediates between the large repository of documents and user search queries. The scoring system processes documents by evaluating search term frequencies in weighted sections, creating an intermediate ranked list that facilitates efficient access without requiring users to search through the entire repository manually.
2Loss of information
If full-length research documents are made accessible, then the completeness of information increases, but the complexity of evaluating accuracy and relevance increases
Solution Approach 1:
The patent applies local quality by assigning different weighting factors to different sections of documents based on their relevance to research accuracy. Sections like results, methods, and discussion receive higher weights than title or abstract, allowing the system to focus evaluation on locally important areas rather than treating the entire document uniformly, thus reducing evaluation complexity while maintaining information completeness.
3Measurement precision
If intelligent filtering and priority scoring are implemented, then the accuracy of document retrieval increases, but the processing time and computational resources increase
Solution Approach 1:
The patent implements partial action by focusing processing only on specific weighted sections of documents rather than analyzing entire documents. By concentrating computational resources on high-weight sections (results, methods, discussion) and using predefined weighting schemes, the system achieves accurate retrieval without the excessive processing time that would result from analyzing every portion of every document in the repository.
Data Source
AI summary
Computer-based methods, systems, and computer readable media for managing documents within a content repository or documents within the document subsets are provided. Documents may be pre-processed to be machine readable and classified within the content repository into one or more categories, based upon a number of times classification terms appear in a specific section of the document or based on an article type tag. Document subsets may be generated based on user-defined terms. Documents may be associated with specific cancer-types, genes, gene variants and drugs by comparing relevant search terms to specific sections of the documents. A request for processing the documents may include one or more of the search terms, pertaining to one or more from a group of gene, gene variant, drug, and cancer terms. A priority score may be determined for documents based on a frequency of one or more of the search terms in each of the specific sections, and the documents may be ranked from highest total priority score to lowest total priority score.


