Search Engine Ranking via NLP Entity Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search tools, particularly in life sciences, are inadequate for scientists to find and access necessary information and products for experiments, as they lack efficient search capabilities and objective rating systems for scientific tools, leading to cumbersome searches through numerous articles and unreliable sourcing.
Innovation Solution
The implementation of Natural Language Processing (NLP), Named-Entity Recognition (NER), and machine learning to analyze and structure scientific data from published sources, creating a knowledge base that enables efficient searching and provides unbiased, objective ratings for scientific tools, improving decision-making in experiments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If scientists manually search through hundreds of articles to find relevant specimens, then comprehensive information can be found, but the time and effort required increases significantly
Solution Approach 1:
The system pre-processes and structures scientific data from published sources before users need to search. Natural Language Processing and Named-Entity Recognition are applied to extract and organize information about specimens, storage conditions, and product characteristics in advance, creating a ready-to-query knowledge base that eliminates the need for manual article searching.
Solution Approach 2:
The patent introduces an intermediary search system that mediates between the vast corpus of published scientific literature and the user's information needs. This intermediary automatically queries structured data from multiple sources, aggregates relevant information, and presents synthesized results, replacing the manual process of reading through hundreds of articles.
2Loss of information
If traditional search tools focus on identifying research literature and trends, then broad research insights are provided, but specific experimental product information remains inaccessible
Solution Approach 1:
The system segments the broad corpus of scientific literature into specific, structured data elements relevant to experimental products. Instead of treating all research literature uniformly, it extracts and categorizes specific information about specimens, storage conditions, product characteristics, and experimental outcomes into discrete, searchable fields that provide reliable product information.
Solution Approach 2:
The patent applies local quality by providing different types of information quality for different needs. While maintaining broad research insights from literature analysis, it simultaneously provides detailed, reliable product-specific information (storage conditions, availability, characteristics) extracted and structured from relevant sources, allowing users to access both macro-level research trends and micro-level product details.
3Quantity of substance
If commercial suppliers and non-commercial laboratories provide specimen information, then product availability is increased, but objective ratings of quality and storage conditions are lacking
Solution Approach 1:
The system enables self-service by automatically collecting and structuring information from multiple sources including commercial suppliers and non-commercial laboratories. Rather than relying on manual evaluation, the system uses Natural Language Processing to extract quality metrics, storage conditions, and availability information directly from published sources, creating objective ratings without human intervention.
Solution Approach 2:
The patent implements feedback mechanisms by continuously querying structured data from diverse sources and using this information to generate and update objective ratings. The system processes information from multiple suppliers and laboratories, aggregates the data, and provides feedback in the form of standardized quality ratings and storage condition assessments that improve with each query cycle.
Data Source
AI summary
A search engine for objects in a corpus of document dynamically evaluates search rank of the objects through Natural Language Processing and machine learning. When a search query is received for a first object, the search engine identifies search results including a plurality of source values that are tied to the first object in the corpus of published documents. A search rank is computed for each identified search result based on content of direct textual references to each of the plurality of source values within the corpus of published documents, as well as a weight assigned to each published document. The identified search results are returned according to the computed search rank.


