Automated Keyword Mapping via Feature Vector Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing search engine technologies lack a flexible method to determine the meaning of Web pages and documents, relying on predetermined conceptual maps that fail to adapt to document content and cultural differences, leading to inaccurate targeted advertising and search results.
Innovation Solution
An automated system and method for mapping keywords and key phrases to documents, which analyzes a corpus of documents to generate feature vectors that weight related words and phrases, allowing for real-time mapping and relevance determination based on document content, without requiring predetermined relationships between concepts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If predetermined conceptual maps are used to determine meaning of documents, then the search engine can operate with fixed conceptual relationships, but the system fails to adapt to document content and cultural differences
Solution Approach 1:
The system performs preliminary analysis of documents to extract key terms and phrases before the actual search query is processed. This pre-processing creates a foundation of document-specific terminology that enables accurate matching without requiring complex real-time conceptual mapping, thus resolving the contradiction between adaptability and complexity.
Solution Approach 2:
The search engine automatically extracts and uses key terms from the document content itself to determine relevance, rather than relying on external predetermined conceptual maps. The system serves its own need for accurate meaning determination by leveraging the document's actual terminology, achieving both adaptability and simplicity.
2Measurement precision
If hard match between search query and indexed key terms is required, then the search service provides precise keyword matching, but the system cannot handle spelling mistakes, partial queries, or variations in word order
Solution Approach 1:
The system segments the search query into individual words and phrases, then analyzes each segment separately against the document's extracted key terms. This segmentation allows the system to handle partial matches, spelling variations, and different word orders by evaluating each segment independently, thus achieving both precision and flexibility.
Solution Approach 2:
The system accepts partial matches between query terms and document key terms, rather than requiring complete exact matches. By allowing partial overlaps and variations, the system can handle spelling mistakes and partial queries while maintaining sufficient matching precision through the extracted key term methodology.
3Reliability
If manual submission of key terms by website owners is required, then the search service can provide targeted advertising, but the process is time-consuming and does not scale well
Solution Approach 1:
The system automatically extracts key terms from document content without requiring manual input from website owners. The automated extraction process analyzes document text to identify relevant key terms and phrases, enabling rapid processing while maintaining advertising relevance through the extracted terminology, thus resolving the contradiction between reliability and productivity.
Solution Approach 2:
The system replaces the manual mechanical process of key term submission with an automated computational extraction process. By using algorithms to automatically identify and extract key terms from document text, the system eliminates manual labor while maintaining the relevance needed for effective targeted advertising, significantly improving processing speed.
Data Source
AI summary
A method for automated mapping of key terms to documents. Preferably a feature vector is generated for the key terms. Preferably such feature vectors are automatically generated by analyzing a corpus, but may optionally be generated manually, or using a combination of automated and manual processes. Next, preferably such feature vectors are weighted. Such weighting may optionally be performed manually, but more preferably is performed automatically. Next, a feature vector is optionally and preferably generated for the document, which preferably includes words and phrases that were extracted from the document but may optionally include words and phrases that do not appear in the document, such as synonyms and related. Each element in the document feature vector is preferably weighted. The feature vectors of the key terms and the feature vectors of the documents are compared, in order to produce relevancy scores, which are used to produce mapping between documents and key terms. This mapping may optionally be used for a wide variety of applications, such as for targeted advertising for example.


