Document Analyzer for Targeted Advertising via Statistical Theme Ranking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for advertising in network environments often fail to accurately target advertisements due to improper keyword tagging of documents, leading to irrelevant ads being displayed to users, which reduces the effectiveness of pay-per-click models.
Innovation Solution
A document analyzer system that performs statistical and semantic analysis to generate 'essence' metadata, such as keywords or categories, for entire documents or their sections, to improve targeted advertising by identifying the most relevant themes and keywords that capture the document's essence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional keyword tagging methods are used to associate advertisements with documents, then the advertising system can operate with simple implementation, but the keyword accuracy deteriorates leading to irrelevant ads being displayed
Solution Approach 1:
The patent introduces an intermediary component (document analyzer with statistical and semantic analysis engines) that mediates between the document content and the keyword tagging system. This intermediary performs sophisticated analysis to generate accurate keywords without requiring the entire advertising system to become complex, thereby improving keyword accuracy while isolating the complexity to a specific module.
Solution Approach 2:
The system performs preliminary statistical and semantic analysis on documents before advertisement matching occurs. By pre-processing documents to extract essence metadata and ranked keywords, the system prepares accurate tagging information in advance, improving keyword accuracy without adding complexity to the real-time advertisement serving process.
2Measurement precision
If sophisticated statistical and semantic analysis is performed to generate essence metadata, then the targeted advertising accuracy is improved, but the processing time increases
Solution Approach 1:
The document analyzer performs sophisticated statistical and semantic analysis in advance to generate essence metadata, ranked keywords, and theme information. By completing this computationally intensive work before advertisement matching, the system achieves high targeting accuracy without delaying the actual advertisement serving process.
Solution Approach 2:
The system extracts only the most essential and relevant keywords and themes from documents using statistical analysis to identify top-ranked terms. By taking out only the critical essence metadata rather than processing all document content in real-time, the system maintains high accuracy while reducing processing time for advertisement matching.
3Quantity of substance
If multiple advertisers compete for keyword rights, then the advertising revenue increases, but the relevance of displayed advertisements deteriorates when keywords are improperly tagged
Solution Approach 1:
The patent replaces the mechanical/manual keyword tagging system with an automated statistical and semantic analysis engine. This substitution generates objective, data-driven keywords based on actual document content rather than subjective manual tagging, ensuring that multiple advertisers competing for keywords receive accurate, relevant impressions that match their intended target audience.
Solution Approach 2:
The system uses statistical analysis to identify the most representative keywords and themes from document content, creating a feedback loop where the essence metadata continuously reflects the actual document essence. This feedback mechanism ensures that even with multiple advertisers competing for keywords, the tagging remains accurate and relevant to the actual document content.
Data Source
AI summary
A document analyzer receives a collection of text-based terms associated with a document. The document analyzer performs a statistical analysis on the text-based terms to identify a distribution of where the text-based terms appear in the document and relative frequency indicating how often the text-based terms appear in the document. The document analyzer utilizes the distribution and relative frequency information derived from the statistical analysis to rank multiple themes associated with the document. For example, a received listing of multiple themes may not be presented in any useful order, although it can be assumed that the themes in the listing are present in the document. Based on application of distribution and relative frequency information derived from the analysis, the document analyzer can identify which themes are most relevant to the document as a whole and/or which of themes correspond to different portions (e.g., pages or sections) of the document.


