Web Content Analysis for New Business Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying new business openings are inefficient due to reliance on structured published information and time-consuming web searches, which lack reliability and provide little competitive advantage, especially when dealing with unstructured web content that varies in quality.
Innovation Solution
A system and method for automatically identifying new businesses by programmatically analyzing content from online sources, using source-specific algorithms to extract references such as business names, locations, and opening dates, and storing historical data to refine search queries and access high-quality content sources, thereby reinforcing source and business discovery processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated web content analysis is implemented to identify new businesses, then identification accuracy and speed improve, but system complexity increases
Solution Approach 1:
The system segments the web content analysis process into distinct modules: content crawling, text processing, entity recognition, and verification. Each module handles a specific aspect of the analysis, making the overall complex system manageable and maintainable while achieving high identification speed through parallel processing of these segmented tasks.
Solution Approach 2:
The patent introduces intermediary components such as trained classifiers and entity recognition models that act as mediators between raw web content and business identification results. These intermediaries process and transform unstructured content into structured information, reducing the complexity of direct analysis while improving identification accuracy and speed.
2Measurement precision
If source-specific algorithms are used to analyze web content, then measurement precision of business references improves, but device complexity increases
Solution Approach 1:
The system applies source-specific algorithms tailored to the characteristics of different online sources (e.g., review sites, social media, news portals). Each source type receives customized processing rules and entity recognition models optimized for its specific content format, improving reference accuracy while managing algorithm complexity through localized rather than universal processing.
Solution Approach 2:
The patent employs trained classifiers that adjust processing parameters based on the specific source being analyzed. The system dynamically changes analysis parameters such as entity recognition thresholds, text processing depth, and verification criteria according to source characteristics, achieving high precision without requiring equally complex algorithms for all sources.
3Productivity
If historical business data is stored and reused in search queries, then discovery efficiency improves, but information storage requirements increase
Solution Approach 1:
The system performs preliminary actions by pre-processing and storing historical business data in structured formats with extracted key attributes (business names, categories, locations, opening dates). This pre-processed data is readily available for quick comparison and query matching, improving discovery efficiency while minimizing storage requirements through efficient data structuring and selective retention of only essential attributes.
Data Source
AI summary
In general, embodiments of the present invention provide systems, methods and computer readable media for identifying a new business based on programmatically analyzing content received from online sources and, as a result, discovering one or more references to the business. In embodiments, the system stores historical data representing previously identified new businesses and then uses attributes of those businesses in search queries to receive related content. Additionally or alternatively, the system stores data representing online sources that historically provided content containing references to new businesses and then continues to access those sources for additional content. In embodiments, the system performs content analysis on structured and/or unstructured content. In some embodiments, analysis of content received from a particular online source includes a source-specific algorithm that takes a source-specific representation of the content as input and produces a result indicating the likelihood that the content includes a new business reference.


