Crowd-Sourced Native App Crawling via User Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity of content access mechanisms in mobile-based architectures makes traditional crawling methods inefficient, as content providers offer varying content based on geographic region, device type, and operating system, making it difficult for search engines to maintain accurate search indexes.
Innovation Solution
A crowd-sourced crawling method that utilizes user devices to perform crawling tasks by determining installed native applications and sending work requests to a content acquisition server, which then assigns crawling tasks and receives content from content servers through native applications, allowing for distributed and region-specific content collection without circumventing anti-crawling mechanisms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional crawlers are used to access content, then the crawling process is simple and direct, but the crawler cannot access content through native applications and cannot obtain region-specific or device-specific content
Solution Approach 1:
The patent introduces a crowd-sourced crawling system where user devices with native applications act as intermediaries between the search engine and content providers. These intermediary devices can access content through native applications in a natural way, obtaining region-specific and device-specific content that traditional crawlers cannot access, while the search engine maintains control through task assignment and coordination.
Solution Approach 2:
The crawling function is segmented and distributed across multiple user devices rather than being centralized in a single crawler. Each user device with relevant native applications performs specific crawling tasks, dividing the overall crawling workload and enabling access to content that requires specific device types, operating systems, or geographic regions.
2Measurement precision
If content providers provide different content based on geographic region, device type, and operating system, then content accuracy for specific users is improved, but crawling difficulty increases due to multiple access conditions
Solution Approach 1:
The patent applies local quality by assigning specific crawling tasks to user devices based on their local characteristics such as geographic region, device type, and operating system. Each device crawls content relevant to its local context, ensuring that the search engine obtains accurate, region-specific and device-specific content while simplifying the crawling process for each individual device.
Solution Approach 2:
User devices with native applications naturally access content through their installed applications without requiring special crawling tools or workarounds. The devices self-serve by leveraging their own capabilities to obtain content, eliminating the need for complex anti-crawling mechanism circumvention while maintaining content accuracy.
3Productivity
If a crowd-sourced crawling system is implemented, then content collection efficiency is improved, but system coordination and task management complexity increases
Solution Approach 1:
The patent creates a universal crawling system where diverse user devices with different native applications can all participate in content collection. The system accommodates multiple device types, operating systems, and geographic regions through a single unified framework, improving content collection efficiency while managing complexity through standardized task assignment and result aggregation protocols.
Data Source
AI summary
A method for performing crowd-sourced native application crawling is disclosed. The method includes determining a list of installed native applications installed on a user device and determining whether a set of crawling conditions are met. The method includes generating a work request in response to the set of crawling conditions being met by the user device and transmitting the work request to a content acquisition server. The work request includes the list of installed native applications. The method includes receiving a crawling task including an application access mechanism corresponding to a state of a native application. The method include launching the native application and setting the state of the native application based on the application access mechanism. The native application issues a content request to a content server. The method further includes receiving the content from the content server and transmitting the content to the content acquisition server.


