Distributed spider system and periodical increment capture method
A crawler system and distributed technology, which is applied in the field of efficient data collection of Internet big data, can solve the problem of periodic increment of web page repetition, multi-node task distribution, and capture, etc., to reduce development costs and increase Usable, Simple Architecture Effects
- Summary
- Abstract
- Description
- Claims
- Application Information
AI Technical Summary
Problems solved by technology
Method used
Image
Examples
Embodiment Construction
[0038] In order to better understand the technical content of the present invention, specific embodiments are given together with the accompanying drawings for further description.
[0039] figure 1 It is an architecture diagram of a distributed crawler system of the present invention, the system includes three parts: ZooKeeper-based distributed service, system components and database. Among them, the distributed service based on ZooKeeper provides distributed coordination services for each system component; the system components include the system monitoring component Monitor, the coordination component Coordinator, the log collection component Logger, and the basic crawler component Spider; the database includes Redis memory database and other storage capture For the database of web pages, the distributed URL task queue and distributed BloomFilter are stored in the Redis memory database.
[0040] The distributed service based on ZooKeeper coordinates with each system compon...
PUM
Abstract
Description
Claims
Application Information
- R&D Engineer
- R&D Manager
- IP Professional
- Industry Leading Data Capabilities
- Powerful AI technology
- Patent DNA Extraction
Browse by: Latest US Patents, China's latest patents, Technical Efficacy Thesaurus, Application Domain, Technology Topic, Popular Technical Reports.
© 2024 PatSnap. All rights reserved.Legal|Privacy policy|Modern Slavery Act Transparency Statement|Sitemap|About US| Contact US: help@patsnap.com