Crawling algorithm

By executing a web crawling algorithm based on available bandwidth on the data processing hardware, the problem of web page cache keeping fresh and resource consumption control is solved, and efficient web crawling and resource utilization is achieved.

CN120035825APending Publication Date: 2025-05-23GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070163.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-25
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art uses the current version of the web page when updating the index, resulting in high computational costs, especially for a large number of web pages. How to keep the cache of the web page fresh while limiting resource consumption becomes a challenge.

Method used

By executing a computer-implemented method on the data processing hardware, multiple web pages for the web crawler are obtained, the available bandwidth of the web crawler is determined, and the corresponding crawl value for each web page is determined based on the available bandwidth, ensuring that the crawl value meets the threshold, and the web page in the cache memory is updated.

Benefits of technology

This method can keep the freshness of web page cache while limiting resource consumption, and improves the efficiency and resource utilization of web crawling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035825A_ABST
    Figure CN120035825A_ABST
Patent Text Reader

Abstract

A method (300) for a crawling algorithm includes obtaining a plurality of web pages (152) for a web crawler (160) to crawl. The method also includes determining an available bandwidth of the web crawlers (155). The method includes, for each respective web page of the plurality of web pages, determining a respective crawling value for the respective web page based on the available bandwidth (153), and determining that the respective crawling value for the respective web page satisfies a threshold (162). The method includes, in response to determining that the respective crawling value of the respective web page satisfies a threshold, updating the respective web page in the cache memory (150).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to crawling algorithms. Background Art

[0002] A web crawler is a software application that systematically browses the World Wide Web to index the current versions of web pages, usually for use by search engines. In theory, a web crawler should have a recent copy of each web page that can be used by a search engine. However, updating the index with the current version of a web page can be computationally expensive, particularly for a large number of web pages. Therefore, one problem with web crawling is to keep the cache of web pages fresh while limiting the expense on available resources. Summary of the invention

[0003] One aspect of the present disclosure provides a computer-implemented method for a crawling algorithm. The computer-implemented method is performed by data processing hardware, which causes the data processing hardware to perform operations, the operations including obtaining a plurality of web pages for a web crawler to crawl. The operations include determining an available bandwidth for the web crawler. The operations also include, for each corresponding web page in the plurality of web pages, determining a corresponding crawl value for the corresponding web page based on the available bandwidth, determining that the corresponding crawl value of the corresponding web page satisfies a threshold, and in response to determining that the corresponding crawl value of the corresponding web page satisfies the threshold, updating the corresponding web page in a cache memory.

[0004] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, the corresponding crawl value is based on a change indication signal received from the corresponding web page. In these implementations, the change indication signal may include a true positive signal or a false positive signal. Also in these implementations, the change indication signal may include a delayed signal.

[0005] The available bandwidth may satisfy a bandwidth threshold. In addition, determining the corresponding crawl value may include determining the corresponding crawl value using a machine learning engine. In some implementations, the operation includes dividing a plurality of web pages into a plurality of shards. In these implementations, the operation includes, for each respective shard in the plurality of shards, determining a corresponding shard crawl value of the respective shard based on the available bandwidth, determining that the corresponding shard crawl value of the respective shard satisfies a threshold shard value, and in response to determining that the corresponding shard crawl value of the respective shard satisfies the threshold shard value, updating each web page of the respective shard in the cache memory. In these implementations, a corresponding shard bandwidth including a portion of the available bandwidth may be assigned to each respective shard.

[0006] In some implementations, the operations include, for each corresponding web page in the plurality of web pages, estimating an update time for a corresponding crawl value of the corresponding web page, and updating the corresponding crawl value of the corresponding web page at the estimated update time. The operations may also include, for each corresponding web page in the plurality of web pages, updating the corresponding crawl value for the corresponding web page at each time step of a discrete time interval.

[0007] Another aspect of the present disclosure provides a system for a crawling algorithm. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include obtaining a plurality of web pages for a web crawler to crawl. The operations include determining the available bandwidth of the web crawler. The operations also include, for each corresponding web page in the plurality of web pages, determining a corresponding crawl value for the corresponding web page based on the available bandwidth, determining that the corresponding crawl value of the corresponding web page satisfies a threshold, and in response to determining that the corresponding crawl value of the corresponding web page satisfies the threshold, updating the corresponding web page in the cache memory.

[0008] This aspect may include one or more of the following optional features. In some implementations, the corresponding crawl value is based on a change indication signal received from the corresponding web page. In these implementations, the change indication signal may include a true positive signal or a false positive signal. Also in these implementations, the change indication signal may include a delayed signal.

[0009] The available bandwidth may satisfy a bandwidth threshold. In addition, determining the corresponding crawl value may include using a machine learning engine to determine the corresponding crawl value. In some implementations, the operation includes dividing a plurality of web pages into a plurality of shards. In these implementations, the operation includes, for each respective shard in the plurality of shards, determining a corresponding shard crawl value for the respective shard based on the available bandwidth, determining that the corresponding shard crawl value for the respective shard satisfies a threshold shard value, and in response to determining that the corresponding shard crawl value for the respective shard satisfies the threshold shard value, updating each web page of the respective shard in the cache memory. In these implementations, a corresponding shard bandwidth including a portion of the available bandwidth may be assigned to each respective shard.

[0010] In some implementations, the operations include, for each corresponding web page in the plurality of web pages, estimating an update time for a corresponding crawl value of the corresponding web page, and updating the corresponding crawl value of the corresponding web page at the estimated update time. The operations may also include, for each corresponding web page in the plurality of web pages, updating the corresponding crawl value for the corresponding web page at each time step of a discrete time interval.

[0011] The details of one or more implementations of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a diagram of an example system for a crawling algorithm.

[0013] Figure 2 is an example diagram of crawling algorithm events over time.

[0014] Figure 3 is a flow diagram of an example operational arrangement of a method for a crawling algorithm.

[0015] Figure 4 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.

[0016] Like reference numbers in the various drawings indicate like elements. DETAILED DESCRIPTION

[0017] Efficient web crawling is one of the most fundamental data management problems in web search. Web pages are the simplest and most common source of information on the Internet, therefore, extracting and organizing information content from web pages remains a crucial task. Web crawlers crawl web pages regularly to ensure that the information available to the corresponding search engine or web browser is up to date. However, given that there are hundreds of trillions of web pages available online, resource-aware crawling algorithms have great practical significance.

[0018] Web crawlers typically poll web pages to determine whether the web page has changed. When a web page has changed, the web crawler may perform a crawl event (i.e., update the web page information in the memory cache). In an ideal solution, the web crawler would continuously poll all web pages for changes, and then perform crawl events accordingly. However, polling uses network resources, and therefore continuous polling is impractical. In addition, scheduling crawl events at predefined discrete time intervals is also an imperfect solution, because a web page may change multiple times during the time interval without being updated in the cache, resulting in the web page in the memory cache being stale / outdated.

[0019] The implementation of this paper includes a web crawler, which includes an algorithm for efficient and effective web crawling, which is designed to optimize the use of resources while maximizing the probability that the available web page in the memory cache is the current web page. As described herein, an efficient web crawler is intended to schedule crawling events (i.e., updating the web page in the cache) between the content change of the web page and the content request event for the corresponding web page. In other words, the web crawler is intended to minimize the amount of unnecessary crawling events (and polling). In some implementations, the crawling algorithm currently applied optimizes crawling events using one or more change indication signals received from the web page, while maintaining fully updated web pages in the memory cache. For example, the web crawler schedules crawling events based at least in part on the change indication signal received. In some implementations, the web crawler is optimized according to the available bandwidth. For example, the web crawler is optimized based on bandwidth constraints that are constant or non-uniform (i.e., variable over time) over time.

[0020] Figure 1 1 is a schematic diagram of an example system 100 for a crawling algorithm. The system 100 includes a client 12 that uses a client device 10 to access a cloud computing environment 140. The client device 10 includes data processing hardware 16 and memory hardware 18. The client device 10 can be any computing device capable of communicating with the cloud computing environment 140 via, for example, a network 112. The client device 10 can transmit a search engine request 20 to the cloud computing environment 140 and receive search results 22. The search results can be based on a plurality of web pages 152 152a-n stored at a cache memory 150 that can be used by the search engine. The client device 10 includes, but is not limited to, desktop computing devices and mobile computing devices, such as laptops, tablet computers, smart phones, smart speakers / displays, smart appliances, Internet of Things (IoT) devices, and wearable computing devices (e.g., headphones and / or watches).

[0021] In some implementations, client device 10 communicates with a cloud computing environment 140 (also referred to herein as remote system 140) via network 112. The cloud computing environment 140 can be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) with scalable / elastic resources 142, which include computing resources 144 (e.g., data processing hardware) and / or storage resources 146 (e.g., memory hardware). A cache memory 150 can be used to store at least portions of web pages 152 152a-n. In some implementations, the cache memory 150 is available to a search engine such that the search engine can quickly scan the multiple web pages 152 to rapidly return results to a user. The cache memory 150 can store a copy of each web page 152 as a file copy, index, snapshot, etc. Each web page 152 can also include a corresponding crawl value 153 153a-n indicating the priority of the web page 152 to be refreshed. The cloud computing environment 140 can execute some or all of a web crawler 160. Typically, the web crawler 160 is executed remotely (e.g., at remote system 140). However, in some implementations, at least portions of the web crawler 160 are executed locally on client device 10 (e.g., on data processing hardware 16). Similarly, the cache memory 150 can be stored locally at client device 10 or at remote system 140 or any combination thereof.

[0022] The web crawler 160 can manage a cache memory 150 storing a plurality of web pages 152 for a search engine or a web browser. In some implementations, the purpose of the web crawler 160 is to maximize the expected number of web page requests served with fresh copies of the corresponding web pages, while also optimizing the use of system resources. The web crawler 160 can schedule crawl events for the web pages 152 based on the corresponding crawl values ​​153 associated with the web pages 152. The web crawler 160 can determine the crawl value 153 based on various information. For example, the web crawler 160 uses the available bandwidth 155 to at least partially determine the crawl value 153. For example, a greater available bandwidth corresponds to a greater crawl value 153, and therefore corresponds to more frequent crawl events (i.e., because a greater available bandwidth supports more frequent crawling). The crawl value 153 can always be based at least in part on the available bandwidth 155. In some implementations, the web crawler 160 includes a machine learning engine 161 that calculates the crawl value 153. For example, the machine learning engine 161 includes a model that is trained to determine the crawl value 153 based on the available bandwidth 155, the frequency of changes to the web page, the frequency of requests for the web page 152, etc. The web crawler 160 can continuously update the crawl value 153 for the web page 152 as new information is received.

[0023] The web crawler 160 may implement a discrete model in which crawling events are triggered uniformly at fixed time steps within a time interval. For example, within a time interval T, the web crawler 160 triggers crawling events at periodic time steps t 1 ,t 2 ,t 3 ,...,t n Generate a crawl event. In this example, the web crawler 160 compares the crawl value 153 of each web page 152 with the threshold 162. For each web page 152 having a corresponding crawl value 153 that meets (e.g., exceeds) the threshold 162, the web crawler 160 schedules a crawl event for the web page 152n at the next time step. The web crawler 160 can initiate a crawl event by sending a request 161 to an external server 170 corresponding to the web page 152 being updated. In turn, the web crawler 160 can receive the updated web page 152U from the remote server 170. In some implementations, the web crawler 160 estimates an update time for the crawl value 153 for the corresponding web page 152. The estimated update time can be based on the current crawl value 153. For example, the web crawler 160 determines that the web page 152 will be updated at time step t based on the current crawl value 153. it has a crawl value 153 that satisfies threshold 162. Therefore, web crawler 160 is subsequently scheduled to crawl at time step t i The crawl value 153 of the web page 152 is recalculated.

[0024] The web crawler 160 may also determine the crawl value 153 based on a received change event indicator 172 (e.g., a change event indicator received from a remote server 170 corresponding to the web page 152). The change event indicator 172 indicates that a change has occurred at the web page 152 (e.g., content on the web page 152 has been added, removed, or modified). Based on the change event indicator 172, the web crawler 160 may recalculate the crawl value 153 for the web page 152. In some implementations, the change event indicator 172 is a delayed change event indicator. For example, the change event indicator 172 is sent after a non-insignificant time has passed since the time when the web page 152 changed. In this case, the web crawler 160 may or may not adjust the crawl value 153 for the web page 152. For example, if the web crawler 160 updates the web page 152 after the web page 152 changes but before the change event indicator 172 is received, the web crawler 160 does not recalculate the crawl value 153. However, if the web crawler 160 does not update the web page 152 after the web page 152 changes (according to the change event indicator 172), the web crawler 160 can recalculate the crawl value 153 and / or initiate a crawl event. In some implementations, the change event indicator 172 is a true positive signal or a false positive signal. In other words, the change event indicator 172 may not always accurately indicate that a change has occurred at the corresponding web page 152. Therefore, the web crawler 160 can take into account the false positive rate of the received change event indicator 172 when determining the crawl value 153 (e.g., by using a machine learning engine).

[0025] In some implementations, the web crawler 160 bundles multiple web pages 152 into multiple slices. Each slice may include one or more web pages 152. In addition, each slice may be defined based on a certain feature of the web page 152 so that similar web pages 152 are grouped together in the corresponding slice. For example, the slice may be based on a refresh rate, the type of web page, an access rate, etc. In some of these implementations, a portion of the bandwidth of the available bandwidth 155 is assigned to each slice. In some implementations, an equal portion of the available bandwidth is assigned to each slice (i.e., the same amount of bandwidth is assigned to each slice). In other implementations, a corresponding portion of the bandwidth is assigned to each slice based on the features of the web pages included in the slice. For example, a slice that includes a frequently updated web page is assigned a larger portion of the available bandwidth relative to another slice that includes a less frequently updated web page.

[0026] Figure 1 The system 100 is presented for illustration purposes only and is not intended to be limiting. For example, although only a single example of each component is shown, the system 100 includes any number of components 10, 140, 150, 160, and 170. In addition, although some components are described as being located in the cloud environment 140, in some implementations, some or all of the components may be hosted locally on the client device 10. In addition, in various implementations, some or all of the components 150 and 160 are hosted locally on the client device 10, remotely (such as in the cloud environment 140), or some combination thereof.

[0027] Figure 2 2 is an example diagram 200 of crawling algorithm events over time. Timelines 210, 220 represent event actions caused by the web crawler 160 or from the web page 152 (i.e., the server 170 corresponding to the web page 152), respectively. As time progresses, the web crawler 160 may receive information and generate one or more crawling events 212 (indicated by black circles in the diagram 200). For example, the web crawler 160 generates a crawling event 212 by sending a request to an external server 170 corresponding to the web page 152 (i.e., Figure 1The web crawler 160 sends an update request 161 to the web page 152 to initiate a crawl event 212. In turn, the web crawler 160 receives the updated web page 152U as a response. In some implementations, when the corresponding crawl value 153 of the web page 152 meets the threshold 162, the web crawler 160 initiates a crawl event 212 for the web page 152. The threshold 162 can be a predefined value. Alternatively, the threshold 162 is based on the highest percentile of the crawl value 153 of the corresponding web page 152 in the cache memory 150. In some implementations, the crawl event 212 is initiated in response to the change event indicator 172.

[0028] The web crawler 160 may receive information corresponding to the web page 152 (i.e., by a change event indicator 172 generated in response to a change in the web page). The change in the web page is described as a change event 222. When a change event 222 for the web page 152 occurs, the server 170 corresponding to the web page 152 may or may not generate a change event indicator 172. For example, each change event 222 corresponds to one of the following: a change event 222 222a that does not include a corresponding change event indicator 172 (indicated by a white circle in the schematic diagram 200), a change event 222 222b in which a true positive change event indicator 172 is transmitted (indicated by a black square in the schematic diagram 200), a change event 222 222c corresponding to a change event indicator 172 that is a false positive (indicated by a black X in the schematic diagram 200), or a change event 222 222d corresponding to a delayed change event indicator 172 (indicated by a white square in the schematic diagram 200). Here, change event 222a corresponds to a change made to a website without a change event indicator 172. Additionally, true positive change event 222b may correspond to a change event indicator 172 sent when a corresponding web page 152 changes. False positive change event 222c may correspond to a change event indicator 172 sent when a corresponding web page 152 has not changed. Delayed change event 222d corresponds to a change in a web page 152 that occurs a non-insignificant amount of time before the server 170 transmits the change event indicator 172. As described above, the web crawler 160 may use, in part, the change event indicator 172 to determine the crawl value 153 for the web page 152.

[0029] In addition, the machine learning model 161 can be trained to calculate crawl values ​​based on various change events 222. For example, the machine learning model 161 receives a batch of labeled training data, the batch of labeled training data including multiple web pages 152 with corresponding change events 222 and change event indicators 172. The machine learning model 161 can be trained to predict the frequency of updating the web page 152 (i.e., how often the change event 222 occurs). In addition, the machine learning model 161 can also be trained to predict the frequency and type of change events 222. For example, the machine learning model 161 predicts the frequency of change events 222a-d and adjusts the calculation of the crawl value 153 accordingly.

[0030] As described herein, the web crawler 160 provides significant improvements over conventional web crawling solutions. For example, conventional web crawling systems that employ discrete crawling events without a change indication signal (i.e., a change event indicator 172) have significantly lower accuracy than a web crawler 160 that receives a change event indicator 172. In addition, by calculating the crawl value 153 based at least in part on the available bandwidth 155, the web crawler 160 of the present disclosure also uses system resources more efficiently while maintaining fresher data (i.e., updated web pages 152) in the memory cache. In addition, when the bandwidth 155 changes, the web crawler 160 can adjust the crawl value 153 accordingly, further resulting in optimized use of system resources. Additionally, when the web crawler 160 implements machine learning techniques to improve crawl value 153 calculations, the web crawler 160 further improves upon previous web crawling systems by efficiently and effectively executing web crawling events to maintain the latest data (i.e., web pages 152) in a memory cache while reducing consumption of system resources.

[0031] Figure 3 is a flow chart of an exemplary operational arrangement of a method 300 for a crawling algorithm. The method 300 may be, for example, Figure 1 The various components of system 100 or Figure 4The method 300 is performed by a computing device 400 of a remote system 140. For example, the method 300 may be performed on the data processing hardware 144 of the remote system 140, on the data processing hardware 16 of the client device 10, on the data processing hardware 410 of the computing device 400, or some combination thereof. At operation 302, the method 300 includes obtaining a plurality of web pages 152 for crawling by the web crawler 160. At operation 304, the method 300 includes determining an available bandwidth 155 for the web crawler 160. For each respective web page in the plurality of web pages 152, the method 300 performs operations 306, 308, and 310. At operation 306, the method 300 includes determining a respective crawl value 153 for the respective web page 152 based on the available bandwidth 155. At operation 308, the method 300 includes determining that the respective crawl value 153 of the respective web page 152 satisfies the threshold 162. At operation 310 , the method 300 includes updating the web page 152 in the cache memory 150 in response to determining that the corresponding crawl value 153 of the corresponding web page 152 satisfies the threshold 162 .

[0032] Figure 4 4 is a schematic diagram of an example computing device 400 that can be used to implement the systems and methods described in this document. Computing device 400 is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only and are not intended to limit implementations of the inventions described and / or claimed in this document.

[0033] The computing device 400 includes a processor 410, a memory 420, a storage device 430, a high-speed interface / controller 440 connected to the memory 420 and a high-speed expansion port 450, and a low-speed interface / controller 460 connected to a low-speed bus 470 and the storage device 430. Each of the components 410, 420, 430, 440, 450, and 460 is interconnected using various buses and can be installed on a common motherboard or installed in other ways as appropriate. The processor 410 can process instructions for execution within the computing device 400, including instructions stored in the memory 420 or on the storage device 430, to display graphical information of a graphical user interface (GUI) on an external input / output device such as a display 480 coupled to the high-speed interface 440. In other implementations, multiple processors and / or multiple buses and multiple memories and multiple types of memories can be used as appropriate. In addition, multiple computing devices 400 can be connected, each of which provides a portion of the necessary operations (for example, as a server group, a group of blade servers, or a multi-processor system).

[0034] Memory 420 stores information non-temporarily within computing device 400. Memory 420 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Non-temporary memory 420 may be a physical device for temporarily or permanently storing programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 400. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware, such as bootloaders). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disk or tape.

[0035] Storage device 430 is capable of providing mass storage for computing device 400. In some implementations, storage device 430 is a computer-readable medium. In various implementations, storage device 430 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable medium or a machine-readable medium, such as memory 420, storage device 430, or memory on processor 410.

[0036] The high-speed controller 440 manages bandwidth-intensive operations of the computing device 400, while the low-speed controller 460 manages less bandwidth-intensive operations. Such a division of responsibilities is exemplary only. In some implementations, the high-speed controller 440 is coupled to the memory 420, the display 480 (e.g., through a graphics processor or accelerator), and a high-speed expansion port 450 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 460 is coupled to the storage device 430 and the low-speed expansion port 490. The low-speed expansion port 490, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device, such as a switch or a router, for example, through a network adapter.

[0037] As shown, the computing device 400 may be implemented in a variety of different forms. For example, the computing device may be implemented as a standard server 400a or multiple times in a group of such servers 400a, as a laptop computer 400b, or as part of a rack server system 400c.

[0038] Various implementations of the systems and techniques described herein can be implemented in digital electronic and / or optical circuitry, integrated circuit systems, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system that includes at least one programmable processor that can be either special purpose or general purpose and can be coupled to receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.

[0039] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages ​​and / or in assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0040] Software applications (i.e., software resources) may refer to computer software that enables a computing device to perform tasks. In some examples, software applications may be referred to as "applications," "apps," or "programs." Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0041] The processes and logic flows described in this specification may be performed by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to perform functions by operating on input data and generating outputs. The processes and logic flows may also be performed by a dedicated logic circuit system, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). For example, processors suitable for executing computer programs include both general-purpose microprocessors and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as a magnetic disk, a magneto-optical disk, or an optical disk, or may be operably coupled to receive data from one or more mass storage devices or to transfer data to one or more mass storage devices or both. However, a computer need not have such a device. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0042] To provide interaction with a user, one or more aspects of the present disclosure may be implemented on a computer having a display device for displaying information to the user—e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen—and possibly a keyboard and pointing device—e.g., a mouse or trackball—through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, speech, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser on a user's client device in response to a request received from the web browser.

[0043] A variety of implementations have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Therefore, other implementations are within the scope of the following claims.

Claims

1. A computer-implemented method (300) which, when executed by data processing hardware (16, 144), causes the data processing hardware (16, 144) to perform operations, It is characterized in that The operations include: Obtaining a plurality of web pages (152) for crawling by a web crawler (160); determining available bandwidth (155) for the web crawler (160); and For each respective web page (152) in the plurality of web pages (152): determining a corresponding crawl value (153) for the corresponding web page (152) based on the available bandwidth (155); Determining that the corresponding crawl value (153) of the corresponding web page (152) satisfies a threshold value (162); and In response to determining that the corresponding crawl value (153) of the corresponding web page (152) satisfies the threshold (162), the corresponding web page (152) in the cache memory (150) is updated.

2. The method (300) according to claim 1, It is characterized in that The corresponding crawled value (153) is based on a change indication signal (172) received from the corresponding web page (152).

3. The method (300) according to claim 2, It is characterized in that The change indication signal (172) includes a true positive signal or a false positive signal.

4. The method (300) according to claim 2 or 3, It is characterized in that The change indication signal (172) includes a delayed signal.

5. The method (300) according to any one of claims 1 to 4, It is characterized in that The available bandwidth (155) satisfies a bandwidth threshold.

6. The method (300) according to any one of claims 1 to 5, It is characterized in that Determining the corresponding crawl value (153) includes using a machine learning engine (161) to determine the corresponding crawl value (153).

7. The method (300) according to any one of claims 1 to 6, It is characterized in that The operations also include: dividing the plurality of web pages (152) into a plurality of slices; and For each corresponding shard in the plurality of shards: determining a corresponding shard crawl value for the corresponding shard based on the available bandwidth (155); Determining that the corresponding shard crawl value of the corresponding shard satisfies a threshold shard value; and In response to determining that the corresponding shard crawl value of the corresponding shard satisfies the threshold shard value, each web page (152) of the corresponding shard in the cache memory (150) is updated.

8. The method (300) according to claim 7, It is characterized in that Each respective slice is assigned a respective slice bandwidth, the respective slice bandwidth comprising a portion of the available bandwidth (155).

9. The method (300) according to any one of claims 1 to 8, It is characterized in that The operations also include, for each respective web page (152) in the plurality of web pages (152): estimating an update time of the corresponding crawled value (153) of the corresponding web page (152); as well as The corresponding crawl value (153) of the corresponding web page (152) is updated at the estimated update time.

10. The method (300) according to any one of claims 1 to 9, It is characterized in that The operations also include, for each respective web page (152) of the plurality of web pages (152), updating the respective crawl value (153) for the respective web page (152) at each time step of a discrete time interval.

11. A system (100), It is characterized in that include: Data processing hardware (16,144); as well as Memory hardware (18, 146) in communication with the data processing hardware (16, 144), the memory hardware (18, 146) storing instructions that, when executed on the data processing hardware (16, 144), cause the data processing hardware (16, 144) to perform operations, the operations comprising: Obtaining a plurality of web pages (152) for crawling by a web crawler (160); determining available bandwidth (155) for the web crawler (160); and For each respective web page (152) in the plurality of web pages (152): determining a corresponding crawl value (153) for the corresponding web page (152) based on the available bandwidth (155); Determining that the corresponding crawl value (153) of the corresponding web page (152) satisfies a threshold value (162); and In response to determining that the corresponding crawl value (153) of the corresponding web page (152) satisfies the threshold (162), the corresponding web page (152) in the cache memory (150) is updated.

12. The system (100) according to claim 11, It is characterized in that The corresponding crawled value (153) is based on a change indication signal (172) received from the corresponding web page (152).

13. The system (100) according to claim 12, It is characterized in that The change indication signal (172) includes a true positive signal or a false positive signal.

14. The system (100) according to claim 12 or 13, It is characterized in that The change indication signal (172) includes a delayed signal.

15. The system (100) according to any one of claims 11 to 14, It is characterized in that The available bandwidth (155) satisfies a bandwidth threshold.

16. The system (100) according to any one of claims 11 to 15, It is characterized in that Determining the corresponding crawl value (153) includes using a machine learning engine (161) to determine the corresponding crawl value (153).

17. The system (100) according to any one of claims 11 to 16, It is characterized in that The operations also include: dividing the plurality of web pages (152) into a plurality of slices; and For each corresponding shard in the plurality of shards: determining a corresponding shard crawl value for the corresponding shard based on the available bandwidth (155); Determining that the corresponding shard crawl value of the corresponding shard satisfies a threshold shard value; and In response to determining that the corresponding shard crawl value of the corresponding shard satisfies the threshold shard value, each web page (152) of the corresponding shard in the cache memory (150) is updated.

18. The system (100) according to claim 17, It is characterized in that Each respective slice is assigned a respective slice bandwidth, the respective slice bandwidth comprising a portion of the available bandwidth (155).

19. The system (100) according to any one of claims 11 to 18, It is characterized in that The operations also include, for each respective web page (152) in the plurality of web pages (152): estimating an update time of the corresponding crawled value (153) of the corresponding web page (152); as well as The corresponding crawl value (153) of the corresponding web page (152) is updated at the estimated update time.

20. The system (100) according to any one of claims 11 to 19, It is characterized in that The operations also include, for each respective web page (152) of the plurality of web pages (152), updating the respective crawl value (153) for the respective web page (152) at each time step of a discrete time interval.