Data capture method, device, storage medium and electronic device

By obtaining user data crawling request information and formulating crawling tasks based on data characteristics and website characteristics, the problem of low efficiency of large-scale data crawling in the existing technology is solved, and efficient data crawling is achieved through the collaborative work of multiple crawler programs.

CN117972182BActive Publication Date: 2025-05-16新余市万邦科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410133270.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-05-16
Estimated Expiration
2044-01-30

AI Technical Summary

Technical Problem

When facing large-scale data, the data grabbing efficiency is low, and special analysis of each data object is required.

Method used

By obtaining the user's data crawling request information, the first crawling task is determined based on the data characteristics of the data to be crawled and the website characteristics of the website to be crawled, and it is decomposed into the second crawling task of multiple crawling programs on the time wheel node, so as to realize the data crawling of the coordinated work of multiple crawler programs.

Benefits of technology

It improves the efficiency of data crawling, can quickly formulate preliminary crawling tasks for large-scale data to be crawled, and reasonably allocate crawling tasks through time round scheduling to achieve efficient collaborative crawling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117972182B_ABST
    Figure CN117972182B_ABST
Patent Text Reader

Abstract

The present application provides a data crawling method, device, storage medium and electronic device, which relate to the field of Internet technology. The method includes: obtaining the user's data crawling request information, determining the data to be crawled and the website to be crawled where the data to be crawled is located; determining the first crawling task for the data to be crawled according to the data features corresponding to the data to be crawled and the website features corresponding to the website to be crawled; decomposing the first crawling task to determine the second crawling tasks of multiple crawler programs on the time wheel node; controlling each crawler program to crawl the data to be crawled according to the second crawling task. For large-scale data to be crawled, the present application quickly formulates a preliminary first crawling task through data features and website features, and then reasonably allocates the second crawling task to multiple crawler programs through the time wheel, so as to realize the collaborative and efficient data crawling of multiple crawler programs, thereby effectively improving the efficiency of data crawling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a data capture method, device, storage medium and electronic device. Background Art

[0002] In recent years, users' demand for crawling data of different formats and types from major websites has increased significantly, because these network data can be used for important purposes such as data analysis and business decision-making.

[0003] At present, data capture technology mostly uses methods based on deep learning, which can train special neural network models for different types of data, analyze each data object one by one through the model, and achieve accurate capture. However, when facing large-scale data, this kind of method has low data capture efficiency because it needs to conduct special analysis on each data object. Summary of the invention

[0004] The present application provides a data capture method, device, storage medium and electronic device, which can improve data capture efficiency.

[0005] In a first aspect, the present application provides a data capture method, the method comprising:

[0006] Obtaining data crawling request information from a user, determining the data to be crawled and the website to be crawled where the data to be crawled is located;

[0007] Determine a first crawling task for the data to be crawled according to data features corresponding to the data to be crawled and website features corresponding to the website to be crawled;

[0008] Decomposing the first crawling task to determine second crawling tasks for multiple crawler programs on the time wheel node;

[0009] Control each of the crawler programs to crawl the data to be crawled according to the second crawling task.

[0010] By adopting the above technical solution, the data crawling request information is obtained to clarify the data and website to be crawled; a reasonable first crawling task is formulated according to the data characteristics and website characteristics, and a preliminary one can be formulated for large-scale data to be crawled; the first crawling task is decomposed into a second crawling task that the crawler program can work with through time wheel scheduling. For large-scale data to be crawled, a preliminary first crawling task is quickly formulated based on the data characteristics and website characteristics, and then the second crawling task is reasonably allocated to multiple crawler programs through the time wheel, so that multiple crawler programs can crawl data efficiently in collaboration, thereby effectively improving the efficiency of data crawling.

[0011] Optionally, the obtaining of the user's data crawling request information and determining the data to be crawled and the website to be crawled where the data to be crawled is located includes:

[0012] Obtaining data crawling request information from a user, and extracting key crawling information from the data crawling request information;

[0013] A user portrait is established based on the key crawled information, and the data to be crawled and the website to be crawled where the data to be crawled is located are determined.

[0014] By adopting the above technical solution, we first obtain the user's data crawling request information, then extract the key crawling information from the information, and then build a user interest portrait based on the key crawling information, and finally determine the specific data content that needs to be crawled and the target website where the data is located. By extracting key information and building a user portrait, we can more accurately judge the user's crawling needs, thereby determining a more targeted crawling target, which is helpful for the subsequent formulation of personalized crawling strategies and improving the accuracy and efficiency of data crawling.

[0015] Optionally, determining a first crawling task for the data to be crawled according to data features corresponding to the data to be crawled and website features corresponding to the website to be crawled includes:

[0016] Determining a first data capture mode for the data to be captured according to data features corresponding to the data to be captured;

[0017] Determining a second data crawling mode for the website to be crawled according to website features corresponding to the website to be crawled;

[0018] The first data grabbing mode and the second data grabbing mode are integrated to determine a first grabbing task for the data to be grabbed.

[0019] By adopting the above technical solution, the first data crawling mode is first determined according to the characteristics of the data to be crawled, and then the second data crawling mode is determined according to the characteristics of the website to be crawled, and finally the first and second data crawling modes are integrated to determine the first crawling task for the data to be crawled. Considering both the data characteristics and the website characteristics, the crawling method of the data and the impact on the website are comprehensively determined, so that the first crawling task is both reasonable and controllable, thereby effectively improving the efficiency of data crawling.

[0020] Optionally, the data features include structural features and dynamic features, and determining the first data capture mode of the data to be captured according to the data features corresponding to the data to be captured includes:

[0021] Determine a first initial data capture mode of the data to be captured according to the structural features, wherein the structural features include data format features and nested level features;

[0022] The first initial data grabbing mode is updated according to the dynamic features to obtain the first data grabbing mode of the data to be grabbed, wherein the dynamic features include data storage features, dynamic loading features and content update features.

[0023] By adopting the above technical solution, the first initial crawling mode is first determined according to the structural characteristics of the data to be crawled, and then the first initial mode is updated according to the dynamic characteristics to obtain the first data crawling mode. At the same time, the static structural characteristics and dynamic characteristics of the data are considered, and the initial mode is combined with the update, so that the first data crawling mode is consistent with the data structure and adapts to data changes, so that the target data can be quickly and efficiently crawled, and the data crawling efficiency is improved.

[0024] Optionally, determining the second data crawling mode of the website to be crawled according to the website features corresponding to the website to be crawled includes:

[0025] Calculating the server performance score and data-related score of the website to be crawled;

[0026] Obtaining a limited bandwidth of the website to be crawled, and performing weighted calculation on the limited bandwidth using the server performance score and the data-related score to obtain a crawling rate of the website to be crawled;

[0027] A second data crawling mode of the website to be crawled is determined according to the crawling rate.

[0028] By adopting the above technical solution, the server performance score and data-related score of the website to be crawled are calculated, and the limited bandwidth of the website is taken into account. The crawling rate suitable for the website is obtained through weighted calculation, and then the second data crawling mode is determined according to the rate. Comprehensively considering multiple factors such as server performance, data value and limited bandwidth, a reasonable crawling rate and mode are determined, which can not only obtain data but also will not overload the website server, thereby improving the efficiency of data crawling.

[0029] Optionally, decomposing the first crawling task to determine second crawling tasks for multiple crawler programs on the time wheel node includes:

[0030] Decomposing the first crawling task into a plurality of second crawling tasks, and determining a crawler program of a type corresponding to each of the second crawling tasks;

[0031] The time wheel is used to perform node task scheduling configuration on each of the second crawling tasks, and the second crawling task of each of the crawler programs on the time wheel node is determined.

[0032] By adopting the above technical solution, the first crawling task is first split into multiple second tasks, and then the time wheel scheduling method is used to reasonably configure the second crawling tasks of each crawler program at different time nodes, so as to realize the distribution of task volume and time scheduling, so that multiple crawler programs can work together in an orderly manner according to the nodes set on the time wheel, thereby improving the efficiency of data crawling.

[0033] Optionally, after controlling each of the crawler programs to crawl the to-be-crawled data according to the second crawling task, the method further includes:

[0034] If the actual task time of the crawler program does not conform to the time wheel node time, the second crawling task of other crawlers is adjusted.

[0035] By adopting the above technical solution, after the crawler program actually crawls the data, if it is found that its actual execution time does not match the node time set on the time wheel, the second crawling task of other crawlers on the time wheel will be adjusted accordingly to ensure that multiple crawlers can still work in an orderly and coordinated manner. Dynamic monitoring and scheduling of the crawler crawling process can be achieved, which can ensure that the crawling plan on the time wheel is continuously optimized and adjusted, thereby improving the flexibility and efficiency of the data crawling process.

[0036] In a second aspect, the present application provides a data capture device, the device comprising:

[0037] An information acquisition module is used to obtain the user's data crawling request information, determine the data to be crawled and the website to be crawled where the data to be crawled is located;

[0038] A first task determination module, used to determine a first crawling task for the data to be crawled according to data features corresponding to the data to be crawled and website features corresponding to the website to be crawled;

[0039] A second task determination module, used to decompose the first crawling task and determine second crawling tasks of multiple crawler programs on the time wheel node;

[0040] The data crawling module is used to control each of the crawler programs to crawl the data to be crawled according to the second crawling task.

[0041] In a third aspect, the present application provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing any one of the above methods.

[0042] In a fourth aspect, the present application provides an electronic device, comprising a processor, a memory and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any one of the methods described above.

[0043] In summary, the beneficial effects brought about by the technical solution of this application include:

[0044] By adopting the above technical solution, the data crawling request information is obtained to clarify the data and website to be crawled; a reasonable first crawling task is formulated according to the data characteristics and website characteristics, and a preliminary one can be formulated for large-scale data to be crawled; the first crawling task is decomposed into a second crawling task that the crawler program can work with through time wheel scheduling. For large-scale data to be crawled, a preliminary first crawling task is quickly formulated based on the data characteristics and website characteristics, and then the second crawling task is reasonably allocated to multiple crawler programs through the time wheel, so that multiple crawler programs can crawl data efficiently in collaboration, thereby effectively improving the efficiency of data crawling. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a flow chart of a data capture method according to an embodiment of the present application;

[0046] Figure 2 It is a structural schematic diagram of a data capture device according to an embodiment of the present application;

[0047] Figure 3 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application.

[0048] Description of reference numerals: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION

[0049] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.

[0050] In the description of the embodiments of the present application, words such as "illustrative", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "illustrative", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "illustrative", "for example" or "for example" is intended to present related concepts in a specific way.

[0051] In the description of the embodiments of the present application, the meaning of the term "multiple" refers to two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "include", "comprise", "have" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.

[0052] See also Figure 1 , is a flow chart of a data capture method provided in an embodiment of the present application. The method can be implemented by a computer program, can be implemented by a single-chip microcomputer, or can be run on a data capture device based on the von Neumann system. The computer program can be integrated into an application or run as an independent tool application. The specific steps of a data capture method are described in detail below.

[0053] S101: Obtain data crawling request information from a user, determine the data to be crawled and the website to be crawled where the data to be crawled is located.

[0054] When providing data services, websites will display various structured or unstructured data on the page, such as news content, stock quotes, etc. Users will make data crawling requests for designated pages of a website according to their needs. When crawling data from a website, it is first necessary to clarify which specific data on the website needs to be crawled, such as a certain type of text, table, etc., and which website the data is located on. Obtaining the user's data crawling request information is precisely to achieve this purpose.

[0055] The specific implementation is that the website provides an interface for users to submit data crawling requirements. Users enter the data type description that needs to be crawled in the interface, such as "news text", "stock market table", etc. After receiving the request, the website backend will extract the data description information from it, and then determine what type of data needs to be crawled this time based on the description content, such as text or table, etc. At the same time, it can also parse the website where the data is located.

[0056] Obtaining the data description in the user data crawling request can help the website clarify which specific data in its platform needs to be crawled, such as a certain type of text or table data. At the same time, it also determines that the website where the data is located is itself. In this way, the website can formulate a targeted crawling plan and use the appropriate crawling mode to crawl specific types of data.

[0057] Data crawling request information refers to the relevant information contained in the data crawling requirements submitted by the user, such as the type of data that needs to be crawled, data feature description, crawling time requirements, and other information related to the crawling task.

[0058] In the embodiment of the present application, the data capture request information can be understood as: the description information related to the data to be captured provided by the user when submitting the data capture request, such as the text description of the data content such as "the headline of a news webpage", "the A-share market table of a stock website", etc. This information will be used to determine the specific data type and content characteristics that need to be captured this time.

[0059] The subsequent implementation process is: after receiving the user's crawling request containing data crawling request information, the website backend will extract the description text of the data content from it, and parse out what type of data needs to be crawled based on the description, such as text or table, and the website platform to which these data belong. Then formulate a targeted data crawling plan.

[0060] The data to be crawled refers to the data content in the target website that needs to be obtained through the crawler program when crawling website data. It may include various types of data, such as text, tables, images, etc. Obtaining and parsing the user's crawling request information can help determine the type of data to be crawled for this demand.

[0061] For example, if the description information in the user data crawling request reflects the need to obtain a certain type of latest research paper, then it can be determined that the data to be crawled this time may be text expressing the content of the paper, which may be stored in the database of a paper website server.

[0062] In the embodiment of the present application, the data to be crawled can be understood as: from the request information provided by the user containing the data crawling requirements, by analyzing the request description, it is determined what category or feature of the specified data that the crawler needs to obtain. These analyzed types of data to be crawled will become the targets of subsequent crawlers.

[0063] During implementation, the platform needs to extract data descriptions from user request information and analyze and determine based on the descriptions whether the data to be captured is text, tables, images, or other different types of data. After clarifying the specific categories of data to be captured, the platform can formulate a capture strategy for that data type.

[0064] The website to be crawled refers to the crawling object of the crawler program determined as this data crawling task, that is, the website platform that stores and serves the data to be crawled.

[0065] In an embodiment of the present application, the website to be crawled can be understood as: based on the information provided in the user's data crawling request, the target website for which the crawler program is required to crawl data is determined.

[0066] For example, if the data crawling request reflects the need to crawl the latest papers on a certain paper website, then the paper website is the website to be crawled for this data crawling task.

[0067] In an optional implementation, data crawling request information of a user is obtained, and key crawling information is extracted from the data crawling request information;

[0068] Create a user portrait based on the key crawling information, determine the data to be crawled and the website to be crawled where the data is located.

[0069] In order to further improve the pertinence and effectiveness of crawled data, after obtaining the user's data crawling request information, key crawling information can also be extracted from the data crawling request information.

[0070] The purpose of extracting key crawling information is to accurately judge the user's actual crawling needs and establish a user interest profile so as to determine the data and websites to be crawled in a targeted manner. The data crawling request information contains various data descriptions related to this crawling submitted by the user. This information may be redundant or incomplete, and direct use may not accurately reflect the user's actual crawling purpose. Therefore, it is necessary to extract key crawling information fragments as a basis for judgment.

[0071] In specific implementation, natural language processing technology can be used to extract key words representing the user's crawling interests, such as "paper", "market conditions", etc. from the data description text requested by the user; at the same time, key crawling parameters representing the crawling time range, data quantity requirements, etc. can also be extracted.

[0072] The platform will then combine the extracted key crawling information and use portrait learning technology to build a user interest crawling portrait. This portrait can reflect the user's crawling intention, such as what category of information needs to be crawled, from which websites to crawl, and how long the time range is.

[0073] Finally, the platform can judge the user's real crawling needs based on keywords and portraits, determine the data content that needs to be crawled and the website where the data is located, so as to formulate a highly targeted crawling plan and use appropriate models to crawl the specified data required by the user, thereby improving the accuracy and effectiveness of data crawling.

[0074] Key crawling information refers to the key words and key parameters related to the crawling task extracted from the data crawling request information submitted by the user using natural language processing technology.

[0075] In the embodiments of the present application, key crawling information can be understood as important crawling vocabulary and crawling parameters extracted from the text description of the crawling request information by the platform backend in order to accurately judge the user's crawling needs. These key information will be used to establish a user interest crawling portrait in order to formulate targeted data crawling plans.

[0076] S102: Determine a first crawling task for the data to be crawled according to data features corresponding to the data to be crawled and website features corresponding to the website to be crawled.

[0077] For the specific implementation process, we first need to analyze the data characteristics such as storage format, organizational structure, field relationship, etc. of different types of data to determine the difficulty of crawling various data and the crawling methods to be adopted. For example, for tabular data with simple and clear structure, crawling can be done by matching its regular structure; while for unstructured text data, crawling is more difficult and requires the use of natural language processing and other technologies to analyze and process the text content.

[0078] At the same time, it is necessary to examine the target website's server performance parameters, network bandwidth, crawler restriction policies, and website data update frequency to determine what frequency, quantity, and other restrictions need to be set for data crawling on the website.

[0079] After comprehensively considering the characteristics of various types of data and the characteristics of the target website, a first crawling task can be formulated, which includes the design of reasonable crawling methods for different data types, and also provides restrictions such as frequency and quantity to minimize adverse effects on the target website. The first task formulated in this way can efficiently and reasonably guide the subsequent crawling of the website data.

[0080] Data features refer to the various attribute information of the data to be captured, including the data's storage format, organizational structure, field relationships, and other characteristics.

[0081] In the embodiments of the present application, data features can be understood as features of the format, organization, relationship, etc. of different types of data to be captured, which are used to determine the difficulty of data capture and design a reasonable capture method accordingly.

[0082] For example, for tabular data, its data features are simple storage format and clear field relationship. This data feature is used to determine whether to crawl by matching its structural relationship.

[0083] Website features refer to various attribute information of the website to be crawled, including website server performance, network bandwidth, anti-crawling strategy, data update frequency and other features. In the embodiment of the present application, website features can be understood as features of server parameters, network bandwidth, restriction strategy, etc. for the target website, which are used to determine the frequency and quantity of data crawling that the website can withstand. For example, if the website server has good performance and sufficient network bandwidth, it can withstand data crawling at a higher frequency and magnitude. On the contrary, if the server performance is average and the bandwidth is limited, it is necessary to set a lower crawling frequency and quantity requirement.

[0084] The first crawling task refers to an initial crawling scheme for crawling the data determined based on the characteristics of the data to be crawled. In the embodiment of the present application, the first crawling task can be understood as a crawling method and mode suitable for the data determined based on the data characteristics such as the storage format and organizational structure of different types of data, for specific crawling of data.

[0085] In an optional implementation, a first data capture mode of the data to be captured is determined according to data features corresponding to the data to be captured;

[0086] Determine a second data crawling mode for the website to be crawled according to website features corresponding to the website to be crawled;

[0087] The first data grabbing mode and the second data grabbing mode are integrated to determine a first grabbing task for grabbing data.

[0088] Specifically, firstly, according to the data characteristics such as storage format and organizational structure of different data to be captured, the difficulty of capturing each type of data is analyzed, and the capture method to be adopted is judged, so as to determine the first data capture mode suitable for this type of data. For example, for tabular data, according to its simple format and clear structure, it can be determined to adopt a capture mode that matches its regular structure.

[0089] Then, based on the website characteristics such as the server performance and network bandwidth of the target website, the crawling frequency and quantity that the website can bear are analyzed to determine the second data crawling mode suitable for the website. For example, for a website with strong performance, a higher crawling frequency can be determined.

[0090] Finally, the first data crawling mode and the second data crawling mode are integrated, and the data characteristics and website characteristics are comprehensively considered to formulate the first crawling task which not only adopts reasonable crawling methods for different data types but also gives frequency limits.

[0091] The first data capture mode refers to an initial capture method and capture mode for capturing data of this type, which is determined based on data characteristics such as the storage format and organizational structure of the data to be captured.

[0092] In the embodiment of the present application, the first data capture mode can be understood as judging the difficulty of capturing different types of data to be captured based on their storage format, structural relationship and other data characteristics, and designing a capture method or capture mode suitable for this type of data for subsequent capture of this type of data. For example, for tabular data, based on its simple format and clear field relationship, the first data capture mode that can match its structural relationship for capture is determined.

[0093] The second data crawling mode refers to crawling conditions such as data crawling frequency and quantity limit determined based on website characteristics such as server performance and network bandwidth of the target website.

[0094] In the embodiment of the present application, the second data crawling mode can be understood as judging the crawling load that the website can bear based on the server parameters, network bandwidth and other characteristics of the website to be crawled, and setting the crawling frequency, order of magnitude and other crawling restrictions suitable for the website accordingly, for subsequent data crawling of the website.

[0095] For example, for a website with strong server performance and sufficient network bandwidth, a second data crawling mode with a higher crawling frequency and magnitude can be determined.

[0096] In an optional implementation, a first initial data capture mode of the data to be captured is determined according to structural features, where the structural features include data format features and nested level features;

[0097] The first initial data grabbing mode is updated according to the dynamic features to obtain the first data grabbing mode for the data to be grabbed, wherein the dynamic features include data storage features, dynamic loading features and content update features.

[0098] First, according to the structural features of the data to be captured, including data format features and nested level features, the data organization and capture difficulty are judged, and the first initial data capture mode for capturing the data is determined. For example, the format of tabular data is simple, and the initial mode of matching structure capture can be used.

[0099] Then, according to the dynamic characteristics of the data to be captured, including data storage characteristics, dynamic loading characteristics and content update characteristics, the data changes are analyzed, and the first initial mode is updated and optimized to obtain the final first data capture mode. For example, if the data is found to change frequently, the capture frequency needs to be increased.

[0100] Structural features refer to various attribute information that express the storage format and organizational mode of the data to be captured. In the embodiments of the present application, structural features can be understood as features that represent the static structural characteristics of the data to be captured, such as the storage format, organizational structure, and nested relationships, and are used to determine the difficulty of capturing data. For example, the structural features of tabular data include a simple format and a clear hierarchy. These structural features can determine that it is easier to capture using a matching format. In this way, the structural features reflect the static organization of the data, and the initial capture mode suitable for the data format can be determined by analyzing the structural features.

[0101] The first initial data capture mode refers to the initial capture method and mode for capturing the data determined based on the static structural features such as the storage format and organizational structure of the data to be captured. In the embodiment of the present application, the first initial data capture mode can be understood as only considering the static structural features such as the data storage format and relationship to determine the initial capture method or mode suitable for this type of data, which is used for the optimization and update of the subsequent capture mode.

[0102] The first data capture mode refers to a comprehensive capture method and mode for capturing such data, which is determined based on the data characteristics such as the storage format and organizational structure of the data to be captured, as well as the dynamic characteristics of the data content changes. In the embodiment of the present application, the first data capture mode can be understood as a method and mode that is designed to be overall suitable for capturing such data based on the consideration of the storage format, relationship characteristics, and dynamic update characteristics of the data, and is used to guide the subsequent capture of this type of data.

[0103] Data format features refer to various attribute information that express the storage format of the data to be captured. In the embodiment of the present application, data format features can be understood as features that represent the static format characteristics of the data to be captured, such as the storage format and organization method, and are used to determine which method is easier to capture the data format. For example, the format feature of tabular data is that the storage method is simple and clear. This format feature can determine that it is easier to capture using a matching format.

[0104] Nested level features refer to various attribute information that express the complexity of the hierarchical relationship of the organizational structure of the data to be captured. In the embodiment of the present application, the nested level features can be understood as features that represent static structural characteristics such as the number of nested levels in the organizational structure of the data to be captured, the relationship between the levels, etc., which are used to determine the number of page traversal operations required to capture the data.

[0105] Data storage characteristics refer to various attribute information that express the storage location and storage method of the data to be captured. In the embodiment of the present application, the data storage characteristics can be understood as characteristics of dynamic storage characteristics such as where the data to be captured is stored on the website and in what format, which are used to determine the location that needs to be accessed and parsed to capture the data. For example, if it is determined that certain data is stored in a database and organized in JSON format, then it is necessary to access the database and parse the JSON when capturing. In this way, the data storage characteristics reflect the dynamic storage situation of the data, and by analyzing the data storage characteristics, the specific location and storage format that need to be accessed and processed for the captured data can be determined.

[0106] Dynamic loading features refer to various attribute information that express whether the data to be captured needs to wait for additional dynamic loading. In an embodiment of the present application, the dynamic loading feature can be understood as a feature that indicates whether the data to be captured uses technologies such as AJAX and JavaScript for dynamic loading, and is used to determine whether it is necessary to wait for additional loading before obtaining the data during the capture process. For example, if it is determined that certain data has a JavaScript dynamic loading feature, then it is necessary to wait for the loading to complete before extracting the data during the capture. In this way, the dynamic loading feature reflects whether the data is dynamically loaded. By analyzing the dynamic loading feature, it can be determined whether the capture process needs to process additional dynamic loading.

[0107] Content update characteristics refer to various attribute information that express the content changes and update frequency of the data to be captured. In the embodiment of the present application, the content update characteristics can be understood as characteristics that represent the frequency or time interval of changes in the content of the data to be captured, and are used to determine at what frequency the capture process needs to be performed to obtain real-time updated data. For example, if it is determined that certain data has a second-level content update feature, a higher capture frequency needs to be set during capture to obtain the latest data. In this way, the content update characteristics reflect the changes in the data content, and by analyzing the content update characteristics, it can be determined how high the capture frequency needs to be set to in order to obtain dynamically changing data.

[0108] In an optional implementation, a server performance score and a data-related score of the website to be crawled are calculated;

[0109] Obtain the limited bandwidth of the website to be crawled, use the server performance score and the data-related score to perform weighted calculation on the limited bandwidth, and obtain the crawling rate of the website to be crawled;

[0110] A second data crawling mode for the website to be crawled is determined according to the crawling rate.

[0111] The server performance score refers to a quantitative indicator after evaluating the processing capacity of the server of the website to be crawled. In the embodiment of the present application, the server performance score can be understood as a quantitative score indicating the performance of the server calculated by using a set scoring mechanism by examining the processor parameters, memory capacity, network bandwidth and other indicators of the website server. The higher the server performance score, the stronger the server performance of the website to be crawled, and the stronger the ability to process data crawling requests; conversely, if the server performance score is low, it means that the server performance is average and the ability to process data crawling is weak. The purpose of determining the website server performance score is to evaluate the intensity of crawling requests that the website server can withstand, so as to reasonably determine the crawling parameters.

[0112] The data relevance score refers to a quantitative indicator after evaluating the relevance and value of the data in the website to be crawled. In the embodiment of the present application, the data relevance score can be understood as a quantitative score representing the relevance and value of the website data calculated by a set scoring mechanism by examining factors such as the update frequency of the website data and the size of the data. The higher the data relevance score, the stronger the data relevance of the website to be crawled, and the greater the value of its data; conversely, if the data relevance score is low, it means that the website data relevance is general and the data value is low. The purpose of determining the website data relevance score is to evaluate the relevance and importance of the website data, so as to give priority to obtaining data with higher relevance and value during the crawling process. Websites with high data relevance scores can appropriately relax the frequency or quantity restrictions when crawling, while websites with low relevance scores can appropriately reduce the crawling intensity to obtain more valuable data.

[0113] The limited bandwidth refers to the maximum network bandwidth that the crawler program is allowed to access and use, which is announced by the website to be crawled or obtained through measurement. In the embodiment of the present application, the limited bandwidth can be understood as the size of the network bandwidth resources that the crawler program is allowed to use to the maximum extent when the website server provides data services to the outside. The limited bandwidth is usually the upper limit of bandwidth usage set by the website to control crawler access. During the data crawling process of the website, the crawler program needs to strictly abide by the setting of its limited bandwidth, control the data transmission rate within the allowed range, and avoid overloading the target website server.

[0114] The crawl rate refers to the data transmission speed of the crawler program when crawling data from the crawled website.

[0115] In the embodiment of the present application, the crawling rate can be understood as an indicator representing the reasonable data transmission rate of the crawler program, which is calculated based on factors such as the server performance score, data value score, and limited bandwidth of the website to be crawled.

[0116] The crawl rate takes into account many factors, including the processing capacity of the website server, the value of the website data, and the upper limit of the bandwidth usage defined by the website. It is a quantitative guide for the transmission speed parameter that the crawler program needs to comply with when crawling data from a specified website. Setting a reasonable crawl rate can not only ensure that the crawler program obtains data efficiently, but also prevent the crawler program from excessively occupying website server resources, thereby keeping the impact of data crawling on the target website within an acceptable range.

[0117] The second data crawling mode refers to the data crawling frequency and order of magnitude requirements of the crawler program determined according to the server performance, network bandwidth and other characteristics of the website to be crawled. In the embodiment of the present application, the second data crawling mode can be understood as the quantitative crawling restriction conditions such as the data crawling frequency and the amount of data crawled each time of the crawler program suitable for the situation of the website are determined based on the evaluation of factors such as the server parameters and network bandwidth of the website to be crawled. The second data crawling mode reasonably sets the crawling frequency and order of magnitude requirements according to the bearing capacity of the website, which is the quantitative crawling constraint condition that the crawler program needs to comply with during the website crawling process. Setting the second data crawling mode can ensure that the crawler program crawls data with an intensity suitable for the situation of the website, and obtains data to the maximum extent without causing unnecessary negative impact on the website.

[0118] In order to reasonably determine the second data crawling mode of the website to be crawled, it is necessary to calculate the server performance score and data-related score of the website, and make a comprehensive judgment based on the website's limited bandwidth to obtain the crawling rate of the website, and finally determine the second crawling mode based on the crawling rate.

[0119] Specifically, we first need to calculate an objective server performance score based on the processor speed, memory capacity, network interface rate and other parameters of the website server to be crawled. The higher the score, the better the server performance. At the same time, we also need to calculate a data-related score based on factors such as the update frequency and data volume of the website data. The higher the score, the more relevant the website data is.

[0120] Then, obtain the publicly announced or measurable limited bandwidth value of the website to be crawled, which is the data transmission bandwidth limited by the website to the crawler program.

[0121] Next, the server performance score and data-related score obtained above are used as weights to perform weighted calculation on the limited bandwidth, that is, the higher the server performance score and data-related score, the greater the weight in the calculation. The calculation result is the crawling rate of the website, that is, the reasonable transmission rate when the crawler crawls data from the website.

[0122] Finally, based on the calculated crawling rate, if the rate is high, a second data crawling mode with a high crawling frequency or order of magnitude can be determined; if the rate is average, a lower crawling frequency or order of magnitude requirement needs to be determined to determine a second data crawling mode suitable for the website situation.

[0123] S103: Decompose the first crawling task to determine second crawling tasks for multiple crawler programs on the time wheel node.

[0124] A crawler program refers to an automated program used to crawl data from a specified website on the Internet. In an embodiment of the present application, a crawler program can be understood as a computer program that runs according to a predetermined logic and mode, and is used to access a website server, parse page content, and extract required text, pictures, videos, and other data. The crawler program simulates user browsing through program code and automatically crawls website data in batches, which is an important technical means to achieve large-scale data collection. The use of multiple crawlers working together can improve the efficiency of data crawling. Setting up different types of crawlers, such as a crawler program that specifically crawls tables, can adopt customized crawling methods according to different data types to improve crawling accuracy.

[0125] The time wheel refers to a scheduling method that assigns tasks to different slots by time. In the embodiment of the present application, the time wheel can be understood as a time-based task scheduling mechanism. By setting multiple wheels to represent different time granularities, tasks are assigned to slots on the time wheel node to achieve the effect of executing tasks at different time points. The time wheel can be used to disperse tasks such as data crawling to different time points for execution, avoiding concentrated execution that causes excessive load on the target website. At the same time, it is also possible to ensure task execution in chronological order and realize coordinated scheduling of crawler programs. By using the time wheel to schedule the crawling tasks of multiple crawler programs, the crawling time can be reasonably set according to the website's bearing capacity, avoiding affecting the target website and improving the orderliness and efficiency of data crawling.

[0126] The second crawling task refers to the specific crawling task assigned to each crawler program after splitting the first crawling task. In the embodiment of the present application, the second crawling task can be understood as the specific crawling behavior requirements that are divided and assigned to different crawlers to be executed on the time wheel according to the overall crawling requirements of the website and data. The second crawling task is more specific and operational than the first crawling task, indicating the specific data volume requirement that a crawler program needs to crawl at a certain point in time. The second crawling task is set to clarify the specific execution requirements of each crawler program and improve the targeting of the crawling.

[0127] In order to better coordinate multiple crawlers to crawl data from the same website and improve crawling efficiency, you proposed a method of decomposing the first crawling task and using a time wheel to determine the second crawling task of each crawler at different time nodes.

[0128] In the specific implementation, according to the amount of crawled data and frequency requirements determined in the first crawling task, the overall task volume will be decomposed according to a certain algorithm, split into multiple partial tasks, and an execution time node will be determined for each partial task. At the same time, according to the characteristics of different types of data, a suitable type of crawler program is matched for each partial task, such as table data matching regular crawlers, text data matching NLP crawlers, etc. Then, using the time wheel method, the second crawling tasks of each crawler program are reasonably planned at different time nodes according to the number of splits, so that they can crawl in a coordinated manner to avoid conflicts or omissions. When the crawler program is actually executed, the task allocation on the time wheel will also be dynamically monitored and corrected to ensure that the overall crawling is carried out as planned. In this way, through the decomposition of the first crawling task and the task scheduling of the time wheel, multiple crawlers can work together in an orderly manner, making the overall data crawling more efficient.

[0129] In an optional implementation, the first crawling task is decomposed into a plurality of second crawling tasks, and a crawler program of a type corresponding to each second crawling task is determined;

[0130] The time wheel is used to perform node task scheduling configuration for each second crawling task, and the second crawling task of each crawler program on the time wheel node is determined.

[0131] First, according to the overall crawling data volume requirement in the first crawling task, it is split into several small tasks according to a certain algorithm. At the same time, according to the characteristics of different types of data, the appropriate crawler program is determined for each small task, such as determining to use a regular expression crawler for the task of crawling table data.

[0132] Next, for each split second crawling task, according to the size of the task and the impact on the website, the corresponding execution time node on the time wheel is reasonably planned to achieve decentralized execution. For example, large tasks are allocated to nodes with sufficient bandwidth during the time period, and small tasks are executed during the time period that is not sensitive to frequency.

[0133] S104: Control each crawler program to crawl the data to be crawled according to the second crawling task.

[0134] After the previous task decomposition and time round scheduling, the second crawling task for each crawler program at different time nodes is obtained. Next, a crawler scheduling control module needs to be established, which can communicate with each crawler program, issue the corresponding second crawling task, and monitor the execution process of the crawler.

[0135] When the crawler program starts, the scheduling control module will first initialize the crawler and issue the corresponding second crawling task on the time wheel, including the type of data, quantity, time node, etc. After receiving the specific second task, the crawler program starts data crawling on the specified website at the required time. At the same time, the scheduling control module monitors the crawler's crawling status in real time, records its progress and compares it with the plan on the time wheel, and sends control instructions for correction when necessary until the second task is completed.

[0136] The following are device embodiments of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0137] See also Figure 2 , which shows a schematic diagram of the structure of a data capture device provided by an exemplary embodiment of the present application. The device can be implemented as all or part of the device through software, hardware or a combination of both.

[0138] An information acquisition module is used to obtain the user's data crawling request information, determine the data to be crawled and the website to be crawled where the data to be crawled is located;

[0139] A first task determination module, used to determine a first crawling task for the data to be crawled according to data features corresponding to the data to be crawled and website features corresponding to the website to be crawled;

[0140] A second task determination module, used to decompose the first crawling task and determine the second crawling tasks of multiple crawler programs on the time wheel node;

[0141] The data crawling module is used to control each crawler program to crawl the data to be crawled according to the second crawling task.

[0142] Optionally, the information acquisition module also includes a portrait unit.

[0143] The portrait unit is used to obtain the user's data crawling request information, extract key crawling information from the data crawling request information; establish a user portrait based on the key crawling information, and determine the data to be crawled and the website to be crawled where the data to be crawled is located.

[0144] Optionally, the first task determination module further includes a task integration unit, a data feature unit and a website feature unit.

[0145] The task integration unit is used to determine the first data crawling mode of the data to be crawled according to the data characteristics corresponding to the data to be crawled; determine the second data crawling mode of the website to be crawled according to the website characteristics corresponding to the website to be crawled; integrate the first data crawling mode and the second data crawling mode to determine the first crawling task for the data to be crawled.

[0146] A data feature unit is used to determine a first initial data crawling mode for the data to be crawled based on structural features, wherein the structural features include data format features and nested level features; the first initial data crawling mode is updated based on dynamic features to obtain a first data crawling mode for the data to be crawled, wherein the dynamic features include data storage features, dynamic loading features and content update features.

[0147] The website feature unit is used to calculate the server performance score and data-related score of the website to be crawled; obtain the limited bandwidth of the website to be crawled, use the server performance score and the data-related score to perform weighted calculation on the limited bandwidth to obtain the crawling rate of the website to be crawled; determine the second data crawling mode of the website to be crawled according to the crawling rate.

[0148] Optionally, the second task determination module further includes a task configuration unit.

[0149] The task configuration unit is used to decompose the first crawling task into multiple second crawling tasks, determine the crawler program of the corresponding type of each second crawling task; use the time wheel to perform node task scheduling configuration on the second crawling task, and determine the second crawling task of each crawler program on the time wheel node.

[0150] Optionally, the data capture module also includes an adjustment unit.

[0151] The adjustment unit is used to adjust the second crawling task of other crawlers if the actual task time of the crawler does not conform to the time of the time wheel node.

[0152] The present application also provides a computer storage medium that can store multiple instructions, which are suitable for being loaded and executed by a processor as described above. Figure 1 A data capture method of the embodiment shown in the figure, the specific execution process can be referred to in Figure 1 The specific description of the illustrated embodiment will not be repeated here.

[0153] See also Figure 3 , is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 3 As shown, the electronic device 300 may include: at least one processor 301 , at least one network interface 304 , a user interface 303 , a memory 305 , and at least one communication bus 302 .

[0154] The communication bus 302 is used to realize the connection and communication between these components.

[0155] The user interface 303 may include a standard wired interface or a wireless interface.

[0156] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).

[0157] Among them, the processor 301 may include one or more processing cores. The processor 301 uses various interfaces and lines to connect various parts in the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 305, and calling data stored in the memory 305. Optionally, the processor 301 can be implemented in at least one hardware form of digital signal processing (Digital Signal Processing, DSP), field programmable gate array (Field-Programmable Gate Array, FPGA), and programmable logic array (Programmable Logic Array, PLA). The processor 301 can integrate one or a combination of a central processing unit (Central Processing Unit, CPU), a graphics processing unit (Graphics Processing Unit, GPU) and a modem. Among them, the CPU mainly processes the operating system, user interface and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 301, and it can be implemented separately through a chip.

[0158] Among them, the memory 305 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read-Only Memory). Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may also be optionally at least one storage device located away from the aforementioned processor 301. As Figure 3 As shown, the memory 305 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program of a data capture method.

[0159] exist Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 301 can be used to call an application program storing a data capture method in the memory 305, which, when executed by one or more processors, enables the electronic device to execute one or more methods in the above-mentioned embodiments.

[0160] An electronic device readable storage medium stores instructions, which, when executed by one or more processors, enable the electronic device to execute one or more methods in the above embodiments.

[0161] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.

[0162] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0163] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0164] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0165] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0166] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: various media that can store program codes, such as USB flash drives, mobile hard drives, magnetic disks or optical disks.

[0167] The above are only exemplary embodiments of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification and practice, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any variation, use or adaptive change of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the technical field not recorded in the present disclosure.

Claims

1. A data capture method, characterized in that: The method comprises: Obtaining data crawling request information from a user, determining the data to be crawled and the website to be crawled where the data to be crawled is located; Determine a first crawling task for the data to be crawled according to data features corresponding to the data to be crawled and website features corresponding to the website to be crawled, wherein the data features include structural features and dynamic features; The determining of the first crawling task for the data to be crawled according to the data features corresponding to the data to be crawled and the website features corresponding to the website to be crawled includes: determining a first data crawling mode for the data to be crawled according to the data features corresponding to the data to be crawled; determining a second data crawling mode for the website to be crawled according to the website features corresponding to the website to be crawled; integrating the first data crawling mode and the second data crawling mode to determine the first crawling task for the data to be crawled; Determining the first data grabbing mode of the data to be grabbed according to the data features corresponding to the data to be grabbed includes: determining the first initial data grabbing mode of the data to be grabbed according to the structural features, the structural features include data format features and nested level features, the first initial data grabbing mode refers to the initial grabbing method and mode for grabbing the data to be grabbed determined according to the static structural features of the data to be grabbed, the static structural features include storage format and organizational structure; updating the first initial data grabbing mode according to the dynamic features to obtain the first data grabbing mode of the data to be grabbed, the dynamic features include data storage features, dynamic loading features and content update features, the first data grabbing mode refers to the comprehensive grabbing method and mode for grabbing such data determined according to the static structural features of the data to be grabbed and the dynamic features of data content changes; Determining the second data crawling mode of the website to be crawled according to the website features corresponding to the website to be crawled includes: calculating the server performance score and data-related score of the website to be crawled, the server performance score refers to a quantitative index after evaluating the processing capacity of the server of the website to be crawled, and the data-related score refers to a quantitative index after evaluating the relevance and value of the data in the website to be crawled; obtaining the limited bandwidth of the website to be crawled, and using the server performance score and the data-related score to perform weighted calculation on the limited bandwidth to obtain the crawling rate of the website to be crawled; and determining the second data crawling mode of the website to be crawled according to the crawling rate; Decomposing the first crawling task to determine second crawling tasks for multiple crawler programs on the time wheel node; Control each of the crawler programs to crawl the data to be crawled according to the second crawling task.

2. The method according to claim 1, characterized in that The step of obtaining the data crawling request information of the user and determining the data to be crawled and the website to be crawled where the data to be crawled is located includes: Obtaining data crawling request information from a user, and extracting key crawling information from the data crawling request information; A user portrait is established based on the key crawled information, and the data to be crawled and the website to be crawled where the data to be crawled is located are determined.

3. The method according to claim 1, characterized in that The step of decomposing the first crawling task to determine the second crawling tasks of multiple crawler programs on the time wheel node includes: Decomposing the first crawling task into a plurality of second crawling tasks, and determining a crawler program of a type corresponding to each of the second crawling tasks; The time wheel is used to perform node task scheduling configuration on each of the second crawling tasks, and the second crawling task of each of the crawler programs on the time wheel node is determined.

4. The method according to claim 1, characterized in that: After controlling each of the crawler programs to crawl the to-be-crawled data according to the second crawling task, the method further includes: If the actual task time of the crawler program does not conform to the time wheel node time, the second crawling task of other crawlers is adjusted.

5. A data capture device, characterized in that: The device comprises: An information acquisition module is used to obtain the user's data crawling request information, determine the data to be crawled and the website to be crawled where the data to be crawled is located; A first task determination module, used to determine a first crawling task for the data to be crawled according to data features corresponding to the data to be crawled and website features corresponding to the website to be crawled, wherein the data features include structural features and dynamic features; The determining of the first crawling task for the data to be crawled according to the data features corresponding to the data to be crawled and the website features corresponding to the website to be crawled includes: determining the first data crawling mode of the data to be crawled according to the data features corresponding to the data to be crawled; determining the second data crawling mode of the website to be crawled according to the website features corresponding to the website to be crawled; integrating the first data crawling mode and the second data crawling mode to determine the first crawling task for the data to be crawled; the determining of the first data crawling mode of the data to be crawled according to the data features corresponding to the data to be crawled includes: determining the first initial data crawling mode of the data to be crawled according to the structural features, the structural features including data format features and nested level features, the first initial data crawling mode refers to the initial crawling method and mode for crawling the data to be crawled determined according to the static structural features of the data to be crawled, the static structural features including storage format and organizational structure; crawling the first initial data according to the dynamic features The mode is updated to obtain the first data crawling mode of the data to be crawled, the dynamic features include data storage features, dynamic loading features and content update features, the first data crawling mode refers to the comprehensive crawling method and mode for crawling such data determined according to the static structural features of the data to be crawled and the dynamic features of data content changes; the second data crawling mode of the website to be crawled is determined according to the website features corresponding to the website to be crawled, including: calculating the server performance score and data related score of the website to be crawled, the server performance score refers to a quantitative index after evaluating the processing capacity of the server of the website to be crawled, and the data related score refers to a quantitative index after evaluating the relevance and value of the data in the crawled website; obtaining the limited bandwidth of the website to be crawled, using the server performance score and the data related score to perform weighted calculation on the limited bandwidth to obtain the crawling rate of the website to be crawled; determining the second data crawling mode of the website to be crawled according to the crawling rate; A second task determination module, used to decompose the first crawling task and determine second crawling tasks of multiple crawler programs on the time wheel node; The data crawling module is used to control each of the crawler programs to crawl the data to be crawled according to the second crawling task.

6. A computer storage medium, characterized in that: The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the method according to any one of claims 1 to 4.

7. An electronic device, characterized in that: It includes a processor, a memory and a transceiver, the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Data acquisition method and device, electronic equipment and storage medium

    CN109829096A

  • Cache updating method, device and equipment and computer readable storage medium

    CN116842047A