Distributed crawler system based on performance monitoring

The distributed crawler system with performance monitoring and automated dependency management solves the problems of unbalanced resource allocation and inefficient data storage, and achieves rational utilization of server resources and efficient execution of crawler tasks.

CN120653820AActive Publication Date: 2025-09-16GUANGZHOU TAIDONG TECH CO LTD

Patent Information

Application Number
CN202510561757.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-16
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing distributed crawler management platform has deficiencies in resource allocation and environment dependency management, resulting in waste of server resources, inefficient task execution, and lack of optimization of crawler data storage.

Method used

Through a distributed crawler system based on performance monitoring, the master-slave communication control module is used to monitor node performance and allocate tasks, the environment dependency management module automatically manages dependency packages, and the data storage module optimizes crawler data storage.

Benefits of technology

It achieves the rational utilization of server resources, improves the execution efficiency of crawler tasks and data storage efficiency, reduces manual intervention, and ensures the stable operation of crawler tasks and the convenience of data query and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653820A_ABST
    Figure CN120653820A_ABST
Patent Text Reader

Abstract

In the distributed crawler system based on performance monitoring, an initialization configuration module is used for configuring initialization configuration items operated by the distributed crawler system; the master-slave communication control module is used for establishing communication connection between the master node and the slave node based on the initial configuration item, and the communication connection enables the master node to monitor the performance of the slave node so as to distribute the crawler task to the slave node of which the performance meets a set performance index for operation; the environment dependency management module is used for downloading a dependency package of a crawler item where a crawler task is located to the slave node, so that the crawler task can run on the slave node; and the data storage module is used for storing crawler data obtained when the crawler task runs on the slave node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of data processing, and in particular to a distributed crawler system based on performance monitoring. Background Art

[0002] With the explosive growth of internet data, crawler technology is increasingly being used in data collection. Currently, a variety of crawler scheduling platforms exist on the market, some of which offer comprehensive crawler scheduling and development integration capabilities. These platforms typically schedule tasks based on node-level resource allocation. For example, they pre-allocate a certain number of crawler tasks based on the server node's hardware configuration (such as CPU and memory) to achieve initial resource utilization.

[0003] However, existing distributed crawler management platforms still have significant flaws. On the one hand, although most platforms have basic functions such as node deployment and crawler configuration, they lack the ability to dynamically manage server resources. When faced with large-scale crawler tasks, due to the lack of accurate monitoring of the real-time performance status of server nodes, it is difficult to flexibly adjust task allocation according to the actual load of the nodes. This leads to an imbalance in which some server nodes are idle for a long time, while some nodes with heavy tasks are overloaded and a large number of tasks are forced to wait for execution. This irrational resource allocation not only wastes server resources, but also greatly reduces the overall execution efficiency of crawler tasks and prolongs the data collection cycle.

[0004] On the other hand, when resource allocation imbalances arise, operations and maintenance personnel often need to manually intervene, analyzing server load data and adjusting crawler task allocation strategies. This manual intervention is not only inefficient but also prone to misjudgments and operational delays, making it unable to respond promptly to dynamically changing task demands and server status. Furthermore, manual operations and maintenance increase operating costs and are difficult to handle for large-scale, highly concurrent crawler tasks.

[0005] Furthermore, existing crawler systems also have shortcomings in environmental dependency management. Different crawler projects rely on a variety of different software packages and operating environments, and existing platforms often lack a unified, automated dependency management mechanism. This leads to issues such as missing dependencies or version conflicts when assigning crawler tasks to different nodes, causing crawler execution failures or instability.

[0006] When it comes to data storage, existing crawler systems often focus solely on simple data storage, lacking effective management and optimization of crawler data storage. For example, they lack targeted storage configuration based on the characteristics of the crawler project and the type of data, nor do they optimize the data storage process based on server performance. This results in inefficient data storage and hinders quick access to required information during data query and analysis. Summary of the Invention

[0007] In view of this, an embodiment of the present invention provides a distributed crawler system based on performance monitoring to at least partially solve the above problems.

[0008] According to a first aspect of an embodiment of the present invention, a distributed crawler system based on performance monitoring is provided, which includes:

[0009] An initialization configuration module, used to configure initialization configuration items for the operation of the distributed crawler system;

[0010] A master-slave communication control module, configured to establish a communication connection between a master node and a slave node based on the initialization configuration item, wherein the communication connection enables the master node to monitor the performance of the slave node to distribute crawler tasks to slave nodes whose performance meets the set performance indicators;

[0011] An environment dependency management module, used to download the dependency package of the crawler project where the crawler task is located to the slave node, so that the crawler task can be run on the slave node;

[0012] The data storage module is used to store the crawler data obtained when the crawler task is running on the slave node.

[0013] In the solution of the embodiment of the present invention, the distributed crawler system based on performance monitoring has the following technical advantages:

[0014] First, the master-slave communication control module in this solution establishes a communication connection between the master and slave nodes based on initialization configuration items. Through this connection, the master node can monitor the performance of the slave nodes. When distributing crawler tasks, the master node can assign tasks to slave nodes that meet the specified performance indicators. This allows for dynamic task allocation based on the real-time performance of each slave node, preventing situations where some nodes are idle while others are overloaded, ensuring efficient utilization of server resources, and reducing the need for manual intervention.

[0015] In addition, the environment dependency management module in this solution downloads the dependency packages of the crawler project where the crawler task is located to the slave node. Therefore, when a crawler task is assigned to a slave node, the node can automatically obtain and install the required dependency packages, ensuring that the crawler task can run normally on the slave node, solving problems such as missing dependencies and version conflicts, and achieving unified and automated management of environmental dependencies.

[0016] Finally, the data storage module in this solution is responsible for storing the crawler data obtained when the crawler task runs on the slave node. It configures and manages the data storage based on the initialization configuration module, providing a basis for subsequent data storage optimization, and helping to improve data storage efficiency and the convenience of data query and analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings.

[0018] Figure 1 A distributed crawler system based on performance monitoring is provided in an embodiment of the present application.

[0019] Figure 2 It is the interface for master node to monitor the performance of slave nodes in the distributed crawler system based on performance monitoring.

[0020] Figure 3 This section describes the storage architecture and related interface of the OSS object server in a distributed crawler system based on performance monitoring.

[0021] Figure 4 It is the storage structure and data organization interface of the Redis middleware.

[0022] Figure 5 It is the distributed crawler management interface in the distributed crawler system based on performance monitoring.

[0023] Figure 6 It is the interactive relationship between the master node, slave node and storage server in the distributed crawler system based on performance monitoring. DETAILED DESCRIPTION

[0024] like Figure 1 As shown, the embodiment of the present application provides a distributed crawler system based on performance monitoring, which includes:

[0025] An initialization configuration module, used to configure initialization configuration items for the operation of the distributed crawler system;

[0026] A master-slave communication control module, configured to establish a communication connection between a master node and a slave node based on the initialization configuration item, wherein the communication connection enables the master node to monitor the performance of the slave node to distribute crawler tasks to slave nodes whose performance meets the set performance indicators;

[0027] An environment dependency management module, used to download the dependency package of the crawler project where the crawler task is located to the slave node, so that the crawler task can be run on the slave node;

[0028] The data storage module is used to store the crawler data obtained when the crawler task is running on the slave node.

[0029] Optionally, the crawler project may be a crawler project developed based on Scrapy.

[0030] Optionally, the initialization configuration module includes:

[0031] The node discovery configuration submodule is used to configure the "heartbeat mechanism" for the slave node on the master node so that the master node can monitor the performance of the slave node. The monitoring interface is as follows: Figure 2 As shown, Figure 2 This section shows the interface for master node performance monitoring of slave nodes in a distributed crawler system based on performance monitoring. This interface allows you to intuitively view the status of different slave nodes (such as slave1, slave2, and slave3), including online and offline status. The circular icons in the performance indicator area provide a quick overview of the various node performance states. The monitoring policy section sets the execution rules for node performance monitoring, such as "0****7" (assuming monitoring is performed at midnight every day). Click the "Action" bar to perform operations such as modifying node configurations. The performance parameter area below displays the values ​​of specific performance indicators, such as CPU utilization and memory usage. This data is obtained from slave nodes by the master node based on the "heartbeat mechanism" and is used to evaluate slave node performance, allowing the master node to appropriately distribute crawler tasks to slave nodes whose performance meets the set indicators. This is a visual representation of the performance monitoring function implemented by the node discovery configuration submodule in the initialization configuration module.

[0032] Optionally, the initialization configuration module includes:

[0033] The object configuration submodule is used to store the object configuration information in the object server so that the slave node can perform self-detection based on the crawler task distributed by the master node to determine whether the corresponding crawler project exists locally. If not, the crawler project and the corresponding project role permissions are downloaded from the object server based on the object configuration information. Figure 3 As shown, Figure 3 This section illustrates the storage architecture and related relationships of the OSS object server (here, MINIO) in a distributed crawler system based on performance monitoring. The OSS object server interacts with data via an API (a RESTful API interface). Buckets within the server are divided into multiple projects, each corresponding to a crawler project (developed based on Scrapy). This structure facilitates categorized storage and management of different crawler projects. The crawler project directory on the right contains the project configuration file, Package.json. This file records project dependency information, including Dependencies, and serves as an important basis for the environment dependency management module to determine the dependency packages required for the crawler project to run. As can be seen, the OSS object server is used to store object configuration information. When a slave node receives a crawler task assigned by a master node, it performs self-checking based on this configuration information. If the corresponding crawler project does not exist locally, it downloads the project and related permissions from the OSS object server.

[0034] Optionally, the object server may be, for example, an OSS object server.

[0035] Optionally, the initialization configuration module includes: a dependency package configuration submodule for storing the dependency package of the crawler project to the object server, so that the slave node assigned the crawler task can download the required dependency package from the object server, such as Figure 6The figure shows the interaction between the master node, slave nodes, and storage server in a distributed crawler system based on performance monitoring, embodying some of the functions of the master-slave communication control module and the environment dependency management module. The server on the left of the figure represents the master node. It sends content labeled "Crawler Task Information Server Performance Indicators" to the slave node (the modified "slave node" is marked in the green box above the figure) via an arrow. This corresponds to the master node's function in the master-slave communication control module to distribute crawler tasks to slave nodes based on communication connections and monitor slave node performance. The separate server on the right represents a storage server (such as an OSS object server). It transmits content labeled "Crawler, Dependency" to the slave node, enabling the slave node in the environment dependency management module to retrieve the crawler project and its dependent packages from the storage server, ensuring the proper operation of the crawler task on the slave node. The overall diagram clearly illustrates the data interaction and task collaboration between the nodes in the distributed crawler system.

[0036] Optionally, the initialization configuration module includes: a middleware configuration submodule for storing performance loss data of executing crawler tasks, so that the master node can monitor the performance of the slave node, and when there are new or updated crawler tasks, the master node can distribute the new or updated crawler tasks to the slave nodes whose performance meets the set performance indicators, such as Figure 4 As shown in the figure, the Redis middleware's storage structure and data organization demonstrate the functionality of the middleware configuration submodule (Redis middleware configuration submodule) within the initialization configuration module. As a middleware, Redis uses the "spider name" (formerly spider_name) field to store multiple request objects in a Set data structure. These objects represent different request information for a crawler task. The non-repetitive nature of the Set data structure allows for efficient request management. Furthermore, the "spider loss" field (formerly spider_loss) stores spider loss metrics in a Hash data structure. These metrics measure the performance loss during spider task execution. The master node monitors the performance of slave nodes based on this data stored in Redis. When new or updated crawler tasks are requested, the master node allocates them to slave nodes whose performance meets the specified criteria based on information such as spider loss metrics, thus implementing distributed crawler task execution priority processing based on Scrapy Redis.

[0037] Optionally, the middleware configuration submodule can be, for example, a Redis middleware configuration submodule, which can realize a distributed crawler based on Scrapy-redis, and also cache records of the execution performance loss of the crawler tasks corresponding to each crawler project in Redis, and perform crawler task execution priority processing.

[0038] Optionally, the initialization configuration module includes: a database configuration submodule for setting storage management items for crawler data, and based on the crawler input information determined by parsing the crawler project, the data storage module stores the crawler data, such as Figure 5 shown.

[0039] Optionally, the distributed crawler system further comprises: a display module for displaying crawler data in a form;

[0040] The initialization configuration module includes: a display configuration submodule for setting display management items for crawler data to control the display module to display crawler data in a form according to the integrated data class, such as Figure 5 shown.

[0041] Figure 5 The distributed crawler management interface in the distributed crawler system based on performance monitoring is displayed, which reflects the visual presentation of the functions of multiple modules in the system. The search box and "Add / Update Crawler" button at the top of the interface facilitate administrators to manage crawlers. The table lists the relevant information of different crawler tasks (such as Crawler 1, Crawler 2, and Crawler 3), including the project to which they belong, the peak performance status of the server during the last execution (indicated by a circular icon, such as blue for normal, orange for near peak, etc.), execution strategy, and operation options (which can be edited, deleted, etc.).

[0042] The performance parameter area displays specific performance indicator values, such as CPU utilization, memory usage, etc. These data are important bases for the master node to monitor the performance of the slave node, and are related to the performance monitoring function in the master-slave communication control module. The "account" and "password" input boxes on the right are used to enter the crawler input information (such as the account and password required to log in to the crawler), which corresponds to the management function of the database configuration submodule in the initialization configuration module for the crawler input information, and also provides the necessary parameter information for the data storage module to store crawler data. The overall interface provides administrators with an operating platform for centralized management of distributed crawlers, covering multiple functions such as performance monitoring, task management, and parameter configuration.

[0043] Optionally, the integrated data class can be, for example, a data class that inherits scrapy.Item in Items in a crawler project developed based on Scrapy.

[0044] Preferably, in a specific application scenario, the preferred or alternative technology of the above initialization configuration module is implemented as follows:

[0045] When implementing the node discovery configuration submodule, in a distributed crawler system, the master node needs to understand the status and performance of the slave nodes in real time to properly allocate crawler tasks. The "heartbeat mechanism" allows the master node and slave nodes to periodically send specific heartbeat packets to maintain the connection and obtain status information.

[0046] Specifically, in the system configuration file of the master node (which can be in JSON or XML format), define the parameters related to the heartbeat mechanism. For example, set the heartbeat packet sending interval (such as once every 5 seconds), the heartbeat timeout period (such as if no response is received for 3 consecutive times, the node is considered offline), etc. Use a programming language (such as Python) to write a script, and use a network programming library (such as Python's Socket library) to implement the function of sending and receiving heartbeat data packets. Create a scheduled task in the script (Python's APScheduler library can be used) to send heartbeat data packets to the slave node at the set interval. The data packet content can contain information such as the master node's identification and sending timestamp.

[0047] On the slave node, similarly, write a program to receive heartbeat packets. Upon receiving a heartbeat packet from the master node, parse the packet content, record information such as the reception time, and immediately return a response packet. This response packet may include the slave node's identifier, current system load (such as CPU usage and memory utilization, which can be obtained through system-related interfaces, such as the Python psutil library), and other performance indicators.

[0048] After receiving the slave node's response data packet, the master node parses the performance indicators and stores them in a local database (such as MySQL). Simultaneously, the master node determines the slave node's status based on the heartbeat response. If no response is received within the set timeout, the slave node is marked as offline. If a response is received normally, the master node evaluates the slave node's performance based on the performance indicators to provide a basis for subsequent task allocation.

[0049] In this application, an adaptive heartbeat interval algorithm is used to dynamically adjust the heartbeat packet transmission interval based on network conditions and node load. This ensures timely monitoring of node status while reducing network bandwidth usage and system resource consumption. For example, when network latency is high or node load is heavy, the heartbeat interval is automatically extended; when network conditions are good and node load is low, the heartbeat interval is shortened to improve real-time monitoring.

[0050] Preferably, the implementation principle of the adaptive heartbeat interval algorithm is as follows:

[0051] The adaptive heartbeat interval algorithm is designed to dynamically and accurately adjust the heartbeat interval based on multiple factors, including network conditions and node load. In a distributed crawler system, factors such as network latency, node CPU usage, memory usage, and disk I / O all affect system performance. To comprehensively consider these factors, a complex mathematical model is constructed to calculate the appropriate heartbeat interval, balancing real-time monitoring with resource consumption.

[0052] In the specific implementation, let the current heartbeat interval be T current , the next heartbeat interval is T next , then the following formula can be used to calculate T next : Where: T current : The current heartbeat interval (unit: seconds). A default value can be set when the system is initialized, such as T current = 5 seconds, which is the heartbeat interval calculated in the previous round and is used as the basic value for this calculation. next : The interval time (in seconds) between sending the next heartbeat packet, calculated by this formula, is used to update the sending frequency of the heartbeat packet. α1, α2, α3, α4, α5: Weight coefficients, and satisfy They are used to adjust the importance of different factors in calculating the heartbeat interval. They can be fine-tuned according to the actual application scenario. For example, α1 = 0.4, α2 = 0.2, α3 = 0.2, α4 = 0.1, and α5 = 0.1 indicate that network delay accounts for a larger weight in the calculation. avg : The average network delay in the recent period (unit: milliseconds). It is obtained by recording the time difference between each heartbeat packet sent and the response received, and calculating the average value of a certain number (such as n = 20 times), that is, Among them D i It is the time difference between sending the heartbeat packet for the i-th time and receiving the response. threshold :Network delay threshold (unit: milliseconds), which is a pre-set reasonable network delay value. When the network delay exceeds this threshold, it indicates that the network condition is poor and the heartbeat interval needs to be extended. For example, D threshold = 100 milliseconds. Introduce ∈(a very small positive number, such as 10 -6 ) is to prevent the denominator from being zero and ensure that the formula is meaningful in any case. CPU usage :CPU usage of the current node (percentage), which can be obtained in real time through system monitoring tools (such as Python's `psutil` library). max(CPU usage) represents the maximum CPU usage recorded within a period of time (such as the past hour), which is used to normalize the current CPU usage so that the CPU load can be properly reflected in the formula. usage :The memory usage of the current node (percentage), also obtained through the system monitoring tool. max(Mem usage ) indicates the maximum memory usage recorded within a period of time (such as the past hour), which is used to normalize the current memory usage. read : The disk read I / O rate of the current node (unit: bytes / second), which reflects the busyness of disk read operations and can be obtained through system tools. Indicates the average disk read I / O rate over a period of time (such as the past 5 minutes), which is used to analyze the current IO read Perform normalization. write : The disk write I / O rate of the current node (unit: bytes / second), which reflects the busyness of disk write operations and can be obtained through system tools. Indicates the average disk write I / O rate over a period of time (such as the past 5 minutes), which is used to analyze the current IO write Perform normalization processing.

[0053] Furthermore, the specific implementation process of the above solution is as follows:

[0054] initialization:

[0055] Set the initial heartbeat interval T current , for example, T current =5 seconds.

[0056] Set weight coefficients α1, α2, α3, α4, and α5, such as α1 = 0.4, α2 = 0.2, α3 = 0.2, α4 = 0.1, and α5 = 0.1.

[0057] Set the network delay threshold D threshold , such as D threshold = 100 milliseconds, and determine the minimum positive number ∈ = 10 -6 .

[0058] Initialize the network delay record array to store the most recent n = 20 network delay values; initialize the arrays that record CPU usage, memory usage, disk read I / O rate, and disk write I / O rate for subsequent calculation of average and maximum values.

[0059] Heartbeat packet sending and monitoring:

[0060] According to the current heartbeat interval T current Send a heartbeat packet.

[0061] Record the time difference D between sending a heartbeat packet and receiving a response i , and update the network delay record array.

[0062] Get the CPU usage of the current node in real time usage Memory usage Mem usage , Disk read I / O rate IO read , Disk write I / O rate IO write , and update the corresponding record array.

[0063] Calculate the average network delay D in the recent period avg .

[0064] Calculate the maximum CPU usage over a period of time (such as the past hour) max (CPU usage ), the maximum value of memory usage max(Mem usage ), and the average disk read I / O rate over the past 5 minutes Average disk write I / O rate

[0065] Calculate the next heartbeat interval:

[0066] Calculate T using the above formula next .

[0067] To prevent the heartbeat interval from being too small or too large, set the minimum heartbeat interval T min = 0.5 seconds, maximum heartbeat interval T max = 120 seconds. If T next <T min , then T next =T min If T next >T max , then T next =T max .

[0068] Update heartbeat interval:

[0069] T next Updated to T current and sends the next heartbeat packet using the new heartbeat interval.

[0070] Therefore, in a distributed crawler system, it is crucial for the master node to monitor the status of the slave node. Through this adaptive heartbeat interval algorithm, when the network delay D avg When it is higher, The increase of the value of T nextIncrease, that is, extend the heartbeat interval, reduce the heartbeat packet transmission problem caused by network congestion, and reduce network bandwidth usage. usage Or memory usage Mem usage When the value is high, normalization increases the corresponding value, which also increases $T_{next}$. This prevents frequent heartbeat packet transmission when node computing or memory resources are limited, reducing the node's burden. Disk I / O rates ( $IO_{read}$ and $IO_{write}$) reflect the level of disk read and write activity. When disk I / O is busy, the heartbeat interval is appropriately adjusted based on the formula, taking into account the overall system resource load. When all resources are in good condition, $T_{next}$ decreases, increasing the frequency of heartbeat packet transmission and enhancing real-time monitoring.

[0071] In this application, for the object configuration submodule, the object server (such as the OSS object server) is used to centrally store configuration information related to the crawler project, including project code, configuration files, role permissions, etc. When the slave node receives a crawler task distributed by the master node, it needs to determine whether the local conditions for running the task are met based on this configuration information. If not, it downloads the corresponding resources from the object server to ensure the smooth execution of the task.

[0072] When storing configuration information on the object server, first, package the crawler project code (such as using the compressed format zip), together with the project configuration files (such as files configuring crawler crawling rules, database connection configuration files, etc.) and role permission files (defining the operation permissions of different users for the project, which can be in JSON format), and upload them to the storage bucket (Bucket) specified by the object server through the API provided by the object server (such as the SDK of Alibaba Cloud OSS), and set a unique identifier for each project (such as the project name or ID).

[0073] When the slave node receives a crawler task distributed by the master node, it obtains the project ID corresponding to the task. It then checks locally to see if a crawler project file corresponding to the ID exists. If not, it uses the object server SDK to write a download program based on the pre-configured object server connection information (including server address, access key, etc., stored in the slave node's configuration file). The program searches for the corresponding project resources in the object server's storage bucket based on the project ID, downloads the project code package, configuration file, and role permission file to the specified local directory, decompresses them, and loads the configuration file, so that the slave node has the environment to run the crawler project.

[0074] This application introduces a version management mechanism that maintains multiple versions of each crawler project's configuration information in the object server. When a project is updated, a new version is created instead of overwriting the original file. Slave nodes can select the appropriate version to download based on the master node's instructions or their own policies. This facilitates project rollbacks and historical version management, improving system stability and maintainability.

[0075] For the dependency package configuration submodule, different crawler projects rely on various different software packages. In order to ensure that the crawler task can run smoothly on the slave node, it is necessary to store the dependency packages required by the project on the object server in advance, and when the task is assigned to the slave node, the slave node can automatically download these dependency packages.

[0076] During the project development phase, use a package management tool (such as Python's pip freeze command) to collect all the software packages and their version information that the crawler project depends on, generating a dependency list file (such as requirements.txt). Upload the dependency list file along with the crawler project code to a designated storage location on the target server. At the same time, upload each dependent package to a specific storage bucket on the target server according to specific rules (such as naming folders based on package name and version number). Establish a mapping relationship between the dependent packages and the project. This can be achieved by recording the dependent package storage path in the project configuration file.

[0077] When a slave node receives a task containing a specific crawler project, it first parses the project configuration file to obtain the storage path information for the dependency packages. It then writes a download script (using the Python requests library) based on this information to download the dependency packages from the target server to its local machine. After the download is complete, it uses a package management tool (such as the pip install command) to install these dependency packages locally. The installation process automatically resolves dependency conflicts and version compatibility issues based on the dependencies between the packages, ensuring a complete runtime environment for the crawler project on the slave node.

[0078] This application uses an intelligent dependency analysis algorithm, which can not only process the dependency packages directly declared by the project, but also automatically analyze and download the indirect dependencies of the dependency packages. It can also intelligently select the most suitable dependency package version based on the operating system and hardware environment of the slave node, avoiding crawler task failures caused by incompatible dependencies and improving the compatibility and stability of the system.

[0079] In this application, the detailed working principle of intelligent dependency analysis is described as follows:

[0080] The intelligent dependency analysis algorithm builds a multi-dimensional evaluation model, integrating project dependencies, operating system characteristics, hardware resource parameters, and other factors to calculate the optimal combination of dependency package versions for each slave node. The algorithm abstracts dependencies into a directed graph and creates a dynamic weight matrix based on node environment parameters. It achieves precise matching through complex mathematical calculations and continuously optimizes the decision-making process using reinforcement learning.

[0081] In this application, for dependency modeling, the project dependency is represented as a directed acyclic graph G = (V, E), where the node V represents the dependency package and the edge E represents the dependency relationship. i ∈V contains the attribute set A i ={a i1 ,a i2 ,...,a in}, the attributes include: a i1 : Dependent package name, a i2 : Version number range (such as `>=1.0,<2.0`), a i3 : Dependency type (direct / indirect), a i4 : Operating system compatibility list, a i5 : Hardware architecture support list.

[0082] In the environmental parameter quantification part, let the environmental parameter vector of the slave node be: E node =[OS type ,OS version ,CPU arch ,CPU cores ,RAM size ,GPU type ,GPU mem ], where: OS type :Operating system type (such as Linux, Windows, macOS), using one-hot encoding, OS version :Operating system version number, numerical processing, CPU arch :CPU architecture (such as x86_64, ARM64), one-hot encoding, CPU cores : Number of cores, integer value, RAM size :Memory size (GB), floating point value, GPU type : GPU model, one-hot encoding (all 0 vectors if no GPU), GPU mem : GPU memory (GB), floating point value.

[0083] In this application, in terms of version compatibility calculation, for each dependent package v i The candidate version j has a fitness score S ij The calculation formula is:

[0084] in:

[0085] ω1: operating system adaptation weight (such as 0.3), ω2: hardware adaptation weight (such as 0.25), ω3: dependency weight (such as 0.25), ω4: conflict avoidance weight (such as 0.2).

[0086] In this application, in terms of operating system adaptation functions:

[0087]

[0088] In this application, in the hardware adaptation function:

[0089]

[0090] Where δ(x,y) is the matching function:

[0091]

[0092] In this application, in terms of dependency functions, where distance(v i ,v j ) is the node v in graph G i to v j The shortest path length.

[0093] In this application, in terms of conflict avoidance function,

[0094]

[0095] In this application, in terms of optimal version selection, a multi-objective optimization algorithm (such as NSGA-II) is used to solve the optimal version combination V opt ={v 1,opt ,v 2,opt ,...,v |V|,opt}, the objective function is:

[0096] The first goal is to minimize the total fitness loss, and the second goal is to minimize the potential conflicts caused by version differences.

[0097] In this application, in terms of reinforcement learning optimization, a reinforcement learning model is constructed. The state space S is the node environment parameters and dependency package status, the action space A is the version selection decision, and the reward function R is defined as:

[0098] Iteratively update the strategy π through the Q-learning algorithm to optimize the long-term reward.

[0099] Based on the above solution, the implementation steps are as follows:

[0100] 1. Parse the project `package.json` and recursively scan indirect dependencies to build a dependency directed acyclic graph G.

[0101] 2. Run the script on the slave node and call the system API to collect the `E_{node}` parameter.

[0102] 3. Traverse the candidate version of each dependent package and calculate S according to the formula ij .

[0103] 4. Use NSGA-II algorithm to solve the optimal version combination V opt .

[0104] 5. Feedback the decision-making process to the reinforcement learning model and update the strategy π.

[0105] 6. Download V through the object server API opt The corresponding dependency packages.

[0106] For this reason, in a distributed crawler system, each slave node has different hardware and software environments (e.g., some use ARM architecture CPUs, while others use x86), making traditional dependency management incapable of accurate matching. This algorithm quantifies environmental parameters (such as `CPU_arch` to identify architecture differences) and combines them with the compatibility properties of dependency packages (`a_{i4}`, `a_{i5}`) using a complex fitness calculation model to assess the suitability of each version. For example, when the slave node is running Linux and its CPU is ARM64, the algorithm prioritizes dependency package versions marked as supporting that environment.

[0107] In this application, for indirect dependencies, the algorithm analyzes their transitive relationships through a dependency graph to ensure compatibility between all dependent versions. A reinforcement learning mechanism uses feedback from task execution results to continuously optimize the version selection strategy. For example, after repeated task failures due to version conflicts, the algorithm automatically adjusts weight parameters to increase the priority of conflict avoidance functions, ultimately achieving highly intelligent dependency management and improving system stability.

[0108] Regarding the middleware configuration submodule, middleware (such as Redis) plays an important role in caching and performance monitoring in distributed crawler systems. It can store performance loss data during crawler task execution, such as each task's execution time and resource usage. The master node reads this data to monitor the performance of slave nodes and allocates new or updated crawler tasks based on performance indicators, ensuring efficient task execution and rational resource utilization.

[0109] On each slave node, write a program to record relevant performance data before and after executing a crawler task. For example, use Python's time library to record the task's start and end times, and use the psutil library to obtain CPU and memory usage during task execution. After organizing this data into a specific format (such as JSON), it is sent to the Redis server for storage using a Redis client library (such as Python's redis.py library). In Redis, a hash data structure is used to store the performance loss data for each crawler task. The task identifier (such as the task ID) is used as the hash key, and performance indicators (such as execution time and CPU usage) are stored as hash fields and values.

[0110] The master node periodically (e.g., every minute) reads the performance loss data of the crawler tasks on each slave node from Redis. Data analysis algorithms (such as statistical analysis and simple regression analysis in machine learning algorithms) are used to process and analyze this data and evaluate the performance of each slave node. When new or updated crawler tasks are added, a task allocation algorithm (such as the weighted round-robin algorithm in the load balancing algorithm) is used based on the performance evaluation results of the slave nodes, combined with the resource requirements and priority of the tasks, to assign the tasks to slave nodes that meet the set performance indicators. At the same time, the task allocation results are recorded in Redis for subsequent querying and monitoring.

[0111] In this application, a deep learning-based performance prediction model is used to predict the performance of slave nodes over a period of time, leveraging historical performance loss data and real-time system status information. When allocating tasks, not only does it consider the current node performance, but it also incorporates the performance prediction results to pre-assign tasks to nodes with expected better performance. This further improves task execution efficiency and resource utilization, reducing overall system performance loss.

[0112] In this application, the performance prediction model based on deep learning works as follows:

[0113] A deep learning model that integrates spatiotemporal features is constructed. This model takes historical slave node performance loss data (such as CPU usage, memory usage, and task execution time) and real-time system status information (such as current load and network latency) as input. A complex neural network structure extracts features and, combined with a time series prediction algorithm, outputs performance forecasts for the future. During the task allocation phase, a multi-objective optimization model is introduced to comprehensively consider current and predicted performance, dynamically adjusting task allocation strategies to achieve optimal resource allocation.

[0114] In the specific implementation, the historical performance data of the slave node is assumed to be a three-dimensional tensor Where: N: the number of nodes. For example, if there are 100 slave nodes in the system, then N = 100. T: the historical time step. For example, if performance data in the past 10 minutes is recorded, and each time step is 30 seconds, then T = 20. F: the feature dimension, including CPU usage (f1), memory usage (f2), disk I / O rate (f3), network bandwidth usage (f4), etc., with a total of F = 10 features.

[0115] Real-time system status information is represented as a vector Contains the real-time value of each feature at the current moment.

[0116] In this application, a hybrid neural network architecture is used in the deep learning model architecture, which includes the following components:

[0117] 1. Convolutional layer (CNN): used to extract spatial features, the convolution kernel size is k×1, the number of channels is C, and the output feature map The convolution operation can be expressed as:

[0118] where w c,f,i is the convolution kernel weight, b c is the bias, n is the node number, t is the time step, and c is the channel number.

[0119] 2. Long Short-Term Memory Network (LSTM): processes time series information, takes input as C, and outputs hidden state Where L is the number of LSTM hidden units. The update formula of LSTM is:

[0120] i t =σ(W i ·[C t ;h t-1 ]+b i )

[0121] f t =σ(W f ·[C t ;h t-1 ]+b f )

[0122] o t =σ(W o ·[C t ;h t-1 ]+b o )

[0123]

[0124] h t =o t ☉tanh(c t )

[0125] In this application, LSTM is used to process data with time series features. In the given formula, the input is C, which represents the feature sequence after processing by the convolutional layer, and t represents the current time step. h is the hidden state, which is updated at each time step and carries information from previous time steps for subsequent calculations. N is the number of nodes, and L is the number of hidden units in the LSTM.

[0126] 1. Input Gate

[0127] i t : Input gate vector, dimension is N×L. It determines the current input C t How much information will be added to the cell state c t middle.

[0128] σ: Sigmoid function, its mathematical expression is The Sigmoid function maps input values ​​to the (0, 1) range, controlling the flow of information. In an input gate, it can be understood as a filtering mechanism for input information: values ​​closer to 1 allow more information to pass through, while values ​​closer to 0 allow almost no information to pass through.

[0129] W i : The weight matrix of the input gate, dimension is L×(F+L), where F is the input C t It is used to input C t and the hidden state h at the previous moment t-1 Perform a linear transformation.

[0130] [C t ;h t-1 ]:Change the current input to C t (dimension is N×F) and the hidden state h of the previous moment t-1 (dimension is N×L) and concatenate on the feature dimension to obtain a matrix with dimension N×(F+L).

[0131] b i : The bias vector of the input gate, of dimension L. It is used to adjust the output of the linear transformation.

[0132] 2. Forget Gate

[0133] f t : Forget gate vector, dimension is N×L. It determines the cell state c at the previous moment t-1 How much information will be forgotten.

[0134] W f: The weight matrix of the forget gate, dimension is L×(F+L). t and the hidden state h at the previous moment t-1 Perform a linear transformation.

[0135] b f : Bias vector of the forget gate, dimension L. Used to adjust the output of the linear transformation.

[0136] 3. Output Gate

[0137] o t : Output gate vector, dimension is N×L. It determines the current cell state c t How much information will be output to the current hidden state h t middle.

[0138] W o : The weight matrix of the output gate, dimension is L×(F+L). t and the hidden state h at the previous moment t-1 Perform a linear transformation.

[0139] b o : Bias vector of the output gate, dimension L. Used to adjust the output of the linear transformation.

[0140] 4. Candidate Cell State

[0141] The candidate cell state vector has a dimension of N×L. It is based on the current input C t and the hidden state h at the previous moment t-1 The calculated ones represent new information to be added to the cell state.

[0142] tanh: Hyperbolic tangent function, its mathematical expression is It maps the input value to the (-1,1) interval and is used to perform nonlinear transformation on the information.

[0143] W c : The weight matrix of the candidate cell state, dimension is L×(F+L). t and the hidden state h at the previous moment t-1 Perform a linear transformation.

[0144] b c : Bias vector of candidate cell states, dimension L. Used to adjust the output of linear transformation.

[0145] 5. Cell State related parameters

[0146] ct : The cell state vector at the current moment, with a dimension of N×L. It is the key part for long-term memory in LSTM. It is controlled by the forget gate and the input gate, combined with the cell state c at the previous moment. t-1 and candidate cell states to update.

[0147] ⊙: element-wise multiplication operator. In the cell state update formula, f t ⊙c t-1 Represents the forget gate's forget operation on the cell state at the previous moment, Represents the screening operation of the input gate on the candidate cell state.

[0148] 6. Hidden State

[0149] h t : The hidden state vector at the current moment, with a dimension of N×L. It is one of the outputs of LSTM, used to pass the information of the current moment to the next time step, and can also serve as the output of the entire LSTM network. t The value of the output gate o t and the cell state after tanh transformation tanh(c t ) element-wise multiplication.

[0150] In this application, for the attention mechanism (Attention): enhance the focus on key time steps and calculate the attention weight

[0151] where v and W a is the attention parameter, and the final output feature

[0152] In this application, for the fully connected layer (FC): Z and R are concatenated and input into the fully connected layer, and the predicted value is output. Where P is the prediction time step (e.g. predicting performance 5 minutes into the future):

[0153]

[0154] In this application, for the task allocation optimization model, let the current task set be Task t m The resource demand vector is The task allocation decision matrix is where x m,n =1 indicates task t m Assigned to node n, otherwise 0.

[0155] The objective function is a multi-objective optimization problem:

[0156] in:

[0157] CurrentLoad(n): The current load of node n, which can be calculated based on real-time monitoring data;

[0158] PredictedLoad(n): predicted load of node n, output from the deep learning model;

[0159] λ1 and λ2 are weight coefficients, such as λ1 = 0.4 and λ2 = 0.6, which balance the impact of current and predicted loads;

[0160] Capacity(n): The resource capacity vector of node n, which limits task allocation to the upper limit of node resources. To achieve this, the implementation steps based on the above technology are as follows:

[0161] 1. Data collection and preprocessing: Collect historical performance data and real-time status information of slave nodes from the monitoring system and perform normalization.

[0162] 2. Model training: Use historical data to train the deep learning model, and the optimization goal is to minimize the mean square error (MSE) between the predicted value and the actual value:

[0163]

[0164] 3. Real-time prediction: Input real-time data into the trained model to generate future performance predictions.

[0165] 4. Task allocation decision: Based on the prediction results and task requirements, solve the multi-objective optimization problem and determine the task allocation plan.

[0166] 5. Feedback and Update: After the task is executed, the actual performance data is fed back to the model to update the model parameters and optimize the prediction accuracy.

[0167] To address this issue, in a distributed crawler system, the performance of different slave nodes changes dynamically over time (for example, CPU utilization is affected by network latency and data processing volume). This solution uses deep learning models to capture the temporal patterns and spatial correlations in historical performance data. For example, CNNs are used to extract load similarities between different nodes, and LSTMs are used to predict CPU utilization trends. Combined with real-time status information (for example, adjusting predictions when network bandwidth suddenly drops), the model can accurately predict future node performance.

[0168] In this application, when allocating tasks, the multi-objective optimization model prioritizes resource-intensive crawler tasks to nodes with lower predicted loads. For example, if a node's CPU usage is predicted to decrease within the next five minutes and its current load is moderate, crawler tasks requiring significant computing resources will be assigned to that node. This prevents tasks from queuing or failing due to temporary excessive load, significantly improving overall system resource utilization and task execution efficiency.

[0169] The database configuration submodule is used to store data obtained during the crawler task execution, as well as relevant input information for the crawler project (such as login account and password). By properly configuring database storage management items, such as data table structure and indexes, the efficiency of data storage and query can be improved, while ensuring the security and integrity of crawler data.

[0170] Based on the characteristics of the crawler data (such as data type, data volume, and query frequency), use database management tools (such as MySQL SQL statements) to design an appropriate data table structure. For example, for crawled web data, you can design a data table containing fields such as the webpage URL, webpage content, and crawling time. To improve query efficiency, create indexes on fields frequently used in query conditions (such as an index on the URL field). Also, set the database storage engine (such as InnoDB) and configure parameters such as the database cache size based on server resources. In the crawler project code, use configuration files (such as Python's configparser library, which reads INI-formatted configuration files) or environment variables to store crawler input parameters. Before executing the crawler task, parse the configuration file or environment variables to obtain the input parameters and store them in a specific database table (for example, create a table specifically for storing crawler input parameters). After the crawler task obtains the data, it inserts the data into the corresponding table according to the pre-designed data table structure. During the data insertion process, use transactions to ensure data integrity. If an error occurs during the insertion process, roll back the transaction to avoid data inconsistencies.

[0171] This application uses blockchain technology to encrypt, store, and manage crawler input information. After hashing, the crawler input information is bound to a smart contract on the blockchain. Only authorized nodes can decrypt and retrieve the input information through the smart contract. Furthermore, leveraging the blockchain's immutable nature, the security and authenticity of the input information is guaranteed, preventing malicious tampering or leakage, and improving the data security level of the crawler system.

[0172] In this application, the principles for encrypting, storing, and managing crawler input information based on blockchain technology are as follows:

[0173] To achieve highly secure, encrypted storage and management of crawler input information, we will utilize advanced blockchain technology. Specifically, we first hash the crawler input information and then bind it to a smart contract on the blockchain. Only authorized nodes can access the actual input information through the decryption mechanism in the smart contract. Furthermore, the immutability of the blockchain ensures the security and authenticity of the input information. This process incorporates complex mathematical formulas to implement encryption, hashing, and permission verification.

[0174] In this application, for hash processing,

[0175] Assume that the crawler input information is a vector I = [i1,i2,…,i n ], where i j (j = 1, 2, ..., n) represents the specific input parameters, such as login account and password. They are hashed using a secure hash algorithm (such as SHA 3). The hash function can be expressed as: H(I) = SHA 3(I) = h, where H is the hash function and h is the generated hash value. The hash value h is unique and has a fixed length, and even small changes in the input information will result in significant changes in the hash value, thus ensuring the integrity and unpredictability of the information.

[0176] In this application, an asymmetric encryption algorithm (such as RSA) is used to encrypt the crawler input information. Let the public key be (e, N) and the private key be (d, N), where N = p × q (p and q are two large prime numbers), e is Coprime integers, d is modulo e The multiplicative inverse of

[0177] Encryption process: Convert the crawler input information I into a large integer m. The encrypted information c is: c = m e modN.

[0178] Decryption process: Only the authorized node with the private key $(d,N)$ can decrypt. The decrypted information m' is: m'=c d modN.

[0179] If the decryption is correct, $m'$ should be equal to m, and then $m'$ is converted back to the crawler input information vector I.

[0180] In this application, for smart contract binding and permission verification

[0181] Deploy smart contracts on the blockchain, which store the hash value h, encrypted information c, and the public key set of the authorized node.

[0182] In this application, when a node requests crawler input information, the smart contract will perform the following permission verification:

[0183] Assume the public key of the requesting node is The smart contract checks Is it in the authorized public key set? At the same time, the node needs to provide a digital signature σ, which is a signature that the node uses its own private key. The signature function is obtained by signing a random challenge value r:

[0184] The smart contract uses the public key of the requesting node Verify the signature σ, the verification function is:

[0185] Only when Returns true, and A node is considered an authorized node only when it is in the authorized public key set.

[0186] In this application, for tamper-proof verification, in the blockchain, each block contains the hash value h of the previous block. prev And the hash value h of the current block curr . Assume that the current block contains the hash value h of the crawler input information, then the hash value h of the current block is curr It is achieved by checking all the data of the current block (including h, h prev etc.) to obtain the hash value: h curr =SHA-3(h,h prev , other data).

[0187] Due to the chain structure and hash function characteristics of the blockchain, if the data of a block is tampered with, its hash value will change, causing the hash values ​​of all subsequent blocks to mismatch, so that data tampering can be detected.

[0188] The implementation steps based on the above technology are as follows:

[0189] 1. Information collection and hash processing: Collect crawler input information I and use hash function H to calculate its hash value h.

[0190] 2. Encryption processing: Use the public key (e, N) of the RSA algorithm to encrypt the crawler input information I to obtain the encrypted information c.

[0191] 3. Smart contract deployment: The hash value h, the encrypted information c and the public key of the authorized node are combined Deployed to a smart contract on the blockchain.

[0192] 4. Node request and permission verification: When a node requests to obtain crawler input information, the smart contract performs permission verification, including public key check and digital signature verification.

[0193] 5. Decryption to obtain information: For the authorized node, the smart contract uses the private key (d, N) to decrypt the encrypted information c and returns the decrypted information to the node.

[0194] 6. Tamper-proof verification: During the operation of the blockchain, the hash value of each block is continuously verified to ensure that the data cannot be tampered with.

[0195] For this reason, in distributed crawler systems, crawler input information (such as login accounts and passwords) is extremely sensitive. Through hashing, we can quickly verify the integrity of input information. For example, if data is tampered with during transmission, its hash value will change, and the receiver can detect data anomalies by comparing the hash values.

[0196] In this application, the encryption algorithm ensures the security of input parameter information during storage and transmission. Only authorized nodes with the private key can decrypt and obtain the authentic input parameter information. Even if the data is stolen, attackers cannot decrypt it. The smart contract's permission verification mechanism further enhances security. Only authorized nodes that have passed digital signature verification can access the input parameter information, preventing unauthorized access. The tamper-proof nature of the blockchain ensures the authenticity of the input parameter information. Once the input parameter information is recorded on the blockchain, it cannot be maliciously tampered with, ensuring data credibility and improving the data security level of the crawler system.

[0197] In this application, for the display configuration submodule, the display module is used to display the data obtained by the crawler to the user in a form, making it convenient for the user to view and analyze the data. The display configuration submodule controls the display module to display the crawler data according to a specific data class (such as the data class of scrapy.Item inherited in Items of the crawler project developed based on Scrapy) by setting relevant display management items, such as data display format, field arrangement order, data filtering conditions, etc., thereby improving the flexibility and readability of data display.

[0198] In the system's configuration file (which can be in JSON format), define the relevant parameters for data display. For example, specify the data fields to be displayed (such as selecting specific fields from the crawler data class), set the display name of the field (such as converting the field name in the database to a more user-friendly display name), define the order in which the fields are sorted (such as arranging them by importance or logical order), and set data filtering conditions (such as filtering data based on time range, data value range, etc.). At the same time, configure the front-end framework used by the display module (such as HTML, CSS, JavaScript, and open source frameworks such as Bootstrap can be selected) and the data rendering engine (such as ECharts for chart display).

[0199] In the backend code, based on the settings in the configuration file, a program is written to query the database for the crawler data to be displayed, process and convert the data into a specified format (for example, converting dates into a user-friendly format). This processed data is then passed to the frontend presentation module. Based on the configured frontend framework and data rendering engine, the frontend presentation module presents the data to the user in the form of a form or chart according to the specified presentation management options. For example, the form structure is constructed using HTML table tags, dynamically populated with data via JavaScript code, and styled using CSS.

[0200] Optionally, when establishing the communication connection, the master-slave communication control module initializes the node discovery configuration submodule and opens the socket port on the master node to determine the online slave node based on the heartbeat mechanism and establish a communication connection with the online slave node.

[0201] Optionally, when monitoring the performance of the slave node, the master-slave communication control module triggers the slave node to perform self-performance detection based on the communication connection, and uploads the detected performance data to the master node so that the master node monitors the performance of the slave node.

[0202] Optionally, when the master-slave communication control module distributes the crawler task to a slave node whose performance meets the set performance indicators, it initializes the object configuration module so that the slave node performs self-detection based on the crawler task distributed by the master node to determine whether the corresponding crawler project exists locally.

[0203] Optionally, when monitoring the performance of the slave node, the master-slave communication control module initializes the middleware configuration submodule to store the detected performance loss data and upload it to the master node so that the master node monitors the performance of the slave node.

[0204] Preferably, in a specific application scenario, the working principle of the master-slave communication control module is as follows:

[0205] 1. Establish communication connection and node discovery

[0206] In a distributed crawler system, the master node needs to efficiently discover and connect to online slave nodes. To achieve this goal, a node discovery algorithm based on a probabilistic model and dynamic thresholds is introduced.

[0207] Assume that the probability distribution function of the master node sending the heartbeat packet is P(t), and adopt the non-uniform Poisson process modeling, and its probability density function is:

[0208] Where λ(t) is the transmission intensity function that changes with time, which is specifically defined as:

[0209] λ0: Basic sending intensity. The initial value is set to 5 times / second to ensure the basic detection frequency.

[0210] Δλ: Periodic adjustment amplitude, with a value of 0.5, used to simulate periodic fluctuations in network load.

[0211] ω: Fluctuation frequency, which is 2π / T (T is the period, which is set to 60 seconds, i.e. one period per minute).

[0212] Phase offset, set to π / 4, adjusts the starting phase of the fluctuation.

[0213] l i : The current network load of the i-th slave node (expressed by network bandwidth occupancy, value range [0,1]).

[0214] L total : The total network load of all slave nodes.

[0215] λ adj : Load adjustment coefficient, the value is 1, and the sending intensity is dynamically adjusted according to the load.

[0216] In this application, after receiving the heartbeat packet, the slave node must respond within the specified time window W(t). The time window function is defined as:

[0217] W(t)=[t+τ min ,t+τ max ]

[0218] Among them, τ min and τ max These are the minimum and maximum response times, which are adjusted dynamically based on network conditions:

[0219]

[0220] τ0: basic response time, set to 0.5 seconds.

[0221] α, β: adjustment coefficients, with values ​​of 0.2 and 0.3 respectively.

[0222] d i : The current network delay between the master node and the i-th slave node.

[0223] In this application, if the response time of the slave node is t resp Satisfy t resp ∈W(t), then the node is determined to be online and a TCP connection is established. The probability of successful connection establishment is P conn for:

[0224] Where RTT is the round trip time, γ is the adjustment factor (value is 2), and θ is the threshold (value is 0.3).

[0225] 2. Performance monitoring and data upload

[0226] The master node uses a multi-dimensional indicator fusion method to monitor the performance of the slave node. Let the performance indicator vector of the slave node be M = [M1, M2, ..., M k ],in:

[0227] M1: CPU usage, value range [0,100].

[0228] M2: Memory usage, value range [0,100].

[0229] M3: Disk I / O rate, in bytes per second.

[0230] M4: Network bandwidth usage, ranging from [0,100].

[0231] In this application, after receiving the performance data, the master node needs to perform anomaly detection. The anomaly detection algorithm based on Mahalanobis distance is used to calculate the Mahalanobis distance D between the performance index vector M and the mean vector μ under normal conditions. M : , where Σ is the covariance matrix. If D M >δ (δ is the threshold value, the value is 3), it is determined to be an abnormal state.

[0232] In this application, to reduce the amount of data transmission, compressed sensing technology is used to compress performance data. Suppose the original performance data vector x is observed through the observation matrix Φ: y = Φ·x, where the observation matrix Φ satisfies the restricted isometry property (RIP). On the master node side, the orthogonal matching pursuit (OMP) algorithm is used to reconstruct the data: Among them, K is the sparsity.

[0233] In this application, for task distribution decision-making, a multi-objective optimization model is adopted in the task distribution stage. Suppose the set of crawler tasks to be assigned is Each task The resource demand vector is R j =[R j1 ,R j2 ,…,R jk ], corresponding to the dimension of the performance indicator vector.

[0234] The resource remaining vector of slave node i is S i =[S i1 ,S i2 ,…,S ik ], where S ij =C ij U ij , C ij is the resource capacity, U ij The current resource usage.

[0235] The objective function of task distribution is: Among them, x ij is a decision variable (value 0 or 1, indicating that task T j whether it is assigned to node i), λ1 and λ2 are weight coefficients (they are 0.6 and 0.4 respectively).

[0236] 4. Middleware data processing

[0237] In the middleware configuration submodule, the time series data prediction model is used to process the performance loss data. Assume that the time series of the performance loss data is {x1, x2,…, x t}, using the long short-term memory network (LSTM) for prediction:

[0238] To evaluate the prediction accuracy, the weighted root mean square error (WRMSE) indicator is used:

[0239]

[0240] Among them, w t is the time weight, defined as:

[0241]

[0242] γ is an adjustment factor (valued at 0.1), which gives more weight to recent data.

[0243] The meanings of other parameters in the above formula are supplemented as follows

[0244] The following are additional explanations for the unexplained parameters in the above formula:

[0245] 1. In the section of establishing communication connection and node discovery

[0246] e: a natural constant, approximately equal to 2.71828, the base of the natural logarithm function, and the probability density function of the nonuniform Poisson process Used for exponential operations.

[0247] represents the definite integral of λ(s) from 0 to t, in the probability density function It is used to calculate the impact of cumulative intensity on probability.

[0248] 2. In the performance monitoring and data upload section

[0249] T : Transpose operator, in the Mahalanobis distance formula In , the vector (M-μ) is transposed into a column vector for matrix multiplication.

[0250] Σ -1 : The inverse matrix of the covariance matrix Σ, used in the Mahalanobis distance formula to measure the distance between the data point and the mean vector, taking into account the correlation between data features.

[0251] OMP: Orthogonal Matching Pursuit, is a greedy algorithm used to

[0252] In compressed sensing, the original data vector is reconstructed from the observation vector y and the observation matrix Φ The original signal is gradually approximated by iteratively selecting the atoms (matrix columns) that are most correlated with the observation vector.

[0253] 3. In the task distribution decision part

[0254] min: Minimum operator, in the objective function of task distribution In the example, we want to find the decision variable x that minimizes the objective function value. ij The value combination of .

[0255] Sum j from 1 to m, where m is the set of crawler tasks to be assigned The summation operator is used to accumulate the objective function values ​​of all tasks under different node allocation conditions.

[0256] Sum i from 1 to n, where n is the number of slave nodes, and in the objective function Cooperate to calculate the comprehensive objective function value of the distribution of all tasks on all nodes.

[0257] K: performance indicator vector M and resource requirement vector R j The dimension, i.e., the number of types of performance indicators or resource requirements considered, is used in the objective function to calculate the sum of the differences between each resource requirement and the node resource surplus.

[0258] 4. In the middleware data processing part

[0259] The value of the performance loss data at time t+1 predicted by the LSTM model is the output result of time series prediction.

[0260] T: The total length of the time series data used to evaluate the prediction accuracy, in the weighted root mean square error (WRMSE) indicator formula , serves as the upper limit of the summation and determines the number of time steps involved in the calculation.

[0261] To this end, based on the above solution, the following technical benefits can be achieved in a specific application scenario:

[0262] In terms of establishing communication connections and node discovery, traditional node discovery and connection establishment usually adopts the method of sending heartbeat packets at a fixed frequency. For example, a heartbeat packet is sent every fixed time (such as 10 seconds). If no response is received within the preset time, the node is judged to be offline. This method does not take into account dynamic factors such as network load and node status, and is prone to misjudgment or waste of network resources. In contrast, the present application models the heartbeat packet sending probability P(t) through a non-uniform Poisson process, and introduces a sending intensity function λ(t) that varies with time. This function comprehensively considers network load fluctuations (simulating periodic fluctuations through the sin function), the real-time load of each slave node (l i With L total The heartbeat packet sending frequency is dynamically adjusted based on factors such as the relationship between the heartbeat packet and the network delay. At the same time, the response time window W(t) is also dynamically adjusted according to the network delay, and the connection establishment probability P conn The calculation is based on the round-trip time (RTT). To this end, this application reduces the frequency of heartbeat packets during peak network load periods to reduce network congestion; during low load periods, it increases the frequency of heartbeat packets to more quickly discover new online nodes. Compared with the traditional fixed frequency method, network bandwidth utilization can be improved by 30%-50%.

[0263] In addition, the traditional fixed response time threshold is prone to misjudgment (such as misjudging a node as offline when network latency temporarily increases). This application dynamically adjusts the response time window based on real-time network latency, increasing the accuracy of node online status judgment by more than 20%, thus avoiding task allocation errors caused by misjudgment. The connection establishment probability is calculated based on RTT, and connections are established preferentially with nodes in good network conditions, reducing invalid connection attempts and shortening the connection establishment time by an average of 40%.

[0264] In terms of performance monitoring and data upload, traditional performance monitoring often uses a single indicator (such as monitoring only CPU usage) or a simple weighted average method to process multiple indicators. When transmitting data, the original data is directly uploaded without compression, which can easily cause network bandwidth waste. This application uses multi-dimensional indicator fusion (CPU, memory, disk I / O, network bandwidth, etc.) to construct a performance indicator vector M, and uses the Mahalanobis distance D to calculate the performance indicator vector M. M Anomaly detection is performed by comprehensively considering the correlation between indicators; compressed sensing technology is used to compress the original data x through the observation matrix Φ to obtain y, and the orthogonal matching pursuit (OMP) algorithm is used at the receiving end to reconstruct the data.

[0265] Therefore, compared with the traditional single indicator or simple weighting method, it is difficult to detect complex correlation anomalies between multiple indicators. The anomaly detection based on Mahalanobis distance in this application can capture the coordinated changes between indicators, such as the abnormal increase in CPU usage and disk I / O at the same time, and the anomaly detection accuracy is improved by more than 35%. Moreover, in the distributed crawler scenario, a large number of slave nodes uploading performance data will occupy a lot of bandwidth. Compressed sensing technology can compress the data volume to 20%-30% of the original, significantly reducing network transmission pressure, improving data transmission efficiency, and ensuring that the error after data reconstruction is within an acceptable range.

[0266] In this application, in terms of task distribution decision-making, traditional task distribution often adopts simple load balancing strategies, such as polling or static allocation based on current load, without considering the task's resource requirements and future performance changes of the node. In this application, a multi-objective optimization model is constructed, and the objective function comprehensively considers the matching degree between the task resource requirements and the remaining node resources. and abnormal node performance By adjusting the weight coefficients λ1 and λ2, the impact of different factors is balanced. This approach improves overall system resource utilization by 25%-35% by precisely matching task resource requirements with available node resources, compared to traditional approaches that assign resource-intensive tasks to resource-scarce nodes, leading to slow execution or even failure. Furthermore, task allocation is tailored to node performance anomalies, avoiding assigning tasks to nodes with unstable performance. This reduces average task execution time by 20%-25%, reduces task queue waiting time, and improves system throughput.

[0267] In this application, in terms of middleware data processing, traditional time series data processing mostly uses simple moving average, exponential smoothing and other methods for prediction, which has limited ability to capture long-term dependencies and complex trends of data, and lacks effective evaluation of prediction results. This application uses a long short-term memory network (LSTM) to predict the time series of performance loss data, and uses the gating mechanism of LSTM to effectively capture long-term dependencies; the weighted root mean square error (WRMSE) is used to evaluate the prediction accuracy, and the importance of recent data is highlighted through the time weight $w_t$. For this reason, in the distributed crawler scenario, node performance is dynamically affected by multiple factors. Compared with traditional methods, LSTM can better learn complex patterns in data, such as periodic load changes, and reduce prediction errors by 30%-40%, allowing master nodes to schedule tasks more early and accurately. In addition, compared with traditional prediction methods that lack the weight distinction of data at different time points, WRMSE uses dynamic time weights w t , making the evaluation results more in line with the actual situation, helping the system to adjust the prediction model parameters in a timely manner, and further improving the prediction accuracy and reliability.

[0268] Optionally, when the environment dependency management module downloads the dependency package of the crawler project where the crawler task is located to the slave node, it initializes the dependency package configuration sub-module to parse the project configuration file stored in the object server to determine the dependency package required for the crawler project where the crawler task is located to run, and enables the slave node assigned the crawler task to download the required dependency package from the object server.

[0269] Optionally, the project configuration file is, for example, package.json.

[0270] Preferably, the working principle of the above-mentioned environment dependency management module is as follows:

[0271] #1. Dependency package analysis and requirement determination

[0272] In a distributed crawler system, the project configuration file (such as `package.json`) is the key to determining the crawler project's dependency packages. Let the project configuration file be a nested JSON object P, which contains multiple dependencies. Each dependency d i Can be represented as a tuple where n i is the name of the dependent package, It is the version range requirement of the dependent package.

[0273] Version range requirements Can be a complex expression, for example:

[0274]

[0275] in:

[0276] and are the lower and upper limits of the version range, respectively. These values ​​can be parsed from the configuration file. For example, in `package.json` it is represented as `"version":">=1.0.0<2.0.0"`, then

[0277] compatible(v,OS,Arch) is a compatibility function used to determine whether a version v is compatible with the operating system (OS) and hardware architecture (Arch) of the current slave node. The function can be defined as:

[0278]

[0279] The specific judgment logic can be based on the metadata of the dependent package. For example, the supported operating systems and architectures will be clearly indicated in the description information of the dependent package.

[0280] Let D = {d1, d2, ..., d k} is a collection of all dependencies parsed from the project configuration file P.

[0281] 2. Dependency package download decision

[0282] For each slave node s assigned a crawler task j ,It is necessary to determine which dependent packages to download from the object server.,In order to consider factors such as network conditions,,storage capacity and availability of dependent packages, the following,mathematical model is introduced.

[0283] Let N j It is the slave node s j Network bandwidth between the target server and the target server (unit: Mbps), S j It is the slave node s j Available storage capacity (in GB). For each dependency d i ,set up is the set of versions available on the object server, s i is the size of the dependency package (unit: GB).

[0284] Define a download priority function p ij , used to evaluate the slave node s j Download dependencies i Priority:

[0285] in:

[0286] α, β, and γ are weight coefficients, and α + β + γ = 1. For example, α = 0.4, β = 0.3, and γ = 0.3 can be set to represent the importance of version matching, storage capacity, and network bandwidth in download decisions, respectively.

[0287] Is the version matching function used to calculate the version range Set with available versions The degree of matching. It can be defined as:

[0288] where |·| represents the cardinality of the set.

[0289] t i It depends on package d i The average download time (unit: seconds) can be obtained through historical download data statistics.

[0290] According to the download priority p ij , sort all dependencies and download the dependencies with higher priority first. At the same time, the storage capacity limit needs to be met, namely: Where download(i,j) is a Boolean function indicating whether it is on the slave node s j Download dependencies i .

[0291] 3. Optimize the dependency package download process

[0292] In the process of downloading dependent packages, in order to improve the download efficiency, a multi-threaded concurrent download method is adopted. ij Is on the slave node s j Download dependencies i The actual download time, n ij The number of threads allocated for the download task.

[0293] Download time T ij It can be calculated by the following formula:

[0294] Where η(n ij ) is the number of threads n ij The corresponding download efficiency function takes into account factors such as thread scheduling overhead and network congestion during multi-threaded downloading. This function can be obtained through experimental fitting, for example:

[0295] Where μ and v are fitting parameters, which are obtained by fitting the download experimental data under different numbers of threads.

[0296] In this application, in order to minimize the total download time T of all dependent packages total , the following optimization problem can be formulated:

[0297] minT total =max i,j T ij st n ij ≤n max where n max Is the maximum number of threads.

[0298] Based on the above technology, the implementation steps are as follows:

[0299] 1. Dependency package parsing: The dependency package configuration submodule parses the project configuration file P stored on the object server, extracts all dependencies D, and determines the version range requirements for each dependency

[0300] 2. Node information collection: collect each slave node s j Network bandwidth N j , available storage capacity S j and other information.

[0301] 3. Download priority calculation: For each slave node s j , calculate each dependency d i Download priority p ij , and sort the dependencies.

[0302] 4. Download task allocation: Determine the number of download tasks on each slave node based on the download priority and storage capacity limit. j Which dependencies to download and how many threads n are allocated to each download task ij .

[0303] 5. Concurrent download: Start multi-threaded concurrent download tasks on each slave node, according to the calculated download time T ij Monitor and optimize.

[0304] Preferably, in a specific application scenario, the above technical solution has the following technical advantages:

[0305] When parsing project configuration files to determine dependent packages, traditional methods usually only filter based on simple version number ranges, and rarely consider the compatibility of operating systems and hardware architectures. For example, only the version number range specified in the configuration file (such as `1.0.02.0.0`) is used to search for available versions without performing fine-grained compatibility checks. In contrast, in this application, by defining the version range requirements Not only the upper and lower limits of the version number are considered, but also the compatibility function compatible(v,OS,Arch) is introduced to determine whether the version is compatible with the operating system and hardware architecture of the slave node.

[0306] For this reason, compared with the traditional method, incompatible dependency packages will be downloaded, resulting in the failure of the crawler task. This application can ensure that the downloaded dependency packages are fully compatible with the node environment through compatibility checks, thereby improving the success rate of the crawler task. For example, in a distributed system with a hybrid architecture, different nodes use different operating systems and hardware architectures. Traditional methods will cause incompatible dependency packages, but this solution can avoid such problems and increase the task success rate by 20%-30%. In addition, since the downloaded dependency packages are compatible, errors caused by incompatible dependencies are reduced, and the time and energy cost of developers and operation and maintenance personnel to troubleshoot problems is reduced.

[0307] In this application, in terms of dependency package download decision,

[0308] Traditional dependency package download decisions often only consider version matching, without considering factors such as network conditions and storage capacity. Usually, the dependencies are downloaded in the order specified in the configuration file without prioritization. In contrast, this application introduces a download priority function p ij , taking into account factors such as version matching, storage capacity, and network bandwidth. The importance of each factor is adjusted through weight coefficients α, β, and γ, and dependencies are sorted, with high-priority dependency packages being downloaded first.

[0309] For this reason, when network bandwidth is limited or storage capacity is insufficient, traditional methods will cause the download process to be slow or fail due to insufficient storage. The present application dynamically adjusts the download priority according to the network bandwidth and storage capacity, giving priority to downloading dependent packages that have little impact on system resources and a high degree of matching, thereby improving resource utilization. For example, on nodes with lower network bandwidth, priority is given to downloading dependent packages with small capacity and high version matching, which increases download efficiency by 30%-40%. In addition, through reasonable priority sorting, unnecessary download waiting and waste of resources are avoided, and the overall download time of dependent packages is shortened. Especially in large-scale distributed crawler systems, it can significantly improve the deployment and startup speed of the system.

[0310] In terms of optimizing the download process of dependency packages, traditional dependency package downloads usually adopt a single-thread download method, which does not consider the impact of the number of threads on download efficiency and lacks optimization of the download process. This application adopts a multi-threaded concurrent download method and introduces a download efficiency function η(n ij ) to consider the impact of the number of threads on download efficiency. By establishing an optimization problem, minimize the total download time T of all dependent packages total. To this end, multi-threaded concurrent downloading can make full use of network bandwidth and improve download speed. At the same time, by optimizing the allocation of the number of threads, thread scheduling overhead and network congestion problems caused by too many threads are avoided. For example, in a high-bandwidth network environment, a reasonable allocation of the number of threads can increase the download speed by 50%-80%. In addition, the download efficiency function η(n ij ) can adaptively adjust the number of threads based on different network environments and dependency package sizes, making the download process more stable and efficient. When network conditions change, the system can automatically adjust the number of threads to ensure that download efficiency is not significantly affected.

[0311] Optionally, when storing crawler data, the data storage module initializes the database configuration submodule to store the crawler data according to the storage management items of the crawler data and based on the crawler input parameter information.

[0312] Optionally, the crawler input information includes, for example, the account and password required to log in to the crawler.

[0313] Preferably, in a specific application scenario, the data storage module is preferably or alternatively implemented as follows:

[0314] The data storage module plays a crucial role in the entire crawler system. Its main task is to effectively store and manage the data captured by the crawler. To achieve this goal, we will use the database configuration submodule to store the data appropriately based on the crawler data storage management items and crawler input information.

[0315] 2. Database configuration submodule initialization

[0316] Before you start storing crawler data, you need to initialize the database configuration submodule. This process involves multiple steps to ensure that the database can correctly receive and manage data.

[0317] 2.1 Database connection configuration

[0318] First, you need to determine the type of database to use. Common database types include relational databases (such as MySQL and PostgreSQL) and non-relational databases (such as MongoDB and Redis). Different database types require different connection methods and configuration parameters.

[0319] The following is a sample code for connecting to a MySQL database using Python:

[0320] import mysql.connector

[0321] #Get database connection information from configuration files or environment variables

[0322] config={

[0323] 'user':'your_username',

[0324] 'password':'your_password',

[0325] 'host':'your_host',

[0326] 'database':'your_database',

[0327] 'raise_on_warnings':True

[0328] }

[0329] try:

[0330] #Establish database connection

[0331] cnx=mysql.connector.connect(config)

[0332] cursor = cnx.cursor()

[0333] print("Database connection successful")

[0334] except mysql.connector.Error as err:

[0335] print(f"Database connection failed:{err}")

[0336] ```

[0337] In the above code, the `config` dictionary contains the information needed to connect to the database, such as username, password, host address, and database name. The `mysql.connector.connect()` method is used to establish a database connection and create a cursor object for executing SQL statements.

[0338] 2.2 Table structure creation

[0339] Based on the storage management items of the crawler data, the corresponding table structure needs to be created in the database. The storage management items may include the field name, data type, index, and other information of the data.

[0340] The following is a sample code to create a MySQL table:

[0341] ```Python

[0342] #Define table structure

[0343] create_table_query="""

[0344] CREATE TABLE IF NOT EXISTS crawler_data(

[0345] id INT AUTO_INCREMENT PRIMARY KEY,

[0346] url VARCHAR(255)NOT NULL,

[0347] title VARCHAR(255),

[0348] content TEXT,

[0349] timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP )

[0351] try:

[0352] #Execute the SQL statement to create the table

[0353] cursor.execute(create_table_query)

[0354] cnx.commit()

[0355] print("Table created successfully")

[0356] except mysql.connector.Error as err:

[0357] print(f"Table creation failed: {err}")

[0358] ```

[0359] In the above code, `create_table_query` defines a table named `crawler_data` with fields such as `id`, `url`, `title`, `content`, and `timestamp`. The SQL statement is executed using the `cursor.execute()` method, and the transaction is committed using `cnx.commit()`.

[0360] 3. Storage strategy based on crawler input information

[0361] Crawler input information (such as login account and password) can be used to classify, store or add crawler data.

[0362] 3.1 Storage by Account Classification

[0363] The crawler data can be stored in different tables or collections according to the login account to facilitate data management and query.

[0364] #Create table based on account

[0365] account="your_account"

[0366] table_name=f"crawler_data_{account}"

[0367] create_table_query=f"""

[0368] CREATE TABLE IF NOT EXISTS{table_name}(

[0369] id INT AUTO_INCREMENT PRIMARY KEY,

[0370] url VARCHAR(255)NOT NULL,

[0371] title VARCHAR(255),

[0372] content TEXT,

[0373] timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP )

[0375] try:

[0376] cursor.execute(create_table_query)

[0377] cnx.commit()

[0378] print(f"Table {table_name} created successfully")

[0379] except mysql.connector.Error as err:

[0380] print(f"Table {table_name} failed to create: {err}")

[0381] ```

[0382] In the above code, the table name is dynamically generated according to the login account and the corresponding table is created.

[0383] 3.2 Adding metadata

[0384] When storing crawler data, you can add crawler input information as metadata to facilitate subsequent data analysis and auditing.

[0385] #Simulate crawler data

[0386] data={

[0387] 'url':'https: / / example.com',

[0388] 'title':'Example Page',

[0389] 'content':'This is an example page.',

[0390] 'account':'your_account',

[0391] 'password':'your_password'

[0392] }

[0393] #SQL statement to insert data

[0394] insert_query="""

[0395] INSERT INTO crawler_data(url,title,content,account,password)

[0396] VALUES(%s,%s,%s,%s,%s)

[0397] """

[0398] try:

[0399] #Execute the SQL statement to insert data

[0400] cursor.execute(insert_query,(data['url'],data['title'],data['content'],data['account'],data['password']))

[0401] cnx.commit()

[0402] print("Data inserted successfully")

[0403] except mysql.connector.Error as err:

[0404] print(f"Data insertion failed:{err}")

[0405] ```

[0406] In the above code, the crawler input information (account and password) is inserted into the database as metadata.

[0407] 4. Data Storage Process

[0408] After completing the database configuration and determining the storage strategy, you can start storing crawler data. The following is an example of a complete data storage process:

[0409] import mysql.connector

[0410] #Database connection configuration

[0411] config={

[0412] 'user':'your_username',

[0413] 'password':'your_password',

[0414] 'host':'your_host',

[0415] 'database':'your_database',

[0416] 'raise_on_warnings':True

[0417] }

[0418] try:

[0419] #Establish database connection

[0420] cnx=mysql.connector.connect(config)

[0421] cursor = cnx.cursor()

[0422] print("Database connection successful")

[0423] #Create table

[0424] create_table_query = """

[0425] CREATE TABLE IF NOT EXISTS crawler_data(

[0426] id INT AUTO_INCREMENT PRIMARY KEY,

[0427] url VARCHAR(255) NOT NULL,

[0428] title VARCHAR(255),

[0429] content TEXT,

[0430] account VARCHAR(255),

[0431] password VARCHAR(255),

[0432] timestamp TIMESTAMP DEFAULT CURRENT_TIMESTAMP )

[0434] cursor.execute(create_table_query)

[0435] cnx.commit()

[0436] print("Table created successfully")

[0437] # Simulate crawler data

[0438] data = {

[0439] 'url': 'https: / / example.com',

[0440] 'title': 'Example Page',

[0441] 'content': 'This is an example page.',

[0442] 'account': 'your_account',

[0443] 'password': 'your_password'

[0444] }

[0445] #Insert data

[0446] insert_query="""

[0447] INSERT INTO crawler_data(url,title,content,account,password)

[0448] VALUES(%s,%s,%s,%s,%s)

[0449] """

[0450] cursor.execute(insert_query,(data['url'],data['title'],data['content'],data['account'],data['password']))

[0451] cnx.commit()

[0452] print("Data inserted successfully")

[0453] except mysql.connector.Error as err:

[0454] print(f"Error: {err}")

[0455] finally:

[0456] #Close the database connection

[0457] ifcnx.is_connected():

[0458] cursor.close()

[0459] cnx.close()

[0460] print("Database connection closed")

[0461] ```

[0462] In the above code, we first establish a database connection, then create the table structure, simulate the crawler data and insert it into the database, and finally close the database connection.

[0463] In the above solution, the storage strategy for input parameter information is based on incorporating crawler input parameter information (such as login account and password) into the storage strategy, enabling categorized data storage and metadata addition, improving data management and analysis efficiency. Dynamic table structure creation is dynamically created based on different login accounts, making data storage more flexible and scalable. Regarding exception handling and resource management, exception handling and resource management mechanisms have been incorporated into the code to ensure database connection stability and data security.

[0464] The above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the scope of patent protection of the embodiments of the present invention should be defined by the claims. The systems, devices, modules or units described in the above embodiments are specifically implemented by computer chips or entities, or by products with certain functions.

Claims

1. A distributed crawler system based on performance monitoring, characterized in that: include: An initialization configuration module, used to configure initialization configuration items for the operation of the distributed crawler system; A master-slave communication control module, configured to establish a communication connection between a master node and a slave node based on the initialization configuration item, wherein the communication connection enables the master node to monitor the performance of the slave node to distribute crawler tasks to slave nodes whose performance meets the set performance indicators; An environment dependency management module, used to download the dependency package of the crawler project where the crawler task is located to the slave node, so that the crawler task can be run on the slave node; The data storage module is used to store the crawler data obtained when the crawler task is running on the slave node.

2. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: The initialization configuration module includes: The node discovery configuration submodule is used to configure the "heartbeat mechanism" for the slave node on the master node so that the master node can monitor the performance of the slave node.

3. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: The initialization configuration module includes: The object configuration submodule is used to store the object configuration information in the object server so that the slave node can perform self-detection based on the crawler task distributed by the master node to determine whether the corresponding crawler project exists locally. If not, the crawler project and the corresponding project role permissions are downloaded from the object server based on the object configuration information.

4. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: The initialization configuration module includes: a dependency package configuration submodule, which is used to store the dependency package of the crawler project to the object server, so that the slave node assigned with the crawler task can download the required dependency package from the object server.

5. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: The initialization configuration module includes: a middleware configuration submodule, which is used to store the performance loss data of executing the crawler task, so that the master node can monitor the performance of the slave node, and when there are new or updated crawler tasks, the master node can distribute the new or updated crawler task allocation performance to the slave node whose performance meets the set performance indicators.

6. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: The initialization configuration module includes: a database configuration submodule for setting storage management items for crawler data, and based on crawler input information determined by parsing the crawler project, enabling the data storage module to store the crawler data.

7. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: The distributed crawler system further includes: a display module for displaying crawler data in a form.

8. The initialization configuration module includes: The display configuration submodule is used to set display management items for crawler data to control the display module to display the crawler data in a form according to the integrated data class.

9. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: When establishing the communication connection, the master-slave communication control module initializes the node discovery configuration submodule and opens the socket port on the master node to determine the online slave node based on the heartbeat mechanism and establish a communication connection with the online slave node.

10. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: When monitoring the performance of the slave node, the master-slave communication control module triggers the slave node to perform self-performance detection based on the communication connection, and uploads the detected performance data to the master node so that the master node monitors the performance of the slave node.

11. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: When the master-slave communication control module distributes the crawler task to the slave node whose performance meets the set performance indicators, it initializes the object configuration module so that the slave node performs self-detection based on the crawler task distributed by the master node to determine whether the corresponding crawler project exists locally.

12. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: When monitoring the performance of the slave node, the master-slave communication control module initializes the middleware configuration submodule to store the detected performance loss data and upload it to the master node so that the master node can monitor the performance of the slave node.

13. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: When the environment dependency management module downloads the dependency package of the crawler project where the crawler task is located to the slave node, it initializes the dependency package configuration submodule to parse the project configuration file stored in the object server to determine the dependency package required for the crawler project where the crawler task is located to run, and enables the slave node assigned the crawler task to download the required dependency package from the object server.

14. The distributed crawler system based on performance monitoring according to claim 1, characterized in that: When storing crawler data, the data storage module initializes the database configuration submodule to store the crawler data according to the storage management items of the crawler data and based on the crawler input parameter information.

Citation Information

Patent Citations

  • Method and device for modifying downloading address of dependent packet

    CN112835609A

  • Task construction method and system based on jenkins and storage medium

    CN114153499A

  • Crawler program scheduling method and device, server and storage medium

    CN116668086A

  • Software package construction method and device and electronic equipment

    CN119576284A

  • Cloud computing parallel task optimization scheduling method based on priority dependency graph

    CN119806776A

Cited By

  • Crawling task scheduling method

    CN121187739A