Management method and device of distributed acquisition system facing supercomputing Internet, and storage medium

By adopting a decentralized management method in the distributed acquisition system of the supercomputing Internet, using heartbeat information and instance list for task management, the high availability problem of the distributed acquisition system in complex environments is solved, fault identification and task transfer are realized, and the system management efficiency and stability are improved.

CN120371614APending Publication Date: 2025-07-25DAWNING INFORMATION IND (BEIJING) CO LTD

Patent Information

Application Number
CN202510376975.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing data acquisition methods are difficult to meet the strict requirements of supercomputing Internet for high availability of distributed acquisition systems. Especially in complex network environments, failures are prone to data interruption and loss, affecting the accuracy and completeness of data processing, and may lead to decision-making errors in scenarios such as high-frequency trading.

Method used

By adopting a decentralized management method in the distributed acquisition system, using the mutual monitoring and management permission acquisition of all acquisition nodes, and combining heartbeat information and instance list for task management, it realizes rapid identification and task transfer of faulty nodes, and avoids management conflicts and data inconsistencies.

Benefits of technology

It improves the management efficiency and stability of the distributed acquisition system, ensures the system's continuity and resource utilization, reduces the risk of single point failure, avoids management logic complexity, and improves the availability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371614A_ABST
    Figure CN120371614A_ABST
Patent Text Reader

Abstract

The invention relates to a management method and device for a distributed collection system facing the supercomputing Internet and a storage medium, and the method comprises the steps: obtaining the heartbeat information of all collection nodes in the distributed collection system from a first database, and obtaining the management authority according to the heartbeat information of all collection nodes; and when it is determined that the management authority is obtained, an instance list corresponding to the distributed acquisition system is obtained from the second database, and acquisition tasks on all the acquisition nodes in the distributed acquisition system are managed according to the heartbeat information of all the acquisition nodes and the instance list. According to the method, in the process of managing the acquisition tasks in the distributed acquisition system by using the method, a decentration effect can be realized by adopting a mode of mutually monitoring all the acquisition nodes, and the management efficiency and stability of the whole system are improved. And each acquisition node performs task management in a manner of obtaining the management authority, so that management conflicts can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a management method, device and storage medium for a distributed acquisition system for supercomputing Internet. Background Art

[0002] In today's digital age, the supercomputing Internet has brought unprecedented development opportunities to many fields such as scientific research, industrial manufacturing, weather forecasting, and financial analysis with its powerful computing power and extensive resource integration characteristics. As various industries continue to rely on the supercomputing Internet, the requirements for its underlying data collection system are becoming more and more stringent. The supercomputing Internet covers a large number of computing nodes, storage devices, and various sensors distributed in different geographical locations with different performance and functions. These diverse components continuously generate massive amounts of data, and the distributed collection system is responsible for collecting this data efficiently and accurately for subsequent analysis, processing, and application. The widespread application of microservice architecture in the supercomputing Internet has further increased the complexity of data collection. Although the microservice architecture, with its modular and loosely coupled advantages, enables the system to respond more flexibly to complex and changing business needs, it also requires the distributed collection system to coordinate more services of different numbers and types to complete data collection tasks. In this distributed environment, the distributed collection system of the supercomputing Internet faces many severe challenges. On the one hand, various unforeseen problems such as network failures, node failures, and software anomalies frequently occur. These failures may cause the interruption of some collection tasks, data loss or incomplete collection, seriously affecting the accuracy and integrity of supercomputing Internet data processing. On the other hand, since supercomputing Internet application scenarios have extremely high requirements for data real-time performance, such as in real-time meteorological monitoring and high-frequency financial transactions, delays or interruptions in data collection may lead to wrong decisions and huge losses. Therefore, how to ensure that the distributed collection system has high availability and can operate continuously, stably, and efficiently in the complex supercomputing Internet environment has become a key issue that needs to be solved urgently.

[0003] Existing data collection methods are difficult to meet the strict requirements of the supercomputing Internet for the high availability of data collection systems. An innovative high-availability method for distributed collection systems for the supercomputing Internet is urgently needed to break through this technical bottleneck. Summary of the invention

[0004] Based on this, it is necessary to provide a management method, device and storage medium for a distributed collection system for supercomputing Internet that can improve the management efficiency of the distributed collection system in response to the above technical problems.

[0005] In a first aspect, the present application provides a management method for a distributed acquisition system for a supercomputer internet, which is applied to any acquisition node in the distributed acquisition system. The method includes:

[0006] Obtain the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, and obtain management authority according to the heartbeat information of all acquisition nodes;

[0007] When it is determined that the management authority is obtained, obtain the instance list corresponding to the distributed acquisition system from the second database, and manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list.

[0008] The management method for the distributed acquisition system for a supercomputer internet provided by the embodiments of the present application obtains the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, obtains management authority according to the heartbeat information of all acquisition nodes, and when it is determined that the management authority is obtained, obtains the instance list corresponding to the distributed acquisition system from the second database, and manages the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list. In the above method, first, in the process of managing the acquisition tasks in the distributed acquisition system by using the above method, the decentralized effect can be achieved by adopting the method of mutual monitoring of all acquisition nodes, which can effectively avoid the disadvantages of the traditional middleware management method, and thus improve the management efficiency and stability of the entire system to a certain extent. For example, the above method can avoid the problem of low management efficiency caused by the complex traditional management logic, and can reduce the risk of system management collapse caused by the single-point failure of the traditional middleware. Second, in the process of mutual monitoring by using multiple acquisition nodes, each acquisition node manages tasks by obtaining management authority, and finally the acquisition node that successfully obtains the management authority manages the tasks, which can avoid the chaotic situation caused by multiple nodes intervening in management at the same time, and can effectively avoid management conflicts and data inconsistency problems. In addition, the acquisition node can effectively manage the acquisition tasks through the heartbeat information in the first database and the instance list in the second database, solving the problem of complex management logic caused by the need to consider the coupling between nodes in the traditional middleware management method. The above method does not need to rely on complex management logic, can further optimize the management process, and thus improve the management efficiency.

[0009] In some embodiments, managing the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list includes:

[0010] Compare the acquisition nodes to which the heartbeat information belongs with the acquisition nodes included in the instance list to obtain a comparison result;

[0011] Manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the comparison result.

[0012] The method described in the embodiments of the present application can quickly discover whether there are faulty acquisition nodes in the distributed acquisition system by comparing the number of acquisition nodes to which the heartbeat information belongs and the number of acquisition nodes in the instance list, improving the management efficiency of the distributed acquisition system.

[0013] In some of these embodiments, managing the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the comparison result includes:

[0014] If the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is inconsistent with the number of acquisition nodes included in the instance list, determine that there are faulty acquisition nodes in the distributed acquisition system, and manage the acquisition tasks on the faulty acquisition nodes and all non-faulty acquisition nodes;

[0015] If the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is consistent with the number of acquisition nodes included in the instance list, determine that there are no faulty acquisition nodes in the distributed acquisition system, and manage the acquisition tasks in the distributed acquisition system according to the heartbeat information of all acquisition nodes.

[0016] For the method described in the embodiments of the present application, for faulty nodes, the remaining acquisition tasks on them can be reallocated to ensure the continuity of task processing. For non-faulty nodes, the allocation of acquisition tasks can be reasonably adjusted according to their load conditions and processing capabilities to ensure that the overall acquisition efficiency of the system is not greatly affected. In the case where there are no faulty acquisition nodes, the system will manage the acquisition tasks in the distributed acquisition system according to the heartbeat information of all acquisition nodes, and preferentially allocate the acquisition tasks to nodes with better performance and lower load, thereby improving the acquisition efficiency and resource utilization rate of the entire system.

[0017] In some of these embodiments, managing the acquisition tasks on the faulty acquisition nodes and all non-faulty acquisition nodes includes:

[0018] Determine the first transfer tasks allocated to each non-faulty acquisition node according to the previous heartbeat information corresponding to the faulty acquisition node and the heartbeat information of each non-faulty acquisition node;

[0019] Update the task tables in the heartbeat information corresponding to each non-faulty acquisition node in the first database according to the first transfer tasks, and each task table is used to instruct each non-faulty acquisition node to process tasks according to the tasks in the corresponding task table.

[0020] The method described in the embodiments of the present application determines the first transfer task by comprehensively considering the heartbeat information of the node before the faulty node and the heartbeat information of each non-faulty node. It can reasonably allocate tasks to each non-faulty node, avoiding the situation where tasks are randomly assigned to a certain non-faulty node, resulting in an overly high load on that node while other node resources are idle. This improves the resource utilization efficiency of the entire system. Moreover, the tasks of the faulty node are quickly transferred to non-faulty nodes. Instead of waiting for the faulty node to be repaired and then resuming data collection, the non-faulty nodes immediately take over through the task transfer mechanism, ensuring that the data collection work of the system will not be interrupted due to the failure of individual nodes, maintaining the continuity of the data collection of the entire system. For application scenarios that rely on real-time data, it can avoid accidents caused by data loss, improving the availability and reliability of the system.

[0021] In some embodiments, after determining that there is a faulty acquisition node in the distributed acquisition system, the method further includes:

[0022] Removing the faulty acquisition node from the instance list to update the instance list.

[0023] In the method described in the embodiments of the present application, after removing the faulty node, when monitoring other nodes, the monitored data can more accurately reflect the operating status of the normal nodes in the system, avoiding the interference of the abnormal data of the faulty node on the monitoring results.

[0024] In some embodiments, managing the acquisition tasks in the distributed acquisition system according to the heartbeat information of all acquisition nodes includes:

[0025] Determining whether there is an overloaded acquisition node among all acquisition nodes according to the heartbeat information of all acquisition nodes;

[0026] If there is, manage the acquisition tasks on the overloaded acquisition node and other acquisition nodes;

[0027] If not, release the management authority.

[0028] In some embodiments, managing the acquisition tasks on the overloaded acquisition node and other acquisition nodes includes:

[0029] Determining the second transfer task assigned to other acquisition nodes according to the heartbeat information of the overloaded acquisition node and the heartbeat information of other acquisition nodes;

[0030] Updating the task tables corresponding to all acquisition nodes in the first database according to the second transfer task.

[0031] In the method described in the embodiments of the present application, since the overloaded acquisition node is prone to performance degradation, response delay, or even crash under high-load operation, by promptly detecting the overloaded node based on the heartbeat information and transferring the excess tasks to other nodes, the load pressure on this node can be effectively reduced, preventing it from crashing due to overwhelming load, thereby ensuring the stable operation of the entire distributed acquisition system. Moreover, the crash of a certain node in the system may trigger a chain reaction, causing the load of other nodes to further increase and even leading to the failure of the entire system. Promptly handling the overloaded node can control the fault risk locally, avoid the spread and contagion of faults, and maintain the overall stability of the system. In addition, by allocating the second transfer task according to the heartbeat information of the overloaded acquisition node and other acquisition nodes, the tasks can be more evenly distributed among the nodes. Transferring some tasks of the overloaded node to the nodes with lighter load can make full use of the idle resources in the system, improve the resource utilization rate of the entire system, and avoid the situation where some nodes are over-stressed while some nodes have idle resources. Finally, task transfer and task list update are performed according to the real-time heartbeat information, realizing the dynamic management of tasks in the distributed acquisition system.

[0032] In some embodiments, obtaining the management authority according to the heartbeat information of all acquisition nodes includes:

[0033] Determining the state of holding the management lock for each acquisition node according to the state information of the management lock in the heartbeat information of each acquisition node;

[0034] If the states of the management locks of all acquisition nodes are in the unheld state, updating the state of its own management lock to the held state to obtain the management authority;

[0035] If there is a target acquisition node with the state of the management lock being in the held state among all acquisition nodes, obtaining the management authority according to the heartbeat information of the target acquisition node.

[0036] In the method described in the embodiments of the present application, in a distributed acquisition system, if multiple nodes attempt to obtain the management authority simultaneously, it may lead to problems such as management instruction conflicts and data inconsistency. By checking the state of the management lock, only when the states of the management locks of all nodes are in the unheld state can a node update the state of its own management lock to the held state and obtain the management authority. This ensures that only one node has the management authority at the same time, avoiding the chaos caused by multiple nodes performing management operations simultaneously and guaranteeing the orderliness of system management.

[0037] In some embodiments, obtaining the management authority according to the heartbeat information of the target acquisition node includes:

[0038] Extracting the lock holding time from the heartbeat information of the target acquisition node;

[0039] If the lock holding time is greater than the preset time threshold, update the status of the management lock of the target acquisition node to the unheld state, and update the status of its own management lock to the held state to obtain management authority;

[0040] If the lock holding time is not greater than the preset time threshold, it is determined that the management authority has not been obtained.

[0041] In the method described in the embodiments of the present application, if the target acquisition node cannot perform management duties normally due to faults, network interruptions, or other abnormal conditions, but the management lock is still in the held state, it will cause the system management to stagnate. By checking the lock holding time, when the lock holding time of the target acquisition node exceeds this threshold, it indicates that the target acquisition node may be abnormal and not suitable for task management currently. Therefore, timely recovering the management authority and reallocating it can effectively avoid this situation and ensure the continuity and stability of system management.

[0042] In a second aspect, the present application also provides a management system for a distributed acquisition system for a supercomputing Internet. The management system includes: a distributed acquisition node cluster, a first database engine, and a second database engine; each acquisition node in the distributed acquisition node cluster is respectively connected to the first database engine and the second database engine; a first database is deployed on the first database engine, and a second database is deployed on the second database engine;

[0043] Any acquisition node in the distributed acquisition node cluster is used to execute the method in any of the embodiments in the first aspect above.

[0044] In a third aspect, the present application also provides a management device for a distributed acquisition system for a supercomputing Internet. The device includes:

[0045] An acquisition module, configured to obtain the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, and obtain management authority according to the heartbeat information of all acquisition nodes;

[0046] A management module, configured to obtain the instance list corresponding to the distributed acquisition system from the second database when it is determined that the management authority has been obtained, and manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list.

[0047] In a fourth aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0048] Obtain the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, and obtain management authority according to the heartbeat information of all acquisition nodes;

[0049] When it is determined that the management permission is obtained, obtain the instance list corresponding to the distributed acquisition system from the second database, and manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list.

[0050] In a fifth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0051] Obtain the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, and obtain the management permission according to the heartbeat information of all acquisition nodes;

[0052] When it is determined that the management permission is obtained, obtain the instance list corresponding to the distributed acquisition system from the second database, and manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list.

[0053] In a sixth aspect, the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0054] Obtain the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, and obtain the management permission according to the heartbeat information of all acquisition nodes;

[0055] When it is determined that the management permission is obtained, obtain the instance list corresponding to the distributed acquisition system from the second database, and manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list.

[0056] The management method, device, and storage medium of the above-mentioned distributed acquisition system for the supercomputer Internet. The method obtains the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, obtains the management authority based on the heartbeat information of all acquisition nodes, and when it is determined that the management authority is obtained, obtains the instance list corresponding to the distributed acquisition system from the second database, and manages the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list. In the above method, first, in the process of managing the acquisition tasks in the distributed acquisition system by using the above method, the decentralized effect can be achieved by adopting the method of mutual monitoring of all acquisition nodes, which can effectively avoid the disadvantages of the traditional middleware management method, and thus improve the management efficiency and stability of the entire system to a certain extent. For example, the above method can avoid the problem of low management efficiency caused by the complex traditional management logic, and can reduce the risk of system management collapse caused by the single-point failure of the traditional middleware. Second, in the process of mutual monitoring by using multiple acquisition nodes, each acquisition node manages tasks by obtaining the management authority, and finally the acquisition node that successfully obtains the management authority manages the tasks, which can avoid the chaotic situation caused by the simultaneous intervention of multiple nodes in management, and can effectively avoid management conflicts and data inconsistency problems. In addition, the acquisition node can effectively manage the acquisition tasks through the heartbeat information in the first database and the instance list in the second database, solving the problem of complex management logic caused by the need to consider the coupling between nodes in the traditional middleware management method. The above method does not need to rely on complex management logic, can further optimize the management process, and thus improve the management efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic structural diagram of the management system of the distributed acquisition system for the supercomputer Internet in some embodiments;

[0058] Figure 2 It is one of the schematic flowcharts of the management method of the distributed acquisition system for the supercomputer Internet in some embodiments;

[0059] Figure 3 It is another schematic flowchart of the management method of the distributed acquisition system for the supercomputer Internet in some embodiments;

[0060] Figure 4 It is still another schematic flowchart of the management method of the distributed acquisition system for the supercomputer Internet in some embodiments;

[0061] Figure 5 It is yet another schematic flowchart of the management method of the distributed acquisition system for the supercomputer Internet in some embodiments;

[0062] Figure 6 It is the fifth flow schematic diagram of the management method for a distributed acquisition system for the supercomputer Internet in some embodiments;

[0063] Figure 7 It is the sixth flow schematic diagram of the management method for a distributed acquisition system for the supercomputer Internet in some embodiments;

[0064] Figure 8 It is the seventh flow schematic diagram of the management method for a distributed acquisition system for the supercomputer Internet in some embodiments;

[0065] Figure 9 It is the eighth flow schematic diagram of the management method for a distributed acquisition system for the supercomputer Internet in some embodiments;

[0066] Figure 10 It is the ninth flow schematic diagram of the management method for a distributed acquisition system for the supercomputer Internet in some embodiments;

[0067] Figure 11 It is the structural block diagram of the management device for a distributed acquisition system for the supercomputer Internet in some embodiments;

[0068] Figure 12 It is the internal structure diagram of a computer device in some embodiments. Detailed implementation manners

[0069] In the embodiments of the present application, the term "and / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0070] In the embodiments of the present application, the term "a plurality of" means two or more, and other quantifiers are similar.

[0071] In the embodiments of the present application, the term "at least one" means one or more. For example, at least one of A, B, and C can represent: A exists alone, B exists alone, C exists alone, A and B exist simultaneously, A and C exist simultaneously, B and C exist simultaneously, and A, B, and C exist simultaneously.

[0072] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0073] In today's digital age, the supercomputing Internet has brought unprecedented development opportunities to many fields such as scientific research, industrial manufacturing, weather forecasting, and financial analysis with its powerful computing power and extensive resource integration characteristics. As various industries continue to rely on the supercomputing Internet, the requirements for its underlying data collection system are becoming more and more stringent. The supercomputing Internet covers a large number of computing nodes, storage devices, and various sensors distributed in different geographical locations with different performance and functions. These diverse components continuously generate massive amounts of data, and the distributed collection system is responsible for collecting this data efficiently and accurately for subsequent analysis, processing, and application. The widespread application of microservice architecture in the supercomputing Internet has further increased the complexity of data collection. Although the microservice architecture, with its modular and loosely coupled advantages, enables the system to respond more flexibly to complex and changing business needs, it also requires the distributed collection system to coordinate more services of different numbers and types to complete data collection tasks. Relying on the supercomputing Internet architecture, the distributed collection system of the supercomputing Internet can efficiently collect, integrate, and preprocess various types of data on multiple geographically dispersed nodes. The system breaks the limitations of the traditional centralized data collection architecture, makes full use of computing resources and data resources distributed in different regions and institutions, and realizes the comprehensive collection of large-scale, multi-source, and heterogeneous data. However, in this distributed environment, the distributed collection system of the supercomputing Internet faces many severe challenges. On the one hand, various unforeseen problems such as network failures, node failures, and software anomalies occur frequently. These failures may cause some collection tasks to be interrupted, data to be lost or incompletely collected, seriously affecting the accuracy and integrity of supercomputing Internet data processing. On the other hand, due to the extremely high requirements of supercomputing Internet application scenarios for data real-time performance, such as in real-time meteorological monitoring, financial high-frequency trading, etc., delays or interruptions in data collection may lead to decision-making errors and huge losses. Therefore, how to ensure that the distributed collection system has high availability in the complex supercomputing Internet environment and can run continuously, stably, and efficiently has become a key issue that needs to be solved urgently. The existing data collection methods are difficult to meet the strict requirements of the supercomputing Internet for the high availability of data collection systems. An innovative high-availability method for distributed collection systems for supercomputing Internet is urgently needed to break through this technical bottleneck.

[0074] In view of this, the embodiments of the present application propose a management method, device and storage medium for a distributed collection system for supercomputing Internet, which can improve the management efficiency of the distributed collection system through decentralized management of the collection node cluster in the distributed collection system.

[0075] It should be noted that the beneficial effects brought about by the embodiments of the present application or the technical problems solved are not limited to this one, but may also include other implicit or related problems. For details, please refer to the description of the following embodiments.

[0076] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the above technical problems. The following several specific embodiments can be combined with each other, and for the same or similar concepts or processes, they may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0077] In some embodiments, the management method of the distributed acquisition system for the supercomputing Internet provided by the embodiments of this application can be applied to, for example, Figure 1 the management system of the distributed acquisition system for the supercomputing Internet as shown. This management system includes a distributed acquisition node cluster 101, a first database engine 102, and a second database engine 103. The distributed acquisition node cluster 101 includes multiple acquisition nodes 1011. Each acquisition node 1011 in the distributed acquisition node cluster 101 is respectively connected to the first database engine 102 and the second database engine 103. Each acquisition node 1011 in the distributed acquisition node cluster 101 can communicate with the first database engine 102 and the second database engine 103 respectively through the network. Each acquisition node can be deployed on a computer device, and this computer device can be a server or a server cluster composed of multiple servers. A first database 1021 is deployed on the first database engine 102, and a second database 1031 is deployed on the second database engine 103. The multiple acquisition nodes 1011 in the distributed acquisition node cluster 101 are used to acquire the working status information in the computing cluster, and specifically execute the management method of the distributed acquisition system for the supercomputing Internet described in any of the following embodiments.

[0078] Those skilled in the art can understand that Figure 1 the structure shown in [the figure] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the management system of the distributed acquisition system for the supercomputing Internet to which the solution of this application is applied. The specific management system of the distributed acquisition system for the supercomputing Internet may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0079] In some embodiments, as Figure 2 shown, a management method of the distributed acquisition system for the supercomputing Internet is provided. Taking the computer device where any acquisition node in the distributed acquisition system shown in [the figure] is located as an example, the method includes the following steps: Figure 1

[0080] S201, obtain the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, and obtain the management authority according to the heartbeat information of all acquisition nodes.

[0081] Among them, the first database is a non-relational database. For example, the first database is a Redis cluster. The first database is used to store the heartbeat information of all collection nodes in the distributed collection system. The heartbeat information includes at least one of the node's load information, the node's collection performance information, the status of the node obtaining the management permission (whether the management permission is obtained), the time when the node obtains the management permission, the node's task list (including the total task volume of the node and the current task processing progress), etc. The status of the heartbeat information includes the status with heartbeat information and the status without heartbeat information. The management permission is the permission to identify nodes and reallocate tasks for all nodes in the distributed collection system, and is the permission to lock the collection process of the distributed collection system. For example, the management permission can be a distributed lock. It should be noted that in the distributed collection system, the timeliness of the collection system to process tasks is crucial, which directly affects the timeliness of data and the overall response ability of the system. And the latest status of each collection node plays a key role in ensuring the timely processing of tasks. Only by accurately grasping the real-time status of each node can tasks be reasonably allocated, and potential problems be discovered and solved in a timely manner. As an effective way to reflect the latest status of each node, the heartbeat information can clearly show whether the node is running normally and whether it has the ability to process tasks. In order to efficiently store and manage this heartbeat information, Redis is selected as the first database in the embodiments of the present application. Redis has excellent read and write performance and high concurrency processing ability, and can quickly store and read heartbeat information to ensure that the system can obtain the node status in real time. At the same time, the in-memory storage feature of Redis makes the data access speed extremely fast, further improving the system's response speed to node status changes. By storing the heartbeat information in Redis, real-time monitoring and rapid response to the status of each collection node can be achieved, and then the task allocation can be dynamically adjusted according to the node status, effectively ensuring the timeliness of the collection system to process tasks, and improving the operation efficiency and stability of the entire system.

[0082] In the embodiment of the present application, the first database can be deployed on the first database engine in advance, and the first database engine is communicatively connected to all nodes in the distributed acquisition system. When any acquisition node in the distributed system manages the acquisition task, the server where any acquisition node is located can regularly (for example, every few seconds or minutes) obtain the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database through the first database engine by means of a monitoring tool or a script written, and then extract the management permissions recorded by each acquisition node from the heartbeat information. Then, according to the management permission status of other nodes, it is determined whether the current management permission of the acquisition system is idle. If the current management permission of the acquisition system is idle, the server can obtain the management permission and record the management permission in the heartbeat information of its own acquisition node for other nodes to view the management permission status in real time; if the current management permission of the acquisition system is not idle, it means that a certain acquisition node in the acquisition system has obtained the management permission and is managing the acquisition task of the distributed acquisition system, so the server does not obtain the management permission this time. Then, after a preset time interval, the server can loop through the steps of obtaining the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database again and obtaining the management permission according to the heartbeat information of all acquisition nodes. Optionally, to ensure that the acquisition node that obtains the management permission manages the task normally, when the server determines that a certain acquisition node has obtained the management permission, it can further determine whether the acquisition node is abnormal according to the heartbeat information of the acquisition node. If the acquisition node is abnormal, the server can take over the management permission on the abnormal acquisition node to obtain the management permission, unlock the management permission of the abnormal acquisition node, and update the status of the management permission in the heartbeat information of the abnormal acquisition node.

[0083] S202. When it is determined that the management permission is obtained, obtain the instance list corresponding to the distributed acquisition system from the second database, and manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list.

[0084] Among them, the second database is a relational database. For example, the second database is MySQL or PostgreSQL. The second database is used to store the instance list corresponding to the distributed acquisition system. The instance list stores the identification information of all acquisition nodes currently working in the distributed acquisition system. When a new acquisition node joins the distributed acquisition system, the instance list will be updated. For example, the instance list stores the ID of acquisition node 1, the ID of acquisition node 2... the ID of acquisition node n. It should be noted that the instance list corresponding to the distributed acquisition system essentially belongs to the stored relational information. In the embodiments of the present application, a relational database is selected as the second database, and they can store data on the disk in a structured manner. Although the read and write speed may be slower than Redis in some scenarios, for data such as the instance list that does not need to be refreshed in real time, it can fully meet the requirements. Moreover, the cost of disk storage is much lower than that of memory storage. Using a relational database to store the instance list can greatly save storage costs on the premise of ensuring the basic functions of the system, and make resources more reasonably allocated and utilized.

[0085] In the embodiments of the present application, the second database can be pre-deployed on the second database engine, and the second database engine is communicatively connected to all nodes in the distributed acquisition system. After a server where a certain acquisition node in the distributed acquisition system obtains the management authority based on the above steps, it can start the acquisition node identification and failover operations in the distributed acquisition system. Specifically, the server can obtain the instance list corresponding to the distributed acquisition system from the second database through the second database engine, and then determine whether there are faulty nodes in the distributed acquisition system according to the consistency between the acquisition nodes in the heartbeat information of all acquisition nodes and the acquisition nodes in the instance list. If it is determined that there are no faulty nodes in the distributed acquisition system, then the supervision of the server this time ends, and the management authority can be returned; if it is determined that there are faulty nodes in the distributed acquisition system, the faulty nodes can be identified, and the acquisition tasks on the faulty nodes can be evenly distributed to other acquisition nodes. When evenly distributing the acquisition tasks on the faulty nodes to other acquisition nodes, the acquisition tasks on the faulty nodes can be evenly distributed to the acquisition nodes near the faulty nodes. Optionally, the acquisition tasks on the faulty nodes can be assigned to other acquisition nodes with good acquisition performance. Optionally, the acquisition tasks on the faulty nodes can be assigned to other acquisition nodes with low acquisition load. Optionally, multiple nodes with better acquisition performance and / or lower load can be selected from all normal nodes, and then from the multiple nodes with better acquisition performance and / or lower load, the acquisition nodes near the faulty nodes can be selected to inherit the acquisition tasks on the faulty nodes for management. Among them, nodes with an acquisition rate greater than the rate threshold can be regarded as nodes with better acquisition performance, and nodes with a load less than the load threshold can be regarded as nodes with lower load.

[0086] The management method of the distributed acquisition system for the supercomputer Internet provided by the embodiments of the present application obtains the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, obtains the management authority according to the heartbeat information of all acquisition nodes, and when it is determined that the management authority is obtained, obtains the instance list corresponding to the distributed acquisition system from the second database, and manages the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list. In the above method, first, in the process of managing the acquisition tasks in the distributed acquisition system by using the above method, the decentralized effect can be achieved by adopting the method of mutual monitoring of all acquisition nodes, which can effectively avoid the disadvantages of the traditional middleware management method, and thus improve the management efficiency and stability of the entire system to a certain extent. For example, the above method can avoid the problem of low management efficiency caused by the complex traditional management logic, and can reduce the risk of system management collapse caused by the single point of failure of the traditional middleware. Secondly, in the process of mutual monitoring by using multiple acquisition nodes, each acquisition node manages tasks by obtaining the management authority, and finally the acquisition node that successfully obtains the management authority manages the tasks, which can avoid the chaotic situation caused by multiple nodes intervening in management at the same time, and can effectively avoid management conflicts and data inconsistency problems. In addition, the acquisition node can effectively manage the acquisition tasks through the heartbeat information in the first database and the instance list in the second database, solving the problem of complex management logic caused by the need to consider the coupling between nodes in the traditional middleware management method. The above method does not need to rely on complex management logic, can further optimize the management process, and thus improve the management efficiency.

[0087] In some embodiments, a specific method for managing the acquisition tasks on all acquisition nodes in the distributed acquisition system is also provided, such as Figure 3 shown, the "managing the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list" in S202 above includes:

[0088] S301, comparing the acquisition nodes to which the heartbeat information belongs with the acquisition nodes included in the instance list to obtain a comparison result.

[0089] Among them, the comparison result includes that the number of acquisition nodes to which the heartbeat information belongs is the same as the number of acquisition nodes included in the instance list, or the comparison result includes that the number of acquisition nodes to which the heartbeat information belongs is different from the number of acquisition nodes included in the instance list.

[0090] In the embodiments of the present application, after the server obtains the heartbeat information and instance list of all the acquisition nodes, it can further determine whether the number of acquisition nodes to which the heartbeat information belongs is consistent with the number of acquisition nodes included in the instance list, and determine whether there are acquisition nodes that are in the instance list but have no heartbeat information.

[0091] S302. Manage the acquisition tasks on all the acquisition nodes in the distributed acquisition system according to the comparison result.

[0092] In the embodiments of the present application, after the server obtains the comparison result based on the above steps, it can determine whether there are faulty acquisition nodes according to the comparison result. If the number of acquisition nodes to which the heartbeat information belongs is consistent with the number of acquisition nodes included in the instance list, it means that all the acquisition nodes included in the instance list have heartbeat information. If the number of acquisition nodes to which the heartbeat information belongs is inconsistent with the number of acquisition nodes included in the instance list, it means that there are nodes in the acquisition nodes included in the instance list that have no heartbeat information.

[0093] The method described in the embodiments of the present application can quickly discover whether there are faulty acquisition nodes in the distributed acquisition system by comparing the number of acquisition nodes to which the heartbeat information belongs and the number of acquisition nodes in the instance list, and improve the management efficiency of the distributed acquisition system.

[0094] In some embodiments, a specific implementation manner for managing the acquisition tasks on all the acquisition nodes in the distributed acquisition system is also provided. For example, Figure 4 as shown, the "managing the acquisition tasks on all the acquisition nodes in the distributed acquisition system according to the comparison result" in the above S302 includes:

[0095] S401. If the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is inconsistent with the number of acquisition nodes included in the instance list, determine that there are faulty acquisition nodes in the distributed acquisition system, and manage the acquisition tasks on the faulty acquisition nodes and all the non-faulty acquisition nodes.

[0096] Among them, the faulty acquisition node refers to the acquisition node without heartbeat information.

[0097] In the embodiments of the present application, when the server obtains the comparison result based on the above steps and determines that the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is inconsistent with the number of acquisition nodes included in the instance list, it determines that there are faulty acquisition nodes in the distributed acquisition system. Then the server can reassign the uncompleted acquisition tasks on the faulty acquisition nodes to the non-faulty acquisition nodes, and manage all the acquisition tasks on the non-faulty acquisition nodes (including the previous acquisition tasks and the newly assigned acquisition tasks).

[0098] S402. If the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is the same as the number of acquisition nodes included in the instance list, it is determined that there are no faulty acquisition nodes in the distributed acquisition system, and the acquisition tasks in the distributed acquisition system are managed based on the heartbeat information of all acquisition nodes.

[0099] In the embodiment of the present application, when the server obtains the comparison result based on the above steps and determines that the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is the same as the number of acquisition nodes included in the instance list, it is determined that there are no faulty acquisition nodes in the distributed acquisition system. Then, it can be considered that the server's supervision ends this time, and the management authority is returned. After a preset interval, the server returns to execute step S201 to manage the acquisition tasks in the distributed acquisition system based on the heartbeat information of all acquisition nodes.

[0100] For the method described in the embodiment of the present application, for faulty nodes, the remaining acquisition tasks on them can be redistributed to ensure the continuity of task processing. For non-faulty nodes, the distribution of acquisition tasks can be reasonably adjusted according to their load conditions and processing capabilities to ensure that the overall acquisition efficiency of the system is not greatly affected. In the case where there are no faulty acquisition nodes, the system manages the acquisition tasks in the distributed acquisition system based on the heartbeat information of all acquisition nodes, and preferentially assigns the acquisition tasks to nodes with better performance and lower load, thereby improving the acquisition efficiency and resource utilization rate of the entire system.

[0101] In some embodiments, a specific implementation manner for managing acquisition tasks in the case where the number of acquisition nodes to which the heartbeat information belongs is different from the number of acquisition nodes included in the instance list is also provided. For example Figure 5 As shown, the step of "determining the first transfer tasks assigned to each non-faulty acquisition node according to the previous heartbeat information corresponding to the faulty acquisition node and the heartbeat information of each non-faulty acquisition node" in the above S401 includes:

[0102] S501. Determine the first transfer tasks assigned to each non-faulty acquisition node according to the previous heartbeat information corresponding to the faulty acquisition node and the heartbeat information of each non-faulty acquisition node.

[0103] Among them, the previous heartbeat information is the last heartbeat information of the faulty acquisition node before the fault occurs. The first transfer task is the task that the non-faulty acquisition node needs to take over from the faulty acquisition node.

[0104] In the embodiments of the present application, when the server determines that there is a faulty collection node in the collection node cluster of the distributed collection system, it can obtain the last heartbeat information of the faulty collection node before the fault as the previous heartbeat information from the first database, and obtain the heartbeat information of all non-faulty collection nodes. Then, the collection tasks that the faulty collection node has not processed are extracted from the previous heartbeat information. Then, according to at least one of the node load information, collection performance information, collection performance information, etc. carried in the heartbeat information of each non-faulty collection node, the collection tasks on the faulty collection node are assigned to appropriate non-faulty collection nodes to obtain the first transfer tasks of the non-faulty collection nodes. Specifically, the collection tasks on the faulty collection node can be assigned to other non-faulty collection nodes with low collection load according to the node load information carried in the heartbeat information of each non-faulty collection node. Optionally, the collection tasks on the faulty collection node can be assigned to other non-faulty collection nodes with better collection performance according to the collection performance information carried in the heartbeat information of each non-faulty collection node. Optionally, the collection tasks on the faulty collection node can be assigned to non-faulty collection nodes near the faulty collection node according to the collection performance information of each non-faulty collection node. Among them, a node with a collection rate greater than the rate threshold is a node with better collection performance, and a node with a load less than the load threshold is a node with a lower load.

[0105] S502, update the task table in the heartbeat information corresponding to each non-faulty collection node in the first database according to the first transfer task. Each task table is used to instruct each non-faulty collection node to process tasks according to the tasks in the corresponding task table.

[0106] Among them, the task table records the collection tasks to be processed and the processing progress of the collection tasks.

[0107] In the embodiments of the present application, after the server obtains the first transfer tasks of each non-faulty collection node based on the above steps, it can update the task table of each non-faulty collection node from the heartbeat information of each non-faulty collection node in the first database, which is used to instruct each non-faulty collection node to process tasks according to the tasks in the corresponding task table. For example, the first transfer task can be added to the collection tasks to be processed in the original task table of each non-faulty collection node.

[0108] In the method described in the embodiments of the present application, by comprehensively considering the heartbeat information of the node before the faulty node and the heartbeat information of each non-faulty node to determine the first transfer task, tasks can be reasonably allocated to each non-faulty node, avoiding the situation where tasks are randomly allocated to a certain non-faulty node, resulting in an overly high load on that node while other node resources are idle, and improving the resource utilization efficiency of the entire system. Moreover, the tasks of the faulty node are quickly transferred to non-faulty nodes, without waiting for the faulty node to be repaired and then restoring data collection. Instead, the non-faulty nodes immediately take over the work through the task transfer mechanism, ensuring that the data collection work of the system will not be interrupted due to the failure of individual nodes, maintaining the continuity of data collection of the entire system. For application scenarios that rely on real-time data, it can avoid accidents caused by data loss, and improve the availability and reliability of the system.

[0109] In some embodiments, after determining that there is a faulty collection node in the distributed collection system, the management method of the distributed collection system further includes: removing the faulty collection node from the instance list to update the instance list.

[0110] In the embodiments of the present application, after the server determines that there is a faulty collection node in the distributed collection system, it can delete the identification information of the faulty collection node from the instance list, or mark it as a faulty state to update the instance list. After the faulty collection node is restored, the restored faulty collection node can be added to the instance list as a new collection node to update the instance list. Optionally, after the faulty collection node is restored, the faulty state of the identification of the restored faulty collection node can be cancelled in the instance list, or marked as a normal state.

[0111] In the method described in the embodiments of the present application, after removing the faulty node, when monitoring other nodes, the monitored data can more accurately reflect the running state of the normal nodes in the system, avoiding the interference of the abnormal data of the faulty node on the monitoring results.

[0112] In some embodiments, a specific implementation manner for managing collection tasks is also provided when the number of collection nodes to which the heartbeat information belongs is consistent with the number of collection nodes included in the instance list, such as Figure 6 As shown, the "managing the collection tasks on all collection nodes in the distributed collection system according to the heartbeat information of all collection nodes and the instance list" in S402 above includes:

[0113] S601, determining whether there is an overloaded collection node among all collection nodes according to the heartbeat information of all collection nodes.

[0114] Among them, the overloaded acquisition node is an acquisition node with a node load less than the load threshold. The load threshold can be determined according to the actual acquisition task requirements and the performance of the acquisition nodes. The load thresholds of each acquisition node can be the same or different.

[0115] In the embodiment of the present application, after the server obtains the heartbeat information of all acquisition nodes, it can extract the load information of all acquisition nodes from the heartbeat information of all acquisition nodes, compare the load of each acquisition node with the preset load threshold, and determine whether the load of each acquisition node is less than the preset load threshold. If there is an acquisition node with a load less than the load threshold, it is determined that there is an overloaded acquisition node among all acquisition nodes, and the acquisition node with a load less than the load threshold is determined as the overloaded acquisition node. If there is no acquisition node with a load less than the load threshold, it is determined that there is no overloaded acquisition node among all acquisition nodes.

[0116] S602, if so, manage the acquisition tasks on the overloaded acquisition nodes and other acquisition nodes.

[0117] In the embodiment of the present application, when the server determines that there is an overloaded acquisition node based on the above steps, it can further reassign the overloaded acquisition tasks on the overloaded acquisition nodes to other acquisition nodes, and manage the acquisition tasks on the overloaded acquisition nodes and other acquisition nodes.

[0118] Optionally, as Figure 7 shown, the "managing the acquisition tasks in the distributed acquisition system according to the heartbeat information of all the acquisition nodes" in S602 above includes:

[0119] S6021, determining a second transfer task assigned to other acquisition nodes according to the heartbeat information of the overloaded acquisition nodes and the heartbeat information of other acquisition nodes.

[0120] Among them, the second transfer task is the task that other acquisition nodes need to take over from the overloaded acquisition nodes.

[0121] In the embodiments of the present application, when the server determines that there is an overloaded acquisition node, it can determine the overloaded acquisition task volume according to the current load information and the load threshold in the heartbeat information of the overloaded acquisition node. Specifically, it determines the overload percentage according to the current load information and the load threshold, and then determines the acquisition task volume corresponding to the percentage as the overloaded acquisition task volume. For example, if the load threshold of the overloaded acquisition node is 80% and the current load information of the overloaded acquisition node is 95%, the overload percentage is 15%. Then, 15% of the total acquisition tasks on the overloaded acquisition node can be determined as the overloaded acquisition task volume. After the server determines the overloaded acquisition task volume of the overloaded acquisition node, it can allocate the acquisition tasks on the overloaded acquisition node to appropriate other acquisition nodes according to at least one of the node load information, acquisition performance information, acquisition performance information, etc. carried in the heartbeat information of each other acquisition node, and obtain the reduced acquisition tasks of the overloaded acquisition node and the newly added second transfer tasks of other acquisition nodes. Specifically, it can allocate the acquisition tasks on the overloaded acquisition node to other acquisition nodes with low other acquisition loads according to the node load information carried in the heartbeat information of each other acquisition node. Optionally, it can allocate the acquisition tasks on the overloaded acquisition node to other acquisition nodes with better other acquisition performance according to the acquisition performance information carried in the heartbeat information of each other acquisition node. Optionally, it can allocate the acquisition tasks on the overloaded acquisition node to other acquisition nodes near the overloaded acquisition node according to the acquisition performance information of each other acquisition node. Among them, a node with a collection rate greater than the rate threshold is a node with better collection performance, and a node with a load less than the load threshold is a node with a lower load.

[0122] S6022. Update the task tables corresponding to all acquisition nodes in the first database according to the second transfer task.

[0123] In the embodiments of the present application, after the server obtains the reduced acquisition tasks of the overloaded acquisition node and the newly added second transfer tasks of other acquisition nodes based on the above steps, it can update the task tables corresponding to other acquisition nodes in the first database according to the second transfer task, and update the task tables corresponding to the overloaded acquisition nodes in the first database according to the reduced acquisition tasks of the overloaded acquisition nodes. After the first transfer tasks of each non-faulty acquisition node, it can update the task tables of each non-faulty acquisition node from the heartbeat information of each non-faulty acquisition node in the first database, so as to instruct each non-faulty acquisition node to process tasks according to the tasks in the corresponding task table. For example, the first transfer task can be added to the acquisition tasks to be processed in the original task tables of each non-faulty acquisition node.

[0124] S603. If not, release the management authority.

[0125] In the embodiment of the present application, when the server determines based on the above steps that there is no overloaded acquisition node, it can be considered that the supervision of the server ends this time. Then, the management authority is returned, and after a preset interval, the step of S201 is executed again to manage the acquisition tasks in the distributed acquisition system according to the heartbeat information of all acquisition nodes.

[0126] In the method described in the embodiment of the present application, since the overloaded acquisition node is prone to performance degradation, response delay, or even crash under high-load operation, by timely detecting the overloaded node based on the heartbeat information and transferring the excess tasks to other nodes, the load pressure on this node can be effectively reduced, preventing it from crashing due to being overwhelmed, thereby ensuring the stable operation of the entire distributed acquisition system. Moreover, the crash of a certain node in the system may trigger a chain reaction, causing the load of other nodes to further increase and even leading to the failure of the entire system. Timely handling of the overloaded node can control the fault risk within a local area, avoid the spread and contagion of the fault, and maintain the overall stability of the system. In addition, by allocating the second transfer task according to the heartbeat information of the overloaded acquisition node and other acquisition nodes, the tasks can be more evenly distributed among the nodes. Transferring some tasks of the overloaded node to the nodes with lighter loads can make full use of the idle resources in the system, improve the resource utilization rate of the entire system, and avoid the situation where some nodes are overly resource-intensive while some nodes have idle resources. Finally, task transfer and task list update are performed according to the real-time heartbeat information, realizing the dynamic management of tasks in the distributed acquisition system.

[0127] In some embodiments, a method for obtaining management authority is also provided, as Figure 8 shown, "obtaining management authority according to the heartbeat information of all acquisition nodes" in S201 above includes:

[0128] S701, determining the state of each acquisition node holding the management lock according to the state information of the management lock in the heartbeat information of each acquisition node.

[0129] Among them, the state information of the management lock includes the state of holding the management lock and the lock holding time, and the state of holding the management lock includes the held state and the unheld state.

[0130] In the embodiment of the present application, after the server obtains the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database through the first database engine, then extracts the state information of the management lock in each acquisition node from the heartbeat information, and then determines the state of each acquisition node holding the management lock according to the state information of the management lock in the heartbeat information of each acquisition node.

[0131] S702. If the status of the management locks of all the acquisition nodes is the unheld status, update the status of its own management lock to the held status to obtain the management authority.

[0132] In the embodiments of the present application, if the server determines that the status of the management locks of all the acquisition nodes is the unheld status, it means that the current management authority of the acquisition system is idle. Then the server can obtain the management authority and update the status of its own management lock to the held status to facilitate other nodes to view the status of the management authority in real time.

[0133] S703. If there is a target acquisition node with the status of the management lock being the held status among all the acquisition nodes, obtain the management authority according to the heartbeat information of the target acquisition node.

[0134] Among them, the target acquisition node is an acquisition node with the status of the management lock being the held status.

[0135] In the embodiments of the present application, if the server determines that there is a target acquisition node with the status of the management lock being the held status among all the acquisition nodes, it means that the current management authority of the acquisition system is not idle, that is, a certain acquisition node in the acquisition system has obtained the management authority and is managing the acquisition tasks of the distributed acquisition system. Then, the server can further obtain the heartbeat information of the target acquisition node, and then obtain the management authority according to the heartbeat information of the target acquisition node.

[0136] In the method described in the embodiments of the present application, in a distributed acquisition system, if multiple nodes attempt to obtain the management authority simultaneously, problems such as management instruction conflicts and data inconsistency may occur. By checking the status of the management locks, only when the status of the management locks of all nodes is the unheld status can a node update the status of its own management lock to the held status and obtain the management authority. This ensures that only one node has the management authority at the same time, avoids the chaos caused by multiple nodes performing management operations simultaneously, and guarantees the orderliness of system management.

[0137] In some embodiments, a specific implementation manner for obtaining the management authority is also provided. As Figure 9 shown, "obtain the management authority according to the heartbeat information of the target acquisition node" in the above S703 includes:

[0138] S801. Extract the lock holding time from the heartbeat information of the target acquisition node.

[0139] Among them, the lock holding time represents the holding time of the management lock, that is, the time of obtaining the management authority.

[0140] In the embodiments of the present application, after the server obtains the heartbeat information of the target acquisition node, it can extract the lock holding time from the heartbeat information of the target acquisition node, and then compare the lock holding time with a preset time threshold to determine whether the lock holding time of the target acquisition node is greater than the preset time threshold.

[0141] S802, if the lock holding time is greater than the preset time threshold, then update the status of the management lock of the target acquisition node to the unheld state, and update the status of its own management lock to the held state to obtain the management authority.

[0142] Among them, the preset time threshold can be determined according to the actual acquisition task requirements and the performance of the acquisition node. The preset time thresholds of each acquisition node can be the same or different.

[0143] In the embodiments of the present application, if the server determines based on the above steps that the lock holding time of the target acquisition node is greater than the preset time threshold, it means that the target acquisition node has low processing efficiency or other abnormalities. Then the server can take over the management authority on the target acquisition node, update the status of the management lock of the target acquisition node to the unheld state to unlock the management authority of the target acquisition node, update the status of its own management lock to the held state, and update the status of the management lock in the heartbeat information of the target acquisition node.

[0144] S803, if the lock holding time is not greater than the preset time threshold, then it is determined that the management authority has not been obtained.

[0145] In the embodiments of the present application, if the server determines based on the above steps that the lock holding time of the target acquisition node is greater than the preset time threshold, it means that the target acquisition node is normally managing the acquisition task. Then it means that the management authority has not been obtained this time. The server can return to execute step S201 after a preset time interval.

[0146] In the method described in the embodiments of the present application, if the target acquisition node cannot perform its management duties normally due to a fault, network interruption, or other abnormal conditions, but the management lock is still in the held state, it will cause the system management to stagnate. By checking the lock holding time, when the lock holding time of the target acquisition node exceeds this threshold, it indicates that the target acquisition node may be abnormal and not suitable for task management currently. Therefore, timely taking back the management authority and reallocating it can effectively avoid this situation and ensure the continuity and stability of system management.

[0147] In summary of all the above embodiments, a management method for a distributed acquisition system for a supercomputer internet is also provided, as Figure 10 shown. This method includes:

[0148] S901. Obtain the heartbeat information of all the acquisition nodes in the distributed acquisition system from the first database, and determine the state of each acquisition node holding the management lock according to the status information of the management lock in the heartbeat information of each acquisition node.

[0149] S902. If the states of the management locks of all the acquisition nodes are in the unheld state, update the state of its own management lock to the held state to obtain the management authority.

[0150] S903. If there is a target acquisition node with the state of the management lock being in the held state among all the acquisition nodes, extract the lock holding time from the heartbeat information of the target acquisition node.

[0151] S904. If the lock holding time is greater than the preset time threshold, update the state of the management lock of the target acquisition node to the unheld state, and update the state of its own management lock to the held state to obtain the management authority.

[0152] S905. If the lock holding time is not greater than the preset time threshold, determine that the management authority has not been obtained, and after a preset time interval, return to execute step S901.

[0153] S906. When it is determined that the management authority has been obtained, obtain the instance list corresponding to the distributed acquisition system from the second database, and compare the acquisition node to which the heartbeat information belongs with the acquisition nodes included in the instance list to obtain a comparison result.

[0154] S907. If the number of acquisition nodes to which the heartbeat information belongs is inconsistent with the number of acquisition nodes included in the instance list, determine that there are faulty acquisition nodes in the distributed acquisition system, and remove the faulty acquisition nodes from the instance list to update the instance list.

[0155] S908. Determine the first transfer tasks assigned to each non-faulty acquisition node according to the previous heartbeat information corresponding to the faulty acquisition node and the heartbeat information of each non-faulty acquisition node.

[0156] S909. Update the task table in the heartbeat information corresponding to each non-faulty acquisition node in the first database according to the first transfer task, and each task table is used to instruct each non-faulty acquisition node to perform task processing according to the tasks in the corresponding task table.

[0157] S910. If the number of acquisition nodes to which the heartbeat information belongs is consistent with the number of acquisition nodes included in the instance list, determine that there are no faulty acquisition nodes in the distributed acquisition system, and determine whether there are overloaded acquisition nodes among all the acquisition nodes according to the heartbeat information of all the acquisition nodes.

[0158] S911, if it exists, determine the second transfer tasks assigned to other acquisition nodes according to the heartbeat information of the overloaded acquisition node and the heartbeat information of other acquisition nodes, and update the task tables corresponding to all acquisition nodes in the first database according to the second transfer tasks.

[0159] S912, if it does not exist, release the management authority, and after a preset time interval, return to execute step S901.

[0160] In the method described in the embodiments of the present application, for a running distributed acquisition system, if an instance in multiple acquisition service instances crashes, the entire distributed system can achieve self-healing and fault transfer, making users outside the distributed system unaware, and all acquisition tasks running normally, with the acquired data being normal, continuous, and stable.

[0161] The methods described in the above steps have been described in the foregoing embodiments. For detailed content, please refer to the foregoing description and will not be repeated here.

[0162] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are displayed in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0163] Based on the same inventive concept, the embodiments of the present application also provide a management device for a distributed acquisition system for a supercomputing Internet, which is used to implement the management method of the distributed acquisition system for a supercomputing Internet described above. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the management device for a distributed acquisition system for a supercomputing Internet provided below can refer to the limitations on the management method of the distributed acquisition system for a supercomputing Internet in the foregoing text and will not be repeated here.

[0164] In some embodiments, as Figure 11 shown, a management device for a distributed acquisition system for a supercomputing Internet is provided, including:

[0165] An acquisition module 11, configured to acquire heartbeat information of all acquisition nodes in the distributed acquisition system from a first database, and obtain management authority according to the heartbeat information of all acquisition nodes.

[0166] A management module 12, configured to, when it is determined that the management authority is obtained, acquire an instance list corresponding to the distributed acquisition system from a second database, and manage acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list.

[0167] In some embodiments, the above management module includes:

[0168] A comparison unit, configured to compare the acquisition node to which the heartbeat information belongs with the acquisition nodes included in the instance list to obtain a comparison result.

[0169] A management unit, configured to manage acquisition tasks on all acquisition nodes in the distributed acquisition system according to the comparison result.

[0170] In some embodiments, the above management unit includes:

[0171] A first management subunit, configured to, if the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is inconsistent with the number of acquisition nodes included in the instance list, determine that there are faulty acquisition nodes in the distributed acquisition system, and manage acquisition tasks on the faulty acquisition nodes and all non-faulty acquisition nodes.

[0172] A second management subunit, configured to, if the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is consistent with the number of acquisition nodes included in the instance list, determine that there are no faulty acquisition nodes in the distributed acquisition system, and manage acquisition tasks in the distributed acquisition system according to the heartbeat information of all acquisition nodes.

[0173] In some embodiments, the above first management subunit is specifically configured to determine a first transfer task assigned to each non-faulty acquisition node according to the previous heartbeat information corresponding to the faulty acquisition node and the heartbeat information of each non-faulty acquisition node; update the task table in the heartbeat information corresponding to each non-faulty acquisition node in the first database according to the first transfer task, and each task table is used to instruct each non-faulty acquisition node to perform task processing according to the tasks in the corresponding task table.

[0174] In some embodiments, the above management unit further includes:

[0175] An elimination unit, configured to eliminate the faulty acquisition nodes from the instance list to update the instance list.

[0176] In some embodiments, the above-mentioned second management subunit is specifically configured to determine whether there are overloaded acquisition nodes among all acquisition nodes according to the heartbeat information of all acquisition nodes; if so, manage the acquisition tasks on the overloaded acquisition nodes and other acquisition nodes; if not, release the management authority.

[0177] In some embodiments, the above-mentioned second management subunit is specifically configured to determine the second transfer tasks assigned to other acquisition nodes according to the heartbeat information of the overloaded acquisition nodes and the heartbeat information of other acquisition nodes; update the task tables corresponding to all acquisition nodes in the first database according to the second transfer tasks.

[0178] In some embodiments, the above-mentioned acquisition module includes:

[0179] A determination unit, configured to determine the state of holding the management lock of each acquisition node according to the state information of the management lock in the heartbeat information of each acquisition node.

[0180] A first acquisition unit, configured to update the state of its own management lock to the held state if the states of the management locks of all acquisition nodes are in the unheld state, so as to obtain the management authority.

[0181] A second acquisition unit, configured to obtain the management authority according to the heartbeat information of the target acquisition node if there is a target acquisition node with the state of the management lock being held among all acquisition nodes.

[0182] In some embodiments, the above-mentioned second acquisition unit includes:

[0183] An extraction subunit, configured to extract the lock holding time from the heartbeat information of the target acquisition node.

[0184] An update subunit, configured to update the state of the management lock of the target acquisition node to the unheld state and update the state of its own management lock to the held state if the lock holding time is greater than a preset time threshold, so as to obtain the management authority.

[0185] A determination subunit, configured to determine that the management authority has not been obtained if the lock holding time is not greater than the preset time threshold.

[0186] Each module in the above-mentioned management device of the distributed acquisition system for the supercomputer Internet can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0187] In some embodiments, a computer device is provided. The computer device can be a terminal or a server, and its internal structure diagram can be asFigure 12 As shown, the computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a management method for a distributed acquisition system for the supercomputing Internet. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0188] Those skilled in the art can understand that Figure 12 the structure shown in

[0189] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0190] In some embodiments, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the above-mentioned management method for the distributed acquisition system for the supercomputing Internet are implemented.

[0191] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps of the above-mentioned management method for the distributed acquisition system for the supercomputing Internet are implemented.

[0192] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0193] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0194] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A management method for a distributed acquisition system for a supercomputing Internet, characterized in that, The management method is applied to any acquisition node in the distributed acquisition system, and the method includes: Obtain the heartbeat information of all acquisition nodes in the distributed acquisition system from the first database, and obtain management authority according to the heartbeat information of all acquisition nodes; When it is determined that the management authority is obtained, obtain the instance list corresponding to the distributed acquisition system from the second database, and manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list.

2. The method according to claim 1, characterized in that, The managing the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the heartbeat information of all acquisition nodes and the instance list includes: Compare the acquisition node to which the heartbeat information belongs with the acquisition nodes included in the instance list to obtain a comparison result; Manage the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the comparison result.

3. The method according to claim 2, wherein The managing the acquisition tasks on all acquisition nodes in the distributed acquisition system according to the comparison result includes: If the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is inconsistent with the number of acquisition nodes included in the instance list, determine that there are faulty acquisition nodes in the distributed acquisition system, and manage the acquisition tasks on the faulty acquisition nodes and all non-faulty acquisition nodes; If the comparison result indicates that the number of acquisition nodes to which the heartbeat information belongs is consistent with the number of acquisition nodes included in the instance list, determine that there are no faulty acquisition nodes in the distributed acquisition system, and manage the acquisition tasks in the distributed acquisition system according to the heartbeat information of all acquisition nodes.

4. The method according to claim 3, characterized in that, The managing the acquisition tasks on the faulty acquisition nodes and all non-faulty acquisition nodes includes: Determine the first transfer tasks assigned to each non-faulty acquisition node according to the previous heartbeat information corresponding to the faulty acquisition node and the heartbeat information of each non-faulty acquisition node; Update the task tables in the heartbeat information corresponding to each non-faulty acquisition node in the first database according to the first transfer tasks, and each task table is used to instruct each non-faulty acquisition node to process tasks according to the tasks in the corresponding task table.

5. The method according to claim 3, wherein After determining that there are faulty acquisition nodes in the distributed acquisition system, the method further includes: Remove the faulty acquisition node from the instance list to update the instance list.

6. The method according to any one of claims 3-5, characterized in that, The managing the acquisition tasks in the distributed acquisition system according to the heartbeat information of all acquisition nodes includes: Determine whether there are overloaded acquisition nodes among all acquisition nodes according to the heartbeat information of all acquisition nodes; If so, manage the acquisition tasks on the overloaded acquisition nodes and other acquisition nodes; If not, release the management authority.

7. The method according to claim 6, wherein The managing the acquisition tasks on the overloaded acquisition nodes and other acquisition nodes includes: Determine a second transfer task assigned to the other collection nodes according to the heartbeat information of the overloaded collection node and the heartbeat information of the other collection nodes; Update the task tables corresponding to all collection nodes in the first database according to the second transfer task.

8. The method according to claim 1, wherein The obtaining of the management authority according to the heartbeat information of all the collection nodes includes: Determine the status of each collection node holding the management lock according to the status information of the management lock in the heartbeat information of each collection node; If the status of the management lock of all the collection nodes is the unheld state, update the status of its own management lock to the held state to obtain the management authority; If there is a target collection node with the status of the management lock being held among all the collection nodes, obtain the management authority according to the heartbeat information of the target collection node.

9. The method according to claim 8, wherein The obtaining of the management authority according to the heartbeat information of the target collection node includes: Extract the lock holding time from the heartbeat information of the target collection node; If the lock holding time is greater than the preset time threshold, update the status of the management lock of the target collection node to the unheld state and update the status of its own management lock to the held state to obtain the management authority; If the lock holding time is not greater than the preset time threshold, determine that the management authority has not been obtained.

10. A management system for a distributed acquisition system for a supercomputing Internet, characterized in that, The management system includes: a distributed collection node cluster, a first database engine, and a second database engine; each collection node in the distributed collection node cluster is respectively connected to the first database engine and the second database engine; a first database is deployed on the first database engine, and a second database is deployed on the second database engine; Any collection node in the distributed collection node cluster is used to execute the method according to any one of claims 1-9.

11. A management device for a distributed acquisition system for a supercomputer Internet, characterized in that, The device includes: An obtaining module, configured to obtain the heartbeat information of all collection nodes in the distributed collection system from a first database, and obtain management authority according to the heartbeat information of all the collection nodes; A management module, configured to, when it is determined that the management authority is obtained, obtain an instance list corresponding to the distributed collection system from a second database, and manage the collection tasks on all collection nodes in the distributed collection system according to the heartbeat information of all the collection nodes and the instance list.

12. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Distributed network task scheduling method and device

    CN111045808A

  • Decentralized HPC computing cluster management method and system based on paxos algorithm

    CN111200518A

  • Method for realizing distributed lock based on database

    CN112241400A

  • Resource scheduling method and device, electronic equipment and storage medium

    CN115686813A

  • Distributed task scheduling method and device, distributed task processing system and medium

    CN119105849A

Cited By

  • Data management system and method for medical affairs

    CN120932792A

  • Data management system and method for the medical community

    CN120932792B