A communication method based on node groups
By adopting a node group-based communication method in a distributed system, the problem of cumbersome and error-prone registration in the existing technology is solved, and efficient system deployment and maintenance is achieved.
Patent Information
- Application Number
- CN202411766695.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-03-04
AI Technical Summary
In existing distributed systems, the automatic registration process of work nodes is cumbersome, prone to errors, and is not conducive to system deployment and maintenance.
Using a communication method based on node groups, the target multicast address and port number are determined through the main center node, listen to and respond to registration requests, update the node allocation map, and determine the sub-center node through preset election operations.
Automatic registration of work nodes is realized, the configuration process is simplified, the system deployment and maintenance efficiency is improved, and the possibility of manual errors is reduced.
Smart Images

Figure CN119652864B_ABST
Abstract
Description
[0001] Description of the case
[0002] This application is a divisional application filed for a Chinese application with a filing date of March 4, 2024, application number 202410240946.9, and invention name “A method for automatic registration of working nodes in a distributed system”. Technical Field
[0003] The present invention relates to the technical field of distributed systems, and in particular to a communication method based on node groups. Background Art
[0004] In distributed systems, automatic registration of working nodes is a common requirement. Currently, automatic registration of working nodes is usually achieved by sending a registration request from the working node to the central node when starting up. After receiving the registration request, the central node stores the information of the working node in the central node's database so that the central node can manage the working node. When starting up a working node, it is necessary to manually configure the central node information, the third-party registration center information, and the node information. The configuration work is cumbersome, error-prone, and not conducive to the deployment and maintenance of distributed systems.
[0005] Therefore, this specification provides a communication method based on node groups to overcome the above problems. Summary of the invention
[0006] One or more embodiments of the present specification provide a communication method based on node groups, which is executed by a main central node, including: determining a target multicast address and a target port number, and monitoring the target multicast address and the target port number; in response to monitoring a pending request: the pending request includes a registration request from the target multicast address and / or the target port number, and the pending request is issued based on a task node, and the task node includes an allocated node and / or a node to be allocated; according to the number of requests for the pending request and the node information of the task node, determining to update a node allocation map, the updated node allocation map includes one or more node groups, and the node group includes a sub-central node and a number of task nodes; the sub-central node is determined by a preset election operation performed according to a preset period, and the preset election operation is performed The operations include: obtaining work record data of each task node in the node group, the work record data including at least one of the registration record, data transmission record, and data processing record of the task node; counting the registration time and deregistration time in the registration record to determine the registration activity of the task node; counting the data transmission speed and transmission stability in the data transmission record to determine the broadband performance of the task node; counting the data processing speed and data processing accuracy in the data processing record to determine the processing performance of the task node; based on the registration activity, broadband performance, and processing performance of the task node, evaluating the robustness of each task node; and determining the sub-center node based on the robustness; and establishing a communication connection with the sub-center node. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] This specification will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:
[0008] Figure 1 It is a module diagram of a system for automatic registration of working nodes of a distributed system according to some embodiments of this specification;
[0009] Figure 2 is an exemplary flow chart of a method for automatically registering a working node of a distributed system according to some embodiments of this specification;
[0010] Figure 3 is a schematic diagram of updating a node allocation map according to some embodiments of this specification;
[0011] Figure 4 is a schematic diagram of a preset election operation process according to some embodiments of this specification;
[0012] Figure 5is an exemplary flow chart of determining and updating a node allocation map according to some embodiments of this specification;
[0013] Figure 6 is a schematic diagram of determining future management load according to some embodiments of this specification;
[0014] Figure 7 is an exemplary schematic diagram of determining a load balancing value according to some embodiments of this specification;
[0015] Figure 8 It is a flowchart of a method for automatic registration of working nodes of a distributed system according to some embodiments of this specification. DETAILED DESCRIPTION
[0016] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following is a brief introduction to the drawings required for the description of the embodiments. Obviously, the drawings described below are only some examples or embodiments of this specification. For ordinary technicians in this field, this specification can also be applied to other similar scenarios based on these drawings without creative work. Unless it is obvious from the language environment or otherwise explained, the same reference numerals in the figures represent the same structure or operation.
[0017] It should be understood that the "system", "device", "unit" and / or "module" used herein are a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.
[0018] As shown in this specification and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "comprise" and "include" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0019] Flowcharts are used in this specification to illustrate the operations performed by the system according to the embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed precisely in order. Instead, the steps may be processed in reverse order or simultaneously. At the same time, other operations may also be added to these processes, or one or more operations may be removed from these processes.
[0020] Figure 1 It is a module diagram of a system for automatic registration of working nodes of a distributed system according to some embodiments of this specification.
[0021] The system 100 for automatically registering working nodes of a distributed system may include a monitoring module 110 , a first determining module 120 and a second determining module 130 .
[0022] The monitoring module 110 is used to determine the target multicast address and the target port number, monitor the target multicast address and the target port number, and respond to the monitored pending request. The pending request may include a registration request from the target multicast address and / or the target port number. The pending request is issued based on the task node, and the task node may include an assigned node and / or a node to be assigned. For more details about the target multicast address, the target port number and the pending request, see Figure 2 and related content.
[0023] The first determination module 120 is used to determine and update the node allocation map according to the number of pending requests and the node information of the task node. The updated node allocation map includes one or more node groups. The node group includes a sub-center node and several working nodes. For more details about the node information, node allocation map, sub-center node and working node, see Figure 2 and related content.
[0024] In some embodiments, the first determination module 120 is used to obtain the work record data of each task node in the node group, the work record data including at least one of the registration record, data transmission record, and data processing record of the task node; to evaluate the robustness of each task node based on the work record data; and to determine the sub-center node based on the robustness. For more details about the work record data and robustness, see Figure 4 And related instructions.
[0025] In some embodiments, the first determination module 120 is used to determine the working performance of the task node based on the node information of the task node, and the working performance includes at least one of broadband performance and processing performance; it is used to predict the future workload of each node group in the basic node allocation map based on the number of requests to be processed; based on the working performance, future workload, and node information, determine to update the node allocation map.
[0026] The second determination module 130 is used to determine the node group where the task node is located based on the updated node allocation map, and forward the pending request issued by the task node to the target sub-center node; the target sub-center node is the sub-center node of the node group where the task node is located.
[0027] The system 100 for automatically registering working nodes of a distributed system may further include a third determining module, a first sending module, and a second sending module.
[0028] The third determination module is used to determine the first heartbeat cycle and the second heartbeat cycle. For more details about the first heartbeat cycle, the second heartbeat cycle and their determination, see Figure 3 And related instructions.
[0029] The first sending module is used to send a first heartbeat cycle to the secondary central node, and send heartbeat information to the secondary central node in the first heartbeat cycle.
[0030] The second sending module is used to send a second heartbeat cycle to the working node, so that the secondary central node sends heartbeat information to the working node in the second heartbeat cycle.
[0031] It should be noted that the above description of the distributed system and its modules is only for convenience of description and does not limit the present specification to the scope of the embodiments. It is understandable that, after understanding the principle of the system, those skilled in the art may arbitrarily combine the modules or form a subsystem to connect with other modules without deviating from the principle. In some embodiments, Figure 1 The monitoring module, the first determination module, and the second determination module disclosed in the specification may be different modules in a system, or one module may realize the functions of two or more modules. For example, each module may share a storage module, or each module may have its own storage module. Such variations are within the protection scope of this specification.
[0032] Figure 2 FIG. 1 is an exemplary flow chart of a method for automatically registering a task node according to some embodiments of this specification. Figure 2 As shown, the process 200 includes the following steps. In some embodiments, the process 200 can be performed by the main central node. For more information about the main central node, see Figure 3 and its related parts.
[0033] Step 210: determine the target multicast address and the target port number, and monitor the target multicast address and the target port number.
[0034] The target multicast address refers to a network address for transmitting content between task nodes in a distributed system by multicast. In some embodiments, the target multicast address may be an IP address.
[0035] A task node refers to a computing device in a distributed system that is responsible for task management, allocation, and execution, etc. In some embodiments, the task node includes a primary central node, a secondary central node, and a working node.
[0036] The port number is a unique digital label used to identify the service process in a computer network. The target port number is the port number corresponding to the task node that needs to be monitored.
[0037] In some embodiments, the main central node may determine the target multicast address and the target port number based on the service situation or by obtaining user input, and monitor the target multicast address and the target port number.
[0038] In some embodiments, in response to monitoring a pending request, the main central node may execute steps 220 to 230, wherein the pending request is issued based on the task node.
[0039] The pending request refers to a request obtained through monitoring and waiting to be forwarded to the secondary central node for registration-related operations. In some embodiments, the pending request may include a registration request from a target multicast address and / or a target port number.
[0040] The sub-center node refers to a node used to process services related to registration requests and manage and maintain the working nodes in the node group. In some embodiments, the sub-center node not only communicates with the working nodes in the node group, but also communicates with the main center node. For more information about the sub-center node, please refer to the following and related content.
[0041] A working node refers to a node for performing actual task execution and data storage. In some embodiments, a working node may include a specific computer or server in a distributed system, and a working node may also include a virtual machine or a cloud server.
[0042] In some embodiments, the working nodes include allocated nodes and / or to-be-allocated nodes, wherein the allocated nodes refer to nodes to which the node group to which they belong has been determined, and the to-be-allocated nodes refer to nodes to which the node group to which they belong has not yet been determined.
[0043] Correspondingly, the pending requests include registration requests of working nodes that have been previously assigned a node group and registration requests of working nodes that have not yet been assigned a node group.
[0044] Step 220, determining to update the node allocation map according to the number of requests to be processed and the node information of the task node.
[0045] The node information may reflect the characteristic information of the task node itself. In some embodiments, the node information may include the node type (e.g., server, virtual machine, computer, etc.), node model, and node address of the task node. For more information about node information, see Figure 4 and its related contents.
[0046] The updated node allocation map refers to a map used to represent the node group to which each task node belongs. The updated node allocation map may include one or more node groups. A node group includes a sub-center node and a plurality of working nodes. A sub-center node and a plurality of working nodes connected thereto constitute a node group.
[0047] In some embodiments, the node group can be determined based on a variety of methods, such as based on node type classification, etc. In some embodiments, the node group can be obtained based on clustering. For more information about clustering, see Figure 4 and its related contents.
[0048] For example, Figure 3 It is an exemplary schematic diagram of updating the node allocation map according to some embodiments of the present specification.
[0049] like Figure 3 The updated node allocation map shown includes three node groups and one main central node 311 . The three node groups are node group 1, node group 2, and node group 3, wherein each node group includes a secondary central node 312 and multiple working nodes 313 .
[0050] The main central node refers to the node responsible for the global planning of the distributed system, which can allocate registration requests, elect sub-central nodes, divide node groups, and plan heartbeat cycles. In some embodiments, the main central node can send heartbeat information to the sub-central nodes in the first heartbeat cycle. For more information about the first heartbeat cycle and heartbeat information, please refer to the following. For more information about node election, please refer to Figure 4 and related instructions.
[0051] The sub-center node refers to a node used to process services related to registration requests and manage and maintain the working nodes in the node group. In some embodiments, the sub-center node can send heartbeat information to the working node in the second heartbeat period. In some embodiments, the main center node can elect a sub-center node from the working nodes in a node group, and the election process can be performed periodically. Among them, the election cycle can be determined by the main center node. For more information about the second heartbeat period, please refer to the following.
[0052] A working node refers to a node used to execute and process actual distributed tasks. In some embodiments, the main central node can cluster the working nodes to obtain one or more node groups.
[0053] In some embodiments, when a communication link is established between two task nodes, the task nodes are connected by edges. The communication connection between the task nodes includes the communication connection between the working node and the sub-center node in the same node group, and the communication connection between the sub-center node and the main center node of each node group.
[0054] In some embodiments, the updated node allocation map is a node allocation map obtained by further determining the node group to which the node to be allocated belongs based on the basic node allocation map. Figure 3 As shown, task node 316 is a to-be-allocated node of a node group to be determined. After the main central node determines that the node group to which task node 316 belongs is node group 1, task node 316 is added to node group 1 as a working node of node group 1. The main central node can then update the basic node allocation map to obtain an updated node allocation map.
[0055] The main central node can obtain the basic node allocation map based on manual input or default settings. For instructions on how to determine the node group to which the node to be allocated belongs to obtain an updated node allocation map, please refer to Figure 5 and its related contents.
[0056] In some embodiments, the main central node can determine to update the node allocation map based on a variety of methods. Exemplarily, when there are task nodes (i.e., nodes to be allocated) whose node groups are to be determined, the main central node can calculate the average value of pending registration requests for each sub-central node based on the number of pending requests and the number of sub-central nodes; and determine the actual value of pending registration requests for each sub-central node based on the number of registration requests currently allocated to each sub-central node.
[0057] The main central node can retain the sub-central nodes whose actual value of pending registration requests is less than the average value of pending registration requests as available nodes. The main central node can calculate the distance between the node to be allocated and all available nodes based on the node information of the node to be allocated, and select the node group where the sub-central node with the shortest distance is located as the node group to which the node to be allocated belongs. The distance can include physical distance, data transmission distance, etc.
[0058] In some embodiments, the master center node can determine the work performance of the task node based on the node information of the task node; predict the future workload of each node group in the basic node allocation map based on the number of requests to be processed; and determine to update the node allocation map based on the work performance, future workload, and node information. For more information about this part, please refer to Figure 5 and related instructions.
[0059] Step 230, based on the updated node allocation map, determine the node group where the task node is located, and forward the pending request issued by the task node to the target sub-center node.
[0060] In some embodiments, the main central node can forward the pending request issued by the task node to the target sub-central node. The target sub-central node is the sub-central node in the node group where the task node is located. As an example only, when the pending request is a registration request, the main central node can forward the registration request issued by the task node to the target sub-central node of the node group where it is located to perform subsequent registration steps.
[0061] In some embodiments of the present specification, by constructing a node allocation map and screening it so that the currently relatively idle sub-center nodes can allocate pending requests, the utilization rate of the distributed system can be effectively improved to ensure that some nodes do not have too many pending requests while some nodes are too idle.
[0062] In some embodiments, process 200 may also include: determining a first heartbeat cycle and a second heartbeat cycle; sending the first heartbeat cycle to the sub-center node, and sending heartbeat information to the sub-center node with the first heartbeat cycle; sending a second heartbeat cycle to the task node; so that the sub-center node sends heartbeat information to the task node with the second heartbeat cycle.
[0063] Heartbeat information refers to confirmation information that can be used to detect whether a task node is online or normal. In some embodiments, a task node can determine whether other task nodes are currently online based on whether it receives heartbeat information sent by other task nodes within a preset time range.
[0064] The heartbeat cycle refers to the time period during which the control task node sends the heartbeat information. In some embodiments, the heartbeat cycle may include a first heartbeat cycle and a second heartbeat cycle.
[0065] The first heartbeat period refers to the heartbeat period used to control the main central node to send heartbeat information to the secondary central node. In some embodiments, the first heartbeat period of the main central node for different secondary central nodes can be the same or different.
[0066] The second heartbeat period refers to the heartbeat period used to control the sub-central node to send heartbeat information to the working node. In some embodiments, the second heartbeat periods of the sub-central node for different working nodes may be the same or different.
[0067] In some embodiments, the main central node may determine the first heartbeat cycle and the second heartbeat cycle in a variety of ways.
[0068] In some embodiments, the first heartbeat period received by the sub-center node of each node group may be negatively correlated to the number of working nodes of the node group. Exemplarily, the more working nodes a node group includes, the more important the sub-center node of the node group is, the more frequent monitoring is required to ensure that management efficiency is not abnormal, and the corresponding first heartbeat period is shorter.
[0069] In some embodiments, the first heartbeat period received by the sub-center node can be dynamically adjusted based on the robustness of the sub-center node. For example, the first heartbeat period received by the sub-center node of each node group can also be related to the robustness of the sub-center node. Among them, the robustness of the sub-center node can reflect the health of the sub-center node. For more information about robustness, please refer to Figure 4 The higher the robustness of the sub-center node, the stronger the bearing capacity of the sub-center node, and the lower the possibility of being offline due to anomalies. In this case, the first heartbeat period corresponding to the sub-center node can be appropriately increased.
[0070] Exemplarily, the first heartbeat period received by the secondary central node of a node group can also be obtained based on the following formula (1):
[0071] C 1 =k 1 ×Ek 2 ×N (1)
[0072] Among them, C 1 represents the first heartbeat cycle, k 1 and k 2 represents the coefficient, N represents the number of working nodes in the node group, and E represents the robustness of the sub-center node. 1 and the coefficient k 2 The value of can be preset based on prior experience.
[0073] In some embodiments of this specification, by making the first heartbeat period related to the robustness of the sub-center node, when the sub-center node has a strong carrying capacity, the first heartbeat period can be increased to reduce the frequency of sending heartbeat information and save computing power. When the sub-center node has a weak carrying capacity, the first heartbeat period can be appropriately reduced to ensure that the current operating status of the sub-center node can be monitored in a timely manner.
[0074] In some embodiments, the second heartbeat period received by the working node may be positively correlated to the average duration of the historical working tasks of the working node. For example, the longer the average duration of the historical working tasks, the more time-consuming the tasks performed by the working node are, and the probability of the working node being offline in a short period of time is not high. At this time, the second heartbeat period corresponding to the working node may be appropriately increased.
[0075] Exemplarily, the second heartbeat period received by the working node may also be obtained based on the following formula (2):
[0076] C 2 = k × T W +d (2)
[0077] Among them, C 2 represents the second heartbeat cycle, k represents the coefficient, T w represents the average duration of the historical work tasks of the work node, and d represents the waiting time. The waiting time refers to the delay time used to ensure that the work of the work node has been completed and there are no subsequent tasks within a period of time. The waiting time is used to implement the work of the work node without immediately judging whether it is working through the heartbeat information after the work is completed. This can avoid the need for the work node to re-register for subsequent tasks within a short period of time after the task is completed, which can save computing resources. Among them, the values of the coefficient k and the waiting time d can be preset based on prior experience.
[0078] In some embodiments, the second heartbeat period received by the working node can be dynamically adjusted based on the active data of the working node. For example, the second heartbeat period received by the working node can also be negatively correlated with the active data of the working node. Among them, the active data of the working node can reflect the load degree brought by different working nodes to the sub-center node. For more information about active data, please refer to Figure 5 and related instructions.
[0079] For example, the larger the active data of the working node, the heavier the load of the current sub-center node, and the higher the possibility of abnormal situations. Therefore, it is necessary to reduce the second heartbeat cycle of the working node to timely deregister inactive working nodes and reduce the load of the sub-center node.
[0080] In some embodiments, by making the second heartbeat period related to the active data of the working node, the second heartbeat period can be increased to save computing power when the sub-center node has a strong carrying capacity. When the sub-center node has a weak carrying capacity, the second heartbeat period can be appropriately reduced to ensure that inactive working nodes can be deregistered in time, reduce the load of the sub-center node, and improve the stability and reliability of the sub-center node.
[0081] In some embodiments, sending heartbeat information based on the first heartbeat cycle and the second heartbeat cycle can avoid the computing power occupation and load increase caused by frequent sending of heartbeat information while monitoring the current operating status of the node, thereby achieving a balance between security and stability.
[0082] Figure 4 This is an exemplary flowchart of the preset election operation shown in some embodiments of this specification.
[0083] In some embodiments, the secondary central node of the node group can be determined based on a preset election operation. The primary central node can perform a preset election operation at a preset period to determine the secondary central node of the node group.
[0084] In some embodiments, the master central node may determine the preset period based on a variety of methods. For example, the master central node may determine the preset period by vector retrieval based on the number of task nodes in the node group and the node information of the task nodes.
[0085] For example, the main central node can construct a vector to be matched using at least the number of task nodes and node information of the task nodes as elements. The main central node can search in a vector database based on the vector to be matched, obtain a reference vector whose vector distance to the vector to be matched is less than a distance threshold, and determine the historical preset period corresponding to the reference vector as the currently required preset period.
[0086] The vector database stores a number of reference vectors and their corresponding historical preset periods. The reference vectors are constructed based on the number of historical task nodes and the node information of the historical task nodes. The vector database can be pre-constructed based on historical data or prior knowledge.
[0087] In some embodiments, the preset period can also be determined based on the increase and decrease information of the node group, the registration efficiency of the current sub-center node, and the management efficiency of the current sub-center node.
[0088] The increase or decrease of personnel information refers to the change information of the increase or decrease of task nodes in the node group. The increase or decrease of personnel information can reflect whether the management load of the current sub-center node is at risk of overload. The main center node can obtain the change information of the increase or decrease of task nodes in the node group through the logs generated when different node groups are running, as the increase or decrease of personnel information.
[0089] In some embodiments, the increase and decrease information may include the net increase in the number of personnel. For example, when the net increase in the number of personnel is too large, it means that a large number of task nodes have been added to the node group where the current sub-center node is located. At this time, the management load of the sub-center node increases, and there may be a risk of management load overload, which requires timely regulation.
[0090] The registration efficiency of the current sub-center node can reflect the success rate of the registration request of the current sub-center node. In some embodiments, the registration efficiency can be negatively correlated with the failure rate of the registration request and the average registration response time. The number of successes and failures of the registration request and the registration response time corresponding to different registration requests can be obtained through the logs generated when different sub-center nodes are running, the failure rate of the registration request is obtained based on the number of successes and failures, and the registration response time corresponding to different registration requests is averaged to obtain the average registration response time.
[0091] The management efficiency of the current sub-center node can reflect the efficiency of the current sub-center node in managing and maintaining the working nodes. In some embodiments, the management efficiency can be negatively correlated to the abnormal response time to the working node and the difference in computing resource occupancy of multiple working nodes in the node group.
[0092] For example, the greater the load difference among multiple working nodes in a node group, the more it means that some working nodes have a higher load, while other working nodes have a lower load, and the task allocation is unreasonable; the longer the response time when an abnormality occurs in a working node, the more it means that the sub-center node does not monitor the working node properly, and cannot detect and repair the abnormality in time, and the management efficiency will be lower.
[0093] In some embodiments, the preset period may be negatively correlated with the increase and decrease in staff information, and positively correlated with the registration efficiency and the management efficiency.
[0094] Exemplarily, the preset period may also be obtained based on the following formula (3):
[0095] S=S 0 ×(1-k 1 ×O+k 2 ×Q+k 3 ×H) (3)
[0096] Where S represents the preset period, S 0 represents the basic cycle, O represents the net increase in the number of employees, Q represents the registration efficiency, H represents the management efficiency, and k 1 , k 2 and k 3 represents the coefficient. Among them, the basic period S 0 and the coefficient k 1 , k 2 , k 3 The value of can be preset based on prior experience.
[0097] In some embodiments of the present specification, a preset period is determined based on the increase and decrease information of the node group, the registration efficiency of the current sub-center node, and the management efficiency of the current sub-center node, and a preset election operation is performed. This can ensure that the preset period will not be too long, resulting in the current sub-center node being unable to cope with it and the election cannot be carried out in time; nor will it be too short, resulting in computing power loss caused by frequent election operations.
[0098] The preset election operation refers to the election step of selecting a sub-center node from all nodes in a node group. In some embodiments, the preset election operation includes: obtaining work record data 410 of each task node in the node group; based on the work record data 410, evaluating the robustness of each task node 450; and determining the sub-center node 460 based on the robustness.
[0099] In some embodiments, the main central node can obtain the work record data of each task node in the node group.
[0100] The work record data 410 refers to log data generated during the operation of the task node. In some embodiments, the work record data may include at least one of the registration record 420, data transmission record 430, and data processing record 440 of the task node.
[0101] Among them, the registration record 420 may include the registration frequency of the task node, the time consumed for each registration, etc.; the data transmission record 430 may include the data transmission volume, data transmission speed, data transmission frequency, data transmission packet loss rate, etc. when the task node performs the task; the data processing record 440 may include the data processing speed, data processing volume, and resource call ratio of the task node when performing the task.
[0102] The resource call ratio refers to the ratio of the computing power required by a work node to the total computing power of the work node when performing a task. The master center node can obtain the computing power required for a task and the total computing power from the log records generated by the work node, and calculate the resource call ratio.
[0103] In some embodiments, the main central node can retrieve log records generated by different working nodes and extract corresponding work record data therefrom.
[0104] In some embodiments, the main central node may evaluate the robustness 450 of each task node based on the work record data.
[0105] The robustness 450 of the task node can be used to characterize the task node's ability to bear and cope with the task. In some embodiments, the robustness can be represented by a level or a value, and the maximum value of the robustness can be 1 and the minimum value can be 0.
[0106] In some embodiments, the master central node may determine robustness in a variety of ways.
[0107] Exemplarily, the main center node can determine the robustness based on the registration record. The higher the registration frequency and the longer the registration duration of the task node in the registration record, the longer the task node is in working state and the more tasks it has, which means that the task node can bear a greater workload and has a higher robustness.
[0108] Exemplarily, the master center node can also determine the robustness based on the data transmission record. The higher the average data transmission speed of the task node in the data transmission record, the higher the data transmission frequency, and the lower the data transmission packet loss rate, the stronger and more stable the data transmission capability of the task node is, and the higher the robustness is.
[0109] Exemplarily, the main center node can also determine the robustness based on the data processing record. The larger the data processing volume of the task node in the data processing record, the faster the processing speed, the stronger the stability, and the lower the proportion of called resources, indicating that the computing resources of the task node are relatively loose, the resource scheduling is reasonable, and the robustness is higher.
[0110] In some embodiments, the main center node evaluates the robustness of each task node and further includes: performing statistics on the registration time and deregistration time in the registration record to determine the registration activity of the task node; performing statistics on the data transmission speed and transmission stability in the data transmission record to determine the broadband performance of the task node; performing statistics on the data processing speed and data processing accuracy in the data processing record to determine the processing performance of the task node; and determining the robustness based on the registration activity, broadband performance, and processing performance of the task node. Among them, the data processing speed can be determined based on the amount of data processed per unit time.
[0111] Registration activity refers to the activity of task nodes in registering; in some embodiments, registration activity may be negatively correlated to the interval between the last logout time and the last registration time. For example, the closer the last logout time of a task node is to the next registration time, the higher the registration activity.
[0112] In some embodiments, the main center node can also count the task nodes, obtain the time intervals between multiple logout times and the next registration, and calculate statistical values such as the average value of multiple time intervals; the smaller the average value, the higher the registration activity.
[0113] The broadband performance can characterize the data transmission capability of the task node. In some embodiments, the faster the data transmission speed of the task node and the higher the stability, the higher the broadband performance.
[0114] In some embodiments, the stability is negatively correlated to the variance of the transmission speed and the average value of the packet loss rate.
[0115] In some embodiments, the master central node may obtain stability based on the following formula (4):
[0116]
[0117] Among them, Z represents the variance of transmission speed, D represents the average packet loss rate, and k 1 and k 2 is the coefficient. Among them, the coefficient k 1 and k 2 The value of can be preset based on prior experience.
[0118] In some embodiments, the main central node may obtain the broadband performance based on the following formula (5):
[0119] H=k 1 ×W+k 2 ×V (5)
[0120] Among them, H represents broadband performance, W represents stability, V represents the average transmission speed, and k 1 and k 2 is the coefficient. Among them, the coefficient k 1 and k 2 The value of can be preset based on prior experience.
[0121] The processing performance can represent the task node's ability to complete the task. In some embodiments, the faster the data processing speed of the task node and the higher the data processing accuracy, the higher the processing performance.
[0122] In some embodiments, the processing performance may be obtained based on the following formula (6):
[0123]
[0124] Where I represents the processing performance, n represents the total number of types of data processing speed, and k i represents the coefficient corresponding to the i-th data processing speed, Q i represents the data processing accuracy corresponding to the i-th data processing speed, P i represents the i-th data processing speed. Among them, the coefficient k i The value of can be preset based on prior experience.
[0125] In some embodiments, robustness is positively correlated with the registration activity, bandwidth performance, and processing performance of the task node.
[0126] In some embodiments, the robustness can be obtained based on the following formula (7):
[0127] E=k 1 ×R+k 2 ×H+k 3 ×I (7)
[0128] Among them, E represents robustness, R represents registration activity, H represents broadband performance, I represents processing performance, and k 1 , k 2 , k 3 is the coefficient. Among them, the coefficient k 1 , k 2 and k 3 The value of can be preset based on prior experience.
[0129] In some embodiments of this specification, by extracting relevant data from registration records, data transmission records, and data processing records, and determining robustness based on the registration activity, broadband performance, and processing performance of the task node, the obtained robustness can better reflect the actual bearing capacity of the task node, thereby improving the credibility and accuracy of the results.
[0130] In some embodiments, the main central node may determine the secondary central node 460 based on the robustness 450 of the task node.
[0131] In some embodiments, the main central node can sort the task nodes in a certain node group according to the level of robustness, select the task nodes with the highest robustness and determine them as the sub-central nodes.
[0132] In some embodiments of the present specification, the required sub-center nodes are selected from the node group based on robustness, which can ensure that the obtained sub-center nodes can respond to registration requests from other task nodes without problems such as overload, thereby improving the stability and reliability of the system.
[0133] In some embodiments, the main center node can determine the number of clusters based on the number of all task nodes; construct a grouping vector based on the node information of the task node and the historical task distribution of the task node; perform clustering based on the vector distance between the grouping vectors and the number of clusters to obtain at least one clustering cluster; determine at least one node group based on at least one clustering cluster. In some embodiments, the node information can include node attributes and node addresses.
[0134] The node address refers to the address where the task node is located. In some embodiments, the node address can affect the data transmission of the node. For example, the farther the address distance between two nodes is, the more unstable the data transmission is and the more resources the transmission consumes. Therefore, clustering working nodes with similar node addresses into a working group can save data transmission resources and improve data transmission efficiency.
[0135] Node attributes refer to information such as the task type and model of the task node. In some embodiments, the task type of the task node may include data storage, image processing, text recognition, etc. In order to save management costs and facilitate unified management of sub-center nodes, task nodes with similar processing tasks and processing logic can be placed in the same work group through clustering.
[0136] The number of clusters refers to the number of clusters required in the clustering process. In some embodiments, the number of clusters is positively correlated to the total number of task nodes and negatively correlated to the maximum number of nodes in the node group. The maximum number of nodes in the node group refers to the maximum number of task nodes that can be included in a node group, which can be obtained based on prior experience or historical data.
[0137] In some embodiments, the main central node may determine the ratio of the total number of task nodes to the maximum number of nodes in the node group as the number of clusters.
[0138] The historical task distribution of a task node refers to the distribution of historical tasks processed by the task node in terms of task types. In some embodiments, since different work nodes are good at different task types, clustering work nodes with similar task types together can improve management efficiency.
[0139] In some embodiments, the main central node can construct a grouping vector based on the node information corresponding to each task node and the historical task distribution. For example, the node information of the task node includes the node address A, the node type B; the historical task distribution is C, then the corresponding grouping vector is [N: (A, B); C].
[0140] In some embodiments, the main central node may cluster each task node according to the corresponding grouping vector based on the vector distance and the number of clusters based on a clustering algorithm to obtain clusters corresponding to the number of clusters.
[0141] There may be multiple types of clustering algorithms, for example, clustering algorithms may include K-Means (K-means) clustering, density-based clustering method (DBSCAN), etc.
[0142] In some embodiments, the main central node may determine each cluster obtained by clustering as a node group.
[0143] In some embodiments of the present specification, a grouping vector is constructed and clustered based on the node information of the task nodes and the historical task distribution of the task nodes to determine the node group. Task nodes with similar functions and similar tasks can be divided into the same node group to facilitate unified management of sub-center nodes and save management costs.
[0144] In some embodiments, clustering also needs to be performed under a saturation constraint. Therefore, during clustering, the saturation of the secondary center node in each cluster obtained cannot exceed the saturation constraint.
[0145] The saturation can reflect the load level when the sub-center node manages the working nodes. In some embodiments, the saturation can be represented by the sum of the load rates of all working nodes managed by the sub-center node. The load rate represents the degree to which each working node occupies the management capacity of the sub-center node.
[0146] The saturation constraint condition refers to a limiting condition that constrains the saturation of the sub-center node of each cluster. In some embodiments, the saturation constraint condition may include that the load rate of the sub-center node does not exceed a load rate threshold. The load rate threshold may be preset based on prior experience.
[0147] For example, the load factor can be obtained based on the following formula (8):
[0148]
[0149] Where η represents the load rate, F represents the management load that each working node brings to the sub-center node, and F ALL Indicates the theoretical maximum load of the sub-center node. The theoretical maximum load corresponding to the strongest working node in the node group can be determined as the theoretical maximum load F of the sub-center node. ALL .
[0150] In some embodiments, when the obtained cluster does not meet the saturation constraint, the working nodes in the cluster can be adjusted. Exemplarily, when the load rate of the sub-center node in a cluster exceeds the load rate threshold, the working node farthest from the cluster center of the cluster can be selected according to the distance of each working node from the cluster center, and it is determined whether the saturation constraint can be met if the working node is divided into other clusters closer to it, and a cluster with the closest distance is selected from the clusters that can meet the saturation constraint after division to complete the division operation.
[0151] In some embodiments of the present specification, saturation constraints also need to be considered during the clustering process, so as to ensure that the load of the sub-center node in each cluster will not be overloaded, thereby ensuring the stability and reliability of the operation process of all sub-center nodes.
[0152] Figure 5 is an exemplary flow chart of determining and updating a node allocation map according to some embodiments of this specification. Figure 5 As shown, when the main central node determines to update the node allocation map, it can determine the work performance 532 of the task node based on the node information 510 of the task node, and predict the future workload 533 of each node group in the basic node allocation map based on the number of requests 521 to be processed, and determine to update the node allocation map 540 based on the work performance 532, the future workload 533, and the node information 510.
[0153] In some embodiments, the main central node may determine the work performance 532 of the task node based on the node information 510 of the task node.
[0154] Work performance 532 may reflect the ability of the task node to solve problems. In some embodiments, work performance includes at least one of bandwidth performance and processing performance. For more information about bandwidth performance and processing performance, see Figure 4 and related instructions.
[0155] In some embodiments, the main central node may use the historical bandwidth performance and historical processing performance of the historical task node closest in model and node address as the working performance of the node to be assigned based on the model and node address of the node to be assigned.
[0156] In some embodiments, the master central node may predict the future workload 533 of each node group in the basic node allocation map based on the number of requests 521 to be processed.
[0157] Future workload 533 refers to the estimated amount of tasks that the node group needs to process in the future. In some embodiments, the main center node can divide the number of pending requests according to the node groups to which they are known to belong; and determine the future workload of the node group according to the number of pending requests corresponding to each node group. Exemplarily, the more pending registration requests in a node group, the more working nodes the node group needs to complete the work, and the heavier the future workload of the node group.
[0158] In some embodiments, the main central node may determine to update the node allocation map 540 based on the work performance 532 , the future workload 533 , and the node information 510 .
[0159] In some embodiments, the main central node can determine the updated node allocation map in a variety of ways. Exemplarily, the main central node can determine the number of task nodes required for different node groups according to the future workload of different node groups and the working performance of task nodes. And according to the node attributes and node addresses of different task nodes, the most appropriate node group is selected and allocated.
[0160] Exemplarily, a node group whose node attributes of the task node are similar to those of the sub-center node and whose distance to the node address of the sub-center node is the shortest may be selected as the node group of the node to be assigned.
[0161] In some embodiments, the main central node can also calculate the total working performance of the nodes to be allocated; calculate the theoretical allocable performance of each node group according to the proportion of future workload; select the node address of the node to be allocated based on the distance between the node address of the secondary central node of the node group and the node address within a preset range, and allocate one (or more) or more nodes whose working performance (or the sum of working performance) is closest to the theoretical allocable performance to the node group.
[0162] In some embodiments, if there are multiple allocation methods that can meet the requirements, an allocation method that is closer to the node address of the secondary center node and has a higher similarity in node attributes is preferentially selected for allocation.
[0163] Exemplarily, the total working performance of the nodes to be allocated is 100, and the future workload of node group 1 accounts for 30% of the future workload of all working groups. According to the proportional allocation, one or more nodes to be allocated with a total working performance of 30 need to be allocated to node group 1. At this time, there are four nodes to be allocated whose node addresses are within the preset range from the node addresses of the nodes to be allocated to the node addresses of the sub-center nodes of the node group, namely: node 1 to be allocated with a working performance of 14, node 2 to be allocated with a working performance of 16, node 3 to be allocated with a working performance of 12, and node 4 to be allocated with a working performance of 18. In order to meet the condition that the working performance (or the sum of the working performance) is closest to the theoretically allocable performance, there are two allocation methods at this time, namely, allocating nodes 1 and 2 to working group 1, and allocating nodes 3 and 4 to working group 1. At this time, the node address distances between the nodes to be allocated and the sub-center nodes of the two allocation methods can be calculated, and the allocation method with the smallest node address distance can be selected for allocation.
[0164] In some embodiments, the updated node allocation map is also updated based on the management load 534 of multiple sub-center nodes. In some embodiments, the main center node can determine the management load of the sub-center node based on the following method. The main center node can obtain the work record data 410 of each task node in multiple node groups; based on the work record data 410, determine the active data 522 of each task node; based on the active data 522, the grouping vector 523 and the vector distance 524 of the sub-center node, determine the management load 534 of the sub-center node.
[0165] The management load 534 of the sub-center node refers to the load caused by the sub-center node processing the registration tasks of the working nodes in the node group and maintaining the work in the working group. In some embodiments, when allocating the nodes to be allocated, in order to ensure that the sub-center nodes can run smoothly and orderly, it is necessary not only to meet the condition that the total working performance of the nodes to be allocated to different node groups is closest to the allocable theoretical performance, but also to make the management load in the node group as low as possible.
[0166] Active data 522 refers to the data generated when the task node interacts with the sub-center node, which can characterize the impact of the task node on the load level of the sub-center node. In some embodiments, active data can include registration activity, work activity, etc. For more information about registration activity, please refer to Figure 4 and related instructions.
[0167] The work activity can reflect the amount of work completed by the task node. In some embodiments, the work activity is positively correlated with the data processing amount and data transmission amount of the task node.
[0168] In some embodiments, the management load may be obtained based on the following formula (9):
[0169] Where G represents the management load, T represents the activity, L represents the vector distance between the grouping vector of the working node and the grouping vector of the sub-center node, and k 1 is the coefficient. Among them, the coefficient k 1 The value of can be preset based on prior experience. The value of activity T can be the average of registration activity and work activity. The main center node can calculate the vector distance between the grouping vector of the working node and the grouping vector of the secondary center node by using methods such as Euclidean distance or Manhattan distance. For more information about grouping vectors, see Figure 4 and related instructions.
[0170] In some embodiments of the present specification, when determining to update the node allocation map, the impact on the management load of the sub-center node is taken into consideration, so as to ensure that when the total working performance of the node group can meet the demand, the management load of the sub-center node will not be too large, thereby effectively avoiding abnormal situations caused by excessive management load of the sub-center node.
[0171] In some embodiments, the main central node can also generate multiple allocation maps to be optimized based on the working performance of multiple nodes to be allocated and the future workload of different node groups to form a race to be optimized; based on the management load and load balance value, determine the evaluation value of the allocation map to be optimized through the objective function; based on the evaluation value, perform at least one round of iterative updates on the allocation map to be optimized to determine the updated node allocation map.
[0172] The allocation map to be optimized refers to the candidate node allocation map that needs to be iteratively updated. In the allocation map to be optimized, each task node / node to be allocated is connected to the corresponding sub-center node through an edge; each sub-center node is connected to the main center node through an edge. The specific composition of the allocation map to be optimized is similar to the updated node allocation map. Please refer to the relevant content of the updated node allocation map above.
[0173] In some embodiments, when the main central node allocates the nodes to be allocated to the node groups according to the above method for determining and updating the node allocation map, there may be a situation where multiple allocation methods can meet the requirements. At this time, a to-be-optimized allocation map can be generated based on each allocation method.
[0174] Exemplarily, when allocating based on the condition that the working performance (or the sum of the working performance) of the node to be allocated is closest to the theoretical allocable performance, and the distance between the node address of the node to be allocated and the node address of the sub-center node of the node group is within a preset range, if there is allocation method 1: nodes 1 and 4 to be allocated are allocated to node group 1, and nodes 2 and 3 are allocated to node group 2; allocation method 2: nodes 1 and 2 to be allocated are allocated to node group 1, and nodes 3 and 4 are allocated to node group 2, there are two allocation methods at this time, and two allocation graphs to be optimized can be generated.
[0175] The load balancing value can represent the uniformity of the loads of multiple nodes. The load balancing value can be represented by a numerical value. The larger the numerical value, the more uneven the loads of multiple nodes.
[0176] In some embodiments, the load balance value is positively correlated to the sum of the overload duration ratios and the sum of the idle duration ratios of multiple nodes in the allocation map to be optimized. For more information about the overload duration and the idle duration, please refer to Figure 7 and related instructions.
[0177] The objective function refers to a function that determines the evaluation value of the allocation map to be optimized. The evaluation value can reflect the quality of the allocation map to be optimized. In some embodiments, the evaluation value can be represented by a numerical value. The smaller the evaluation value, the more reasonable the node allocation of the current allocation map to be optimized and the better the stability.
[0178] In some embodiments, the evaluation value may be obtained based on the following formula (10):
[0179]
[0180] Among them, A represents the evaluation value, It indicates the total management load brought to the sub-center node by the nodes to be allocated in the allocation map to be optimized and the original working nodes of each node group. represents the average value of the load balancing value of all sub-center nodes in the distribution map to be optimized, k 1 and k 2 is the coefficient. Among them, the total management load It can be calculated by formula (9); the load balance value can be calculated by formula (12), the coefficient k 1 and k 2 The value of can be preset based on prior experience. For more information about load balancing values and formula (12), see Figure 7 and related instructions.
[0181] Iterative updating may refer to a process of repeatedly updating the race to be optimized through multiple rounds and finally determining the updated node allocation map. In some embodiments, iterative updating may include the following steps:
[0182] Step A: Sort the node allocation graphs in the optimization race according to the evaluation values from small to large, and select a preset number of allocation graphs to be optimized that are ranked high as candidate allocation graphs.
[0183] Step B: transform one or more nodes to be allocated in the candidate allocation graph, obtain an updated allocation graph after the transformation operation, and calculate the evaluation value of the updated allocation graph. In some embodiments, the transformation operation may be to change the node group to which one or more nodes to be allocated in the optimized allocation graph belong.
[0184] Step C: Sort the candidate allocation maps and the update allocation maps again according to the evaluation values from small to large, and select a preset number of allocation maps to be optimized with the highest ranking as the update optimization race.
[0185] Step D: Determine whether the preset iteration conditions are met. If not, perform subsequent rounds of iterative updates, and use the updated optimization race obtained in step C as the race to be optimized for the next round of iterations. If the preset iteration conditions are met, stop the iteration and obtain the updated optimization race after the iteration is completed.
[0186] The preset iteration condition may be reaching a pre-specified number of iterations, the evaluation value no longer changing, the evaluation value changing less than a preset threshold in multiple consecutive iterations, etc. The preset threshold may be preset based on prior experience.
[0187] In some embodiments, the main central node may sort the update optimization races from small to large according to the evaluation value, and select the allocation map to be optimized with the smallest evaluation value as the update node allocation map.
[0188] In some embodiments of the present specification, the node allocation graph is iteratively updated to determine the required updated node allocation graph. The node allocation graph can be simulated and predicted multiple times by iterative updating, and adjusted repeatedly to find an ideal node allocation method from complex actual situations.
[0189] like Figure 6 The diagram shows a schematic diagram of determining the future management load. In some embodiments, the management load also includes a future management load 650. The future management load refers to the estimated management load at a future time point. The main center node can also construct a load analysis map 620 based on the allocation map to be optimized 610; based on the load analysis map 620, the future active data 640 of each task node is predicted through the active model 630, and the active model is a machine learning model; based on the future active data 640 of each task node, the future management load 650 is determined.
[0190] The load analysis graph 620 refers to a knowledge graph representing the load level of the sub-center node. The load analysis graph 620 can reflect the characteristics of the distribution of different nodes in the distributed system. The load analysis graph can be composed of at least one node and at least one edge.
[0191] In some embodiments, the node includes a main central node 311, a secondary central node 312 and a working node 313, wherein the working node includes a to-be-allocated node and an allocated node. The node attributes of a node may be the node type, node working performance and historical working data of the node.
[0192] In some embodiments, the edge includes an edge 315 between a working node and a corresponding sub-center node, and an edge 314 between a sub-center node and a main center node. The direction of the edge is the direction of data transmission between the nodes. When there is only unidirectional data transmission, the edge is a unidirectional edge, and when there is bidirectional data transmission, the edge is a bidirectional edge. The attributes of the edge include the data transmission frequency, transmission volume, transmission speed, etc. in the historical data transmission record. The distance between two nodes in the load analysis map can reflect the address distance between the two nodes. When there is data transmission between two nodes, the nodes are connected by an edge.
[0193] The activity model 630 refers to a model for estimating activity data at future moments. In some embodiments, the activity model 630 may be a machine learning model. For example, the activity model may include any one or combination of a graph neural network (GNN) model or other custom model structures.
[0194] In some embodiments, the input of the activity model 630 may include a load analysis graph 620 , and the output may include future activity data 640 of each task node (including nodes to be assigned and assigned nodes).
[0195] In some embodiments, the main central node may train an active model based on a large number of first training samples with a first label. In some embodiments, the first training sample may include a load analysis diagram constructed from a node allocation diagram in a first historical period. The first label may be actual active data of each sub-central node in a second historical period. The first historical period is earlier than the second historical period.
[0196] During training, a loss function is constructed based on the output and label of the initial active model, and the parameters of the initial active model are iteratively updated based on the loss function until the preset conditions are met, the training ends, and the trained active model is obtained 630. The preset conditions may include but are not limited to the convergence of the loss function and the reaching of a threshold value of the training cycle.
[0197] In some embodiments, the main central node may also determine the future management load 650 based on the future activity data 640 of each task node. The method for determining the future management load 650 is similar to the method for determining the management load, and reference may be made to the above related content.
[0198] In some embodiments, the evaluation value may also be obtained based on the following formula (11):
[0199]
[0200] Among them, A represents the evaluation value, It represents the total management load brought to the sub-center node by the nodes to be allocated in the allocation graph to be optimized 610 and the original working nodes of the working group, represents the average value of the load balancing values of all sub-center nodes in the allocation map to be optimized 610, W represents the future management load 650, k 1 , k 2 and k 3 is the coefficient. Among them, the total management load and the average of the load balancing values The value of is obtained in a similar way to formula (10), and the relevant content can be found above; the coefficient k 1 and k 2 and k 3 The value of can be preset based on prior experience.
[0201] Some embodiments of the present specification provide a method for determining future management load by inputting a load analysis graph into an active model to obtain future active data, and analyzing and learning a large amount of data through a machine learning model, which can make the prediction of the future management load more accurate.
[0202] In some embodiments of the present specification, by using a load analysis chart to analyze and process the nodes and related data within the system, the load analysis chart is input into a trained model to obtain the required future active data, and used in the calculation of the evaluation value, the evaluation value can include the impact of future management loads and better reflect the load conditions of the sub-center nodes.
[0203] Figure 7 It is an exemplary schematic diagram of determining a load balancing value according to some embodiments of the present specification.
[0204] In some embodiments, the evaluation value of the allocation map to be optimized 610 is positively correlated with the load balancing value of the task node. The determination of the load balancing value includes: the main center node can construct a load balancing map 720 based on the allocation map to be optimized 610; input the load balancing map 720 into the load evaluation model 730 to determine the load balancing value 740, and the load evaluation model 730 is a machine learning model.
[0205] The load balance graph 720 refers to a knowledge graph representing the load balance degree of the secondary center node. The load balance graph may be composed of at least one node and at least one edge.
[0206] In some embodiments, the node includes a main central node 311, a secondary central node 312 and a working node 313, wherein the working node includes a to-be-allocated node and an allocated node. The node attributes of a node may be the node type, node working performance and historical working data of the node.
[0207] In some embodiments, the edge includes an edge 315 between a working node and a corresponding secondary central node, and an edge 314 between a secondary central node and a primary central node. The attributes of the edge include the address distance between the two nodes connected by the edge, the first heartbeat cycle or the second heartbeat cycle, etc. When data is transmitted between two nodes, the edge is used to connect the nodes. For more information about the first heartbeat cycle or the second heartbeat cycle, please refer to Figure 3 and its related contents.
[0208] The load assessment model 730 refers to a model for estimating a load balancing value. In some embodiments, the load assessment model 730 may be a machine learning model. For example, the load assessment model 730 may include any one or combination of a graph neural network (GNN) model or other custom model structures.
[0209] In some embodiments, the input of the load assessment model 730 may include a load balancing map, and the output may include a load balancing value 740 of the load balancing map.
[0210] In some embodiments, the main central node may train an active model based on a large number of second training samples with second labels.
[0211] In some embodiments, the second training sample may include a sample load balance map constructed based on the historical sample in the first historical period. The second label may be a load balance value of the sample load balance map. The historical sample includes the historical main center node, the historical sub-center node, the historical working node corresponding to the historical node type, the historical node working performance, the historical working data and the distance between the nodes.
[0212] In some embodiments, for each historical sample of the first historical period, the main center node may count the overload duration ratio and the idle duration ratio of each task node in the historical sample in the second historical period. Overload may refer to the workload of the node being greater than the first preset threshold; idle may refer to the workload of the node being less than the second preset threshold. The load balance value of the historical sample is calculated based on the overload duration ratio and the idle duration ratio of the historical sample.
[0213] Exemplarily, the load balancing value of the historical sample can be obtained based on the following formula (12):
[0214]
[0215] Where M represents the load balancing value, n represents the total number of node groups, and X i represents the average value of the overload duration of the task nodes in node group i, Y i represents the average value of the idle time ratio of the task nodes in node group i, a i and b i are the weights corresponding to the average value of the overload duration ratio and the average value of the no-load duration ratio of node group i.
[0216] In some embodiments, the weight a i and b i Positively correlated with the number of task nodes in node group i. For example, the node group with more task nodes requires more computing power and time for scheduling, and has a greater impact on load balancing, so its weight can be appropriately increased.
[0217] The training process of the load assessment model 730 is similar to the training process of the active model, and reference may be made to the relevant content above.
[0218] In some embodiments of the present specification, the accuracy and reliability of the result calculation can be improved by extracting and processing the node-related data using the load balancing graph and then obtaining the required load balancing value based on the load evaluation model.
[0219] Figure 8 This is a flow chart of a method for automatically registering a working node of a distributed system according to some embodiments of this specification. In some embodiments, a method for automatically registering a working node of a distributed system is performed by a sub-center node. For more information about the sub-center node, see Figure 2 And related content. The above method comprises the following steps:
[0220] Step 810: In response to receiving a registration request of a node to be allocated sent by a main central node, node information of the node to be allocated is determined according to the registration request.
[0221] The registration request may include a request from the node to be assigned to the main central node to execute a distributed task. The registration request includes node information of the node to be assigned. In some embodiments, the node to be assigned broadcasts the registration request to the network periodically or in real time until feedback of successful registration information is received. For more information about the main central node, see Figure 2 and related details.
[0222] In some embodiments, after receiving the registration request forwarded by the main central node, the secondary central node can determine the node information of the node to be assigned based on the registration request. Figure 2 and related details.
[0223] Step 820: Send registration success information to the node to be assigned, and store the node information of the node to be assigned in the node information database. The node to be assigned serves as a working node to perform distributed tasks.
[0224] The registration success information refers to the information that the sub-center node feeds back to the node to be assigned in response to its registration request. After the sub-center node sends the registration success information to the node to be assigned, it means that the node to be assigned joins the corresponding node group as a working node of the node group to which the sub-center node belongs. For instructions on how to determine the node group to which the node to be assigned belongs, please refer to other parts of this manual, such as Figure 2 The corresponding content.
[0225] The node information database refers to a database storing the node information of the task nodes of each node group. In some embodiments, the sub-center node can store the node information of the node to be assigned in the node information database. In some embodiments, the data in the node information database can be entered by the main center node or the sub-center node.
[0226] In some embodiments, after receiving the registration success information, the node to be assigned can be used as a working node of the corresponding node group to perform distributed tasks. The working node will receive the heartbeat information sent from the sub-center node based on the second heartbeat cycle until the working node goes offline.
[0227] For more information about the second heartbeat cycle and heartbeat information, see Figure 2 and related details.
[0228] In some embodiments of the present description, the automatic registration method executed by the sub-center node can determine the node information of the node to be assigned and store it in the node information library after receiving the registration request from the main center node, so that the distributed system can realize automatic registration of working nodes.
[0229] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only for example and does not constitute a limitation of this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements and corrections to this specification. Such modifications, improvements and corrections are suggested in this specification, so such modifications, improvements and corrections still belong to the spirit and scope of the exemplary embodiments of this specification.
[0230] At the same time, this specification uses specific words to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures or characteristics in one or more embodiments of this specification can be appropriately combined.
[0231] In addition, unless explicitly stated in the claims, the order of the processing elements and sequences described in this specification, the use of alphanumeric characters, or the use of other names are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some invention embodiments that are currently considered useful through various examples, it should be understood that such details are only for illustrative purposes, and the attached claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.
[0232] Similarly, it should be noted that in order to simplify the description disclosed in this specification and thus help understand one or more embodiments of the invention, in the above description of the embodiments of this specification, multiple features are sometimes combined into one embodiment, figure or description thereof. However, this disclosure method does not mean that the features required by the subject matter of this specification are more than the features mentioned in the claims. In fact, the features of the embodiments are less than all the features of the single embodiment disclosed above.
[0233] In some embodiments, numbers describing the number of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise specified, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may change according to the required features of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the setting of such numerical values is as accurate as possible within the feasible range.
[0234] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, etc., cited in this specification are hereby incorporated by reference in their entirety. Except for application history documents that are inconsistent with or conflicting with the contents of this specification, documents that limit the broadest scope of the claims of this specification (currently or later attached to this specification) are also excluded. It should be noted that if the descriptions, definitions, and / or use of terms in the materials attached to this specification are inconsistent or conflicting with the contents described in this specification, the descriptions, definitions, and / or use of terms in this specification shall prevail.
[0235] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, as an example and not a limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly introduced and described in this specification.
Claims
1. A communication method based on node groups, characterized in that: Executed by the main central node, including: Determine the target multicast address and the target port number, and monitor the target multicast address and the target port number; In response to monitoring a pending request: the pending request includes a registration request from the target multicast address and / or the target port number, the pending request is issued based on a task node, and the task node includes an allocated node and / or a node to be allocated; According to the number of requests to be processed and the node information of the task node, an updated node allocation map is determined, wherein the updated node allocation map includes one or more node groups, wherein the node group includes a sub-center node and a plurality of task nodes; the sub-center node is determined by a preset election operation performed according to a preset period, wherein the preset election operation includes: Acquire work record data of each of the task nodes in the node group, wherein the work record data includes at least one of a registration record, a data transmission record, and a data processing record of the task node; Collecting statistics on the registration time and the deregistration time in the registration record to determine the registration activity of the task node; The data transmission speed and transmission stability in the data transmission record are counted to determine the broadband performance of the task node; The data processing speed and the data processing accuracy in the data processing record are counted to determine the processing performance of the task node; Based on the registration activity, the broadband performance, and the processing performance of the task nodes, the robustness of each of the task nodes is evaluated; Based on the robustness, determining the secondary center node; Establish a communication connection with the sub-center node.
2. The method according to claim 1, characterized in that The preset period is determined based on the increase and decrease information of the node group, the registration efficiency of the current sub-center node, and the management efficiency of the current sub-center node.
3. The method according to claim 1, characterized in that Each of the node groups is obtained by clustering and includes: Determine the number of clusters based on the number of all task nodes; Constructing a grouping vector based on the node information of the task node and the historical task distribution of the task node; Performing clustering based on the vector distance between the grouping vectors and the number of clusters to obtain at least one cluster cluster; At least one node group is determined based on the at least one cluster.
4. The method according to claim 3, characterized in that The clustering is performed under a saturation constraint condition, where the saturation constraint condition includes that the load rate of the secondary center node does not exceed a load rate threshold.
5. The method according to claim 1, characterized in that The determining, according to the number of requests to be processed and the node information of the task node, to update the node allocation map comprises: Determine the working performance of the task node based on the node information of the task node, where the working performance includes at least one of broadband performance and processing performance; Based on the number of requests to be processed, predicting the future workload of each of the node groups in the basic node allocation map; Based on the work performance, the future workload, and the node information, determine to update the node allocation map.
6. The method according to claim 5, characterized in that The updating of the node allocation map is also based on the management load update of the plurality of sub-center nodes; The method for determining the management load of the sub-center node includes: Obtain work record data of each task node in multiple node groups; Based on the work record data, determine the activity data of each task node, wherein the activity data includes registration activity and work activity; Based on the active data of each task node, the grouping vector and the vector distance of the sub-center node, the management load of the sub-center node is determined; wherein the grouping vector is constructed based on the node information of the task node and the historical task distribution of the task node.
7. The method according to claim 6, characterized in that The updating of the node allocation map is also based on the management load update of the plurality of sub-center nodes and includes: Based on the working performance of multiple nodes to be allocated and the future workload of different node groups, multiple allocation graphs to be optimized are generated to form races to be optimized; Determining an evaluation value of the allocation map to be optimized based on the management load and the load balancing value; The allocation map to be optimized is updated by performing at least one round of iteration based on the evaluation value to determine the updated node allocation map.
8. The method according to claim 1, characterized in that The method also includes determining the node group where the task node is located based on the updated node allocation map, and forwarding the pending request issued by the task node to the target sub-center node, so that the sub-center node determines the node information of the node to be allocated according to the registration request; the target sub-center node is the sub-center node of the node group where the task node is located.
9. The method according to claim 8, characterized in that The method further comprises: Determine a first heartbeat cycle and a second heartbeat cycle; Sending the first heartbeat cycle to the secondary central node, and sending heartbeat information to the secondary central node in the first heartbeat cycle; Sending the second heartbeat cycle to the task node; so that the sub-center node sends the heartbeat information to the task node in the second heartbeat cycle.
10. The method according to claim 9, characterized in that The first heartbeat period is dynamically adjusted based on the robustness of the secondary central node; The second heartbeat period is dynamically adjusted based on the active data of multiple task nodes in the node group.
Citation Information
Patent Citations
Task processing method and device, computer equipment and storage medium
CN112685157A
Task allocation method and device, electronic equipment and storage medium
CN116185623A