A scheduling method and device based on service node indicators, equipment and storage medium

By automatically scheduling service nodes based on multiple performance indicators in the IoT system, the problems of low efficiency and insufficient accuracy caused by reliance on manual evaluation in existing technologies are solved, and efficient and accurate load scheduling and data processing are achieved.

CN120750936BActive Publication Date: 2025-12-30GUANGZHOU SIE CONSULTING CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511221182.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-12-30
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

In existing IoT systems, the load scheduling of service nodes relies on manual assessment, which is inefficient and cannot guarantee accuracy, leading to traffic overload on hot nodes and data disorder or loss.

Method used

By obtaining resource load metrics of service nodes and IoT devices in the Kubernetes cluster, overloaded nodes are identified based on multiple metric scores, and device data is automatically migrated to low-load nodes. The main queue and buffer queue are used to distinguish consumed data.

Benefits of technology

It achieves efficient and accurate service node load scheduling, reduces the probability of data out-of-order and loss, and improves the stability and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750936B_ABST
    Figure CN120750936B_ABST
Patent Text Reader

Abstract

The application provides a scheduling method and device based on service node indicators, equipment and storage medium. The server resource load indicators of each service node in a Kubernetes cluster used to deploy an IoT system and the device indicators related to the accessed IoT devices in each service node are obtained, the overload score of each service node is determined in combination with the first indicator threshold and the second indicator threshold, if there is a first target service node with an overload score greater than the score threshold, a second target service node with the lowest overload score is determined, and the service nodes with low overload and load are automatically evaluated for migration; the device data of the target IoT device in the first target service node is migrated to the main queue of the second target service node for consumption, and the new device data uploaded by the target IoT device is cached to the buffer queue of the second target service node for consumption, so as to reduce the probability of data disorder or loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet of Things (IoT), and in particular to a scheduling method, apparatus, device, and storage medium based on service node indicators. Background Technology

[0002] Currently, IoT (Internet of Things) systems typically configure multiple nodes, each responsible for data access and data storage for a portion of the IoT devices. The allocation of IoT devices and nodes is usually statically scheduled. For example, new devices are connected to a preset service node by default. If the current preset service node is under high load, it can easily lead to traffic overload on hotspot nodes. This requires manual evaluation and migration of nodes, as well as manual scheduling to select new nodes for overloaded service nodes and reschedule IoT devices from the original overloaded nodes. This process is inefficient, relies heavily on human experience, and its accuracy cannot be guaranteed. Furthermore, even with manual scheduling, data from the original overloaded service node and data from new devices received by the migrated node will be in the same queue, easily leading to data disorder or loss. Summary of the Invention

[0003] This application provides a scheduling method, apparatus, device, and storage medium based on service node indicators to solve at least one problem existing in related technologies. The technical solution is as follows:

[0004] In a first aspect, embodiments of this application provide a scheduling method based on service node metrics, comprising:

[0005] Obtain the server resource load metrics of each service node in the Kubernetes cluster used to deploy the IoT system, as well as the device metrics related to the connected IoT devices in each of the service nodes.

[0006] The overload score of each service node is determined based on the server resource load index of each service node, the first index threshold corresponding to the server resource load index, each device index, and the second index threshold corresponding to the device index.

[0007] If there is a first target service node whose overload score is greater than the score threshold, determine the second target service node with the lowest overload score.

[0008] The device data of the target IoT device in the first target service node is migrated to the main queue of the second target service node for consumption, and the new device data uploaded by the target IoT device is cached in the buffer queue of the second target service node for consumption.

[0009] In one implementation, determining the overload score of each service node based on the server resource load index of each service node, a first index threshold corresponding to the server resource load index, each device index, and a second index threshold corresponding to the device index includes:

[0010] The server resource load indicators of each service node are compared with the first indicator threshold, and the device indicators are compared with the second indicator threshold; wherein, the server resource load indicators include CPU utilization, memory utilization, and node server health value, and the device indicators include main queue backlog, number of online devices, number of active devices, and data throughput.

[0011] If any one of the CPU utilization, memory utilization, and node server health value is greater than the first indicator threshold, or if any one of the main queue backlog, number of online devices, number of active devices, and data throughput is greater than the second indicator threshold, the overload score of the corresponding service node is determined to be the highest score.

[0012] Otherwise, based on the CPU utilization, memory utilization, node server health value and corresponding first sub-weight, combined with the main queue backlog, number of online devices, number of active devices, data throughput and corresponding second sub-weight, the overload score of the corresponding service node is calculated.

[0013] In one implementation, the node server health value is calculated and obtained in the following way:

[0014] Determine the first difference and the second difference between the preset values ​​and the CPU utilization rate and the memory utilization rate, respectively;

[0015] Determine the first difference The difference between the power and the second value The first product of powers of 1, and calculate the first product. The power of this number yields the node server's health value;

[0016] in, As a weighting factor, This is the sensitivity coefficient.

[0017] In one implementation, the calculation of the overload score of the corresponding service node based on the CPU utilization, memory utilization, node server health value, and corresponding first sub-weights, combined with the main queue backlog, number of online devices, number of active devices, data throughput, and corresponding second sub-weights, includes:

[0018] The first load score of the service node is obtained based on the CPU utilization, the first sub-weight corresponding to the CPU utilization, the first sub-threshold corresponding to the CPU utilization, the memory utilization, the first sub-weight corresponding to the memory utilization, the first sub-threshold corresponding to the memory utilization, the node server health value, the first sub-weight corresponding to the node server health value, and the first sub-threshold corresponding to the node server health value.

[0019] The second load score of the service node is obtained based on the main queue backlog, the second sub-weight corresponding to the main queue backlog, the second sub-threshold corresponding to the main queue backlog, the number of online devices, the second sub-weight corresponding to the number of online devices, the second sub-threshold corresponding to the number of online devices, the number of activated devices, the second sub-weight corresponding to the number of activated devices, the second sub-threshold corresponding to the number of activated devices, the data throughput, the second sub-weight corresponding to the data throughput, and the second sub-threshold corresponding to the data throughput.

[0020] The first load score is added to the second load score to obtain the overload score of the corresponding service node;

[0021] The first indicator threshold includes a first sub-threshold corresponding to the CPU utilization rate, the memory utilization rate, and the node server health value; the second indicator threshold includes a second sub-threshold corresponding to the main queue backlog, the number of online devices, the number of activated devices, and the data throughput.

[0022] In one implementation, obtaining the first load score of the corresponding service node based on the CPU utilization rate, the first sub-weight corresponding to the CPU utilization rate, the first sub-threshold corresponding to the CPU utilization rate, the memory utilization rate, the first sub-weight corresponding to the memory utilization rate, and the first sub-threshold corresponding to the memory utilization rate, and based on the node server health value, the first sub-weight corresponding to the node server health value, and the first sub-threshold corresponding to the node server health value, includes:

[0023] Determine a first ratio between the CPU utilization rate and a first sub-threshold corresponding to the CPU utilization rate, and determine a second product of the first ratio and a first sub-weight corresponding to the CPU utilization rate;

[0024] Determine a first ratio between the memory utilization rate and a first sub-threshold corresponding to the memory utilization rate, and determine a third product of the first ratio and a first sub-weight corresponding to the memory utilization rate;

[0025] Determine a first ratio between the node server health value and a first sub-threshold corresponding to the node server health value, and determine a fourth product of the first ratio and a first sub-weight corresponding to the node server health value;

[0026] The sum of the second product, the third product, and the fourth product is calculated to obtain the first load score of the corresponding service node.

[0027] In one implementation, the step of migrating the device data of the target IoT device in the first target service node to the main queue of the second target service node for consumption, and caching the new device data uploaded by the target IoT device to the buffer queue of the second target service node for consumption includes:

[0028] The device data of the target IoT device in the first target service node is migrated to the second target service node, and the device data is cached in the main queue of the second target service node for consumption;

[0029] During the process of caching the device data in the main queue and consuming the device data, the new device data uploaded by the target IoT device is cached in the buffer queue of the second target service node;

[0030] Once the main queue has consumed all the device data, the consumption thread of the buffer queue is started to consume the new device data in the buffer queue.

[0031] In one embodiment, the method further includes:

[0032] If the migration time of the device data exceeds the time threshold, or if the storage capacity of the buffer queue exceeds the storage threshold, the migration will be stopped and an alarm will be issued.

[0033] Secondly, embodiments of this application provide a scheduling device based on service node metrics, comprising:

[0034] The acquisition module is used to acquire the server resource load indicators of each service node in the Kubernetes cluster used to deploy the IoT system, as well as the device indicators related to the connected IoT devices in each service node.

[0035] The scoring module is used to determine the overload score of each service node based on the server resource load index of each service node, the first index threshold corresponding to the server resource load index, each device index, and the second index threshold corresponding to the device index.

[0036] The determination module is used to determine the second target service node with the lowest overload score if there is a first target service node with an overload score greater than the score threshold.

[0037] The migration module is used to migrate the device data of the target IoT device in the first target service node to the main queue of the second target service node for consumption, and to cache the new device data uploaded by the target IoT device to the buffer queue of the second target service node for consumption.

[0038] Thirdly, embodiments of this application provide an electronic device, including: a processor and a memory, wherein the memory stores instructions that are loaded and executed by the processor to implement the methods in any of the above-described embodiments.

[0039] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed, implements the methods in any of the above-described embodiments.

[0040] The beneficial effects of the above technical solution include at least the following:

[0041] By acquiring server resource load metrics and device metrics related to connected IoT devices for each service node in the Kubernetes cluster used to deploy the IoT system, an overload score is determined for each service node based on its server resource load metrics, the first threshold corresponding to the server resource load metrics, and the second threshold corresponding to each device metric. If a first target service node has an overload score greater than the score threshold, a second target service node with the lowest overload score is identified. The system automatically migrates the overloaded first target service node and the underloaded second target service node based on multiple metrics, which is more accurate and efficient than manual operation. Device data of target IoT devices in the first target service node is migrated to the main queue of the second target service node for consumption, while new device data uploaded by the target IoT devices is cached in the buffer queue of the second target service node for consumption. By introducing the buffer queue to cache new device data and distinguish it from the device data migrated from the main queue, the probability of data out of order or loss is reduced.

[0042] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, these aspects, embodiments, and features will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description

[0043] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this application and should not be construed as limiting the scope of this application.

[0044] Figure 1 This is a flowchart illustrating the steps of a scheduling method based on service node metrics according to an embodiment of this application.

[0045] Figure 2 A flowchart illustrating the steps of the scheduling method in this application;

[0046] Figure 3 This is a structural block diagram of a scheduling device based on service node indicators according to an embodiment of this application;

[0047] Figure 4 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0048] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0049] Reference Figure 1 The flowchart illustrates a scheduling method based on service node indicators according to an embodiment of this application. This scheduling method based on service node indicators may include at least steps S100-S400:

[0050] S100: Obtain the server resource load metrics of each service node in the Kubernetes cluster used to deploy the IoT system, as well as the device metrics related to the connected IoT devices in each service node.

[0051] S200. Determine the overload score of each service node based on the server resource load index of each service node, the first index threshold corresponding to the server resource load index, each device index, and the second index threshold corresponding to the device index.

[0052] S300. If there is a first target service node with an overload score greater than the score threshold, determine the second target service node with the lowest overload score.

[0053] S400: Migrate the device data of the target IoT device in the first target service node to the main queue of the second target service node for consumption, and cache the new device data uploaded by the target IoT device to the buffer queue of the second target service node for consumption.

[0054] The technical solution of this application embodiment obtains the server resource load indicators of each service node in the Kubernetes cluster used to deploy the IoT system, as well as the device indicators related to the connected IoT devices in each service node. Based on the server resource load indicators of each service node, the first indicator threshold corresponding to the server resource load indicators, each device indicator, and the second indicator threshold corresponding to the device indicators, the overload score of each service node is determined. If there is a first target service node with an overload score greater than the score threshold, the second target service node with the lowest overload score is determined. The first target service node with overload and the second target service node with low load are automatically evaluated based on multiple indicators and migrated. Compared with manual operation, this helps to ensure accuracy and efficiency. The device data of the target IoT device in the first target service node is migrated to the main queue of the second target service node for consumption, and the new device data uploaded by the target IoT device is cached in the buffer queue of the second target service node for consumption. By introducing the buffer queue to cache the new device data and distinguishing and consuming it from the device data migrated from the main queue, it helps to reduce the probability of data out of order or loss.

[0055] In this embodiment, the IoT system is deployed based on a Kubernetes cluster. The Kubernetes cluster has multiple service nodes. Since the server resources of a single service node are limited, IoT microservices need to be deployed across multiple service nodes. Each service node acts as a microservice, responsible for consuming a portion of the IoT device data, such as data access and data storage. For example, there are two consumer nodes, corresponding to the microservices iiot1 and iiot2 on the service nodes.

[0056] Optionally, each service node incorporates the NodeExport component to collect server resource load metrics and device metrics related to connected IoT devices. For example, server resource load metrics include, but are not limited to, CPU utilization, memory utilization, and node server health; device metrics include, but are not limited to, main queue backlog, number of online devices, number of active devices, and data throughput. It should be noted that the main queue backlog refers to the number of unprocessed messages accumulated in the main queue of all connected IoT devices on the service node; the number of online devices refers to the number of IoT devices currently reporting device data on the service node connected to the IoT system; the number of active devices refers to the number of IoT devices on the service node that have previously connected to the IoT system but have not reported device data; and the data throughput refers to the amount of data reported per second by all IoT devices on the service node.

[0057] In one implementation, the node server health value of each service node in step S100 can be obtained through the following steps:

[0058] S110. Determine the preset values ​​and the first and second differences between CPU utilization and memory utilization, respectively.

[0059] For example, the preset value is 1, and the first difference between the preset value 1 and the CPU utilization rate CPU_Util and the memory utilization rate Mem_Util is determined respectively. Second difference .

[0060] S120, Determine the first difference The difference between the power and the second value The first product of powers of 1, and calculate the first product. The power of 1 yields the node server's health value.

[0061] Specifically, node server health value The calculation formula is:

[0062]

[0063] in, The weighting factor (e.g., 0.5) indicates whether CPU utilization or memory utilization is prioritized. A value >0.5 indicates a greater focus on CPU utilization. <0.5 indicates a greater focus on memory usage; The sensitivity coefficient controls the shape of the health value curve. When β>1, the health value curve is smoother, and when β<1, the health value curve is sharper.

[0064] In one implementation, step S200 includes steps S210-S230:

[0065] S210. Compare the server resource load index of each service node with the first index threshold, and compare each device index with the second index threshold.

[0066] Optionally, the first indicator threshold This includes the first sub-threshold corresponding to CPU utilization, memory utilization, and node server health values, with the first sub-threshold corresponding to CPU utilization. The first sub-threshold corresponding to memory utilization and the first sub-threshold corresponding to the node server health value All are 80%, but this can be adjusted in other embodiments based on actual conditions; second indicator threshold This includes the second sub-threshold corresponding to the main queue backlog, the number of online devices, the number of active devices, and the data throughput. The second sub-threshold corresponding to the main queue backlog is... The second sub-threshold corresponding to the number of online devices The second sub-threshold corresponding to the number of activated devices Both are 10000, the second sub-threshold corresponding to data throughput. The rate is 100,000 messages per second.

[0067] S220. If any one of the following is greater than the first indicator threshold: CPU utilization, memory utilization, and node server health value; or if any one of the following is greater than the second indicator threshold: main queue backlog, number of online devices, number of active devices, and data throughput, the overload score of the corresponding service node is determined to be the highest score.

[0068] Optionally, if any one of the following—CPU utilization, memory utilization, and node server health value—is greater than the first indicator threshold (i.e., any one of the following—CPU utilization, memory utilization, and node server health value—is greater than its corresponding first sub-threshold), or if any one of the following—main queue backlog, number of online devices, number of active devices, and data throughput—is greater than the second indicator threshold (i.e., any one of the following—main queue backlog, number of online devices, number of active devices, and data throughput—is greater than its corresponding second sub-threshold)—then the overload score of the corresponding service node is directly increased. A score of 100 indicates that the system is definitely in an overload state.

[0069] S230. Otherwise, based on CPU utilization, memory utilization, node server health value and corresponding first sub-weight, combined with main queue backlog, number of online devices, number of active devices and data throughput and corresponding second sub-weight, calculate the overload score of the corresponding service node.

[0070] Optionally, if the conditions in S220 are not met, the overload score of each corresponding service node is further analyzed. Specifically, this includes steps S2301-S2303:

[0071] S2301. Based on CPU utilization, the first sub-weight corresponding to CPU utilization, the first sub-threshold corresponding to CPU utilization, memory utilization, the first sub-weight corresponding to memory utilization, the first sub-threshold corresponding to memory utilization, the node server health value, the first sub-weight corresponding to the node server health value, and the first sub-threshold corresponding to the node server health value, obtain the first load score of the corresponding service node.

[0072] 1. Determine the first ratio between CPU utilization and the first sub-threshold corresponding to CPU utilization, and determine the second product of the first ratio and the first sub-weight corresponding to CPU utilization.

[0073] 2. Determine the first ratio between memory utilization rate and the first sub-threshold corresponding to memory utilization rate, and determine the third product of the first ratio and the first sub-weight corresponding to memory utilization rate.

[0074] 3. Determine the first ratio between the node server health value and the first sub-threshold corresponding to the node server health value, and determine the fourth product of the first ratio and the first sub-weight corresponding to the node server health value.

[0075] 4. Calculate the sum of the second, third, and fourth products to obtain the first load score S1 for the corresponding service node. Therefore, the specific formula is:

[0076]

[0077] In the formula, This refers to server resource load metrics, for example, when m is 3 and i = 1, 2, 3. , , These correspond to CPU utilization, memory utilization, and node server health values, respectively. For the first sub-weight, for example, when i=1, 2, 3 , , These correspond to the first sub-weight of CPU utilization, the first sub-weight of memory utilization, and the first sub-weight of node server health value, respectively.

[0078] S2302. Obtain the second load score of the corresponding service node based on the main queue backlog, the second sub-weight corresponding to the main queue backlog, the second sub-threshold corresponding to the main queue backlog, the number of online devices, the second sub-weight corresponding to the number of online devices, the second sub-threshold corresponding to the number of online devices, the number of activated devices, the second sub-weight corresponding to the number of activated devices, the second sub-threshold corresponding to the number of activated devices, the data throughput, the second sub-weight corresponding to the data throughput, and the second sub-threshold corresponding to the data throughput.

[0079] Similarly, the second load score of the service node is calculated based on the formula. :

[0080]

[0081] In the formula, For equipment specifications, such as It is 4. =1, 2, 3, 4 , , , These correspond to the main queue backlog, number of online devices, number of activated devices, and data throughput, respectively. For example, the second sub-weight =1, 2, 3, 4 , , , These correspond to the second sub-weights of the main queue backlog, the second sub-weight of the number of online devices, the second sub-weight of the number of activated devices, and the second sub-weight of the data throughput, respectively.

[0082] S2303. Add the first load score to the second load score to obtain the overload score of the corresponding service node.

[0083] Therefore, the corresponding overload score of the service node is ultimately obtained. =S1+S2.

[0084] In one implementation, a scoring threshold is set, for example, 90. In step S300, the overload score of each service node is calculated. The overload score is compared with a scoring threshold. If a service node's overload score exceeds the threshold, it is designated as the first target service node, indicating that the first target service node is overloaded. Then, the service nodes are sorted from lowest to highest overload score, and the service node with the lowest overload score (ranked first) is automatically selected as the second target service node. For example, this second target service node could be an idle node that is not receiving device data. In some implementations, a Top 3 list of service nodes can be generated, allowing manual selection of a target service node from the list, eliminating the need for manual judgment of which service node is underloaded.

[0085] In one implementation, step S400 includes steps S410-S430:

[0086] S410. Migrate the device data of the target IoT device in the first target service node to the second target service node, and cache the device data in the main queue of the second target service node for consumption.

[0087] Optionally, each service node's IoT microservice has a built-in MQTT server. Taking the first target service node as an example, the device data (e.g., measurement point data) of the target IoT device is cached in the main queue of the first target service node at the device level. Consumer threads are responsible for retrieving and consuming device data from the main queue. When it is determined that the first target service node is overloaded, the device data of the target IoT device in the first target service node is migrated to the second target service node. Then, the second target service node also caches the device data in its main queue at the device level to continue consuming device data. It should be noted that since the second target service node caches data at the device level, for example, each device's device data can be distinguished using a device ID, even when the second target service node involves multiple IoT devices, it is still possible to clearly distinguish which IoT device the device data belongs to.

[0088] S420. During the process of caching device data and consumer device data in the main queue, the new device data uploaded by the target IoT device is cached in the buffer queue of the second target service node.

[0089] Optionally, this embodiment adds a new buffer queue to the service node. This buffer queue is used when the service node needs to receive data migrated from other service nodes. Specifically, during the process of caching device data and consuming device data in the main queue, since the first target service node migrates already received device data, and the target IoT device will also send new device data, the new device data uploaded by the target IoT device is cached in the buffer queue of the second target service node to avoid confusion between old and new device data, which could lead to data out-of-order or loss.

[0090] S430. When the main queue has finished consuming all the device data, start the consumption thread of the buffer queue to consume the new device data in the buffer queue.

[0091] Optionally, the second target service node will receive a corresponding notification after the device data migration is complete and the main queue has consumed all the device data. Specifically, when the main queue has consumed all the device data (or the message counter of the main queue is less than or equal to 0), the consumption thread of the buffer queue is started to consume the new device data in the buffer queue, thereby ensuring the orderly and accurate consumption of device data reported by the target IoT device.

[0092] In one embodiment, the scheduling method based on service node indicators of this application may further include step S500: when the migration time of device data exceeds a time threshold, or the storage amount of the buffer queue exceeds a storage threshold, the migration is stopped and an alarm is issued.

[0093] In this embodiment of the application, a time threshold is set. If the migration time of device data exceeds the time threshold or the storage amount of the buffer queue exceeds the storage threshold, it indicates that there is or will be a migration anomaly. The abnormal circuit breaker is triggered in time to stop the migration and issue an alarm to prompt manual switching. For example, manual switching to other service nodes can be performed based on the Top 3 service node list.

[0094] In this embodiment, each service node (microservice) is equivalent to a replica, realizing Kubernetes replica-level device scheduling and breaking through the limitations of traditional IoT system service-level scheduling. By setting up a dual-queue flow switching mechanism of main queue and buffer queue, the out-of-order pain point of data migration caused by system traffic imbalance is solved, and strong data consistency is guaranteed. Overload scores are evaluated by server resource load indicators and device indicators to determine low-load service nodes for dynamic scheduling, which helps to ensure accuracy and efficiency.

[0095] Reference Figure 2 Taking the microservice of the first target service node as iiot1 service and the microservice of the second target service node as iiot2 service as an example, the MQTT server load in iiot1 service and iiot2 service caches the device data of IoT devices to the main queue at the device level. After the scheduler and the scheduler device load determine iiot1 service and iiot2 service, they switch the connection, send scheduling instructions, and trigger data migration. When the main queue of the iiot1 service experiences backlog, the first target service node may be overloaded (e.g., the overload score exceeds the score threshold). The worker thread pool determines whether a migration backup is currently in progress (when a migration backup is needed, the corresponding device is recorded in the backup record, i.e., the migration backup record set). If no record is found, it means that migration is not needed at present, and the process proceeds normally for downstream business. If so, it indicates an overloaded state, and device data migration is required. At this time, a command is sent via migration sender and RPC communication. After the migration receiver of the iiot2 service receives the migrated device data, it is cached (put) into the main queue at the device dimension of the IoT device. The worker thread pool of the iiot2 service normally processes and consumes the device data in the main queue for downstream business. Meanwhile, the new device data (new device data) of the target IoT device corresponding to the device data migrated by the iiot1 service is cached in the buffer queue of the iiot2 service. When the iiot2 service determines that the migration is complete, the main queue has consumed all device data, or the message counter of the main queue is less than or equal to 0, the threads of the standby thread pool consume the new device data in the buffer queue for downstream business. Specifically, when all target IoT devices go offline or the poll queue remains empty for a certain period of time during the consumption process of the standby thread pool, the migration process is completed, the migration device record Set is cleared, and the standby thread pool is stopped.

[0096] Reference Figure 3 The diagram illustrates a structural block diagram of a scheduling device based on service node metrics according to an embodiment of this application. The device may include:

[0097] The acquisition module is used to acquire the server resource load indicators of each service node in the Kubernetes cluster used to deploy the IoT system, as well as the device indicators related to the connected IoT devices in each service node.

[0098] The scoring module is used to determine the overload score of each service node based on the server resource load index of each service node, the first index threshold corresponding to the server resource load index, each device index, and the second index threshold corresponding to the device index.

[0099] The determination module is used to determine the second target service node with the lowest overload score if there is a first target service node with an overload score greater than the score threshold.

[0100] The migration module is used to migrate the device data of the target IoT device in the first target service node to the main queue of the second target service node for consumption, and to cache the new device data uploaded by the target IoT device to the buffer queue of the second target service node for consumption.

[0101] The functions of each module in the device of this application embodiment can be found in the corresponding description in the above method, and will not be repeated here.

[0102] Reference Figure 4 The diagram illustrates a structural block diagram of an electronic device according to an embodiment of this application. The electronic device includes a memory 310 and a processor 320. The memory 310 stores instructions that can be executed on the processor 320. The processor 320 loads and executes these instructions to implement the scheduling method based on service node metrics in the above embodiment. The number of memories 310 and processors 320 can be one or more.

[0103] In one embodiment, the electronic device further includes a communication interface 330 for communicating with external devices and exchanging data. If the memory 310, processor 320, and communication interface 330 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0104] Optionally, in a specific implementation, if the memory 310, processor 320 and communication interface 330 are integrated on a single chip, the memory 310, processor 320 and communication interface 330 can communicate with each other through an internal interface.

[0105] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the scheduling method based on service node indicators provided in the above embodiments.

[0106] This application also provides a chip, which includes a processor for calling and executing instructions stored in a memory, causing a communication device on which the chip is installed to perform the method provided in this application.

[0107] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in the application embodiment.

[0108] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting the Advanced Reduced Instruction Set Computing (RISC) machine (ARM) architecture.

[0109] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).

[0110] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0111] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0112] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0113] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functionality involved.

[0114] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can fetch and execute instructions from, an instruction execution system, apparatus or device).

[0115] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.

[0116] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.

[0117] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for scheduling based on service node metrics, the method comprising: The method comprises: obtaining server resource load indicators of each service node in a Kubernetes cluster for deploying an IoT system and device indicators related to accessed IoT devices in each service node, the server resource load indicators including CPU usage, memory usage, and node server health value; The node server health value is calculated and obtained by: determining a preset value and a first difference and a second difference between the CPU utilization rate and the memory utilization rate; determining the first difference... The difference between the power and the second value The first product of powers of 1, and calculate the first product. The power of this factor yields the node server's health value; where... As a weighting factor, Sensitivity coefficient; determining an overload score of each service node according to the server resource load indicators, first indicator threshold values corresponding to the server resource load indicators, each device indicator, and second indicator threshold values corresponding to the device indicators; if the CPU usage, the memory usage, and the node server health value are all less than or equal to the corresponding first indicator threshold values, or if all the device indicators are less than or equal to the corresponding second indicator threshold values, the overload score includes a first load score, and the calculation process of the first load score includes: determining a first ratio of the CPU usage to a first sub-threshold value corresponding to the CPU usage, and determining a second product of the first ratio and a first sub-weight corresponding to the CPU usage; determining a first ratio of the memory usage to a first sub-threshold value corresponding to the memory usage, and determining a third product of the first ratio and a first sub-weight corresponding to the memory usage; determining a first ratio of the node server health value to a first sub-threshold value corresponding to the node server health value, and determining a fourth product of the first ratio and a first sub-weight corresponding to the node server health value; calculating a sum of the second product, the third product, and the fourth product to obtain the first load score of the corresponding service node; the first indicator threshold values include the first sub-threshold values corresponding to the CPU usage, the memory usage, and the node server health value; if there is a first target service node with an overload score greater than a score threshold value, determining a second target service node with the lowest overload score; migrating device data of a target IoT device in the first target service node to a main queue of the second target service node for consumption, and caching new device data uploaded by the target IoT device to a buffer queue of the second target service node for consumption.

2. The method of claim 1, wherein: The method further comprises: comparing the server resource load indicators of each service node with the first indicator threshold values, and comparing each device indicator with the second indicator threshold values; wherein the device indicators include main queue backlog, number of online devices, number of active devices, and data throughput. If one of the CPU usage, the memory usage, and the node server health value is greater than the first index threshold value, or one of the master queue backlog, the online device number, the active device number, and the data throughput is greater than the second index threshold value, it is determined that the overload score of the corresponding service node is the highest score. Otherwise, the overload score of the corresponding service node is calculated according to the CPU usage, the memory usage, the node server health value, and the corresponding first sub-weight, in combination with the master queue backlog, the online device number, the active device number, and the data throughput, and the corresponding second sub-weight, wherein the overload score includes a first load score and a second load score.

3. The method of claim 2, wherein: The calculation of the overload score of the corresponding service node according to the CPU usage, the memory usage, the node server health value, and the corresponding first sub-weight, in combination with the master queue backlog, the online device number, the active device number, and the data throughput, and the corresponding second sub-weight, includes: The second load score of the corresponding service node is obtained according to the master queue backlog, the second sub-weight corresponding to the master queue backlog, the second sub-threshold corresponding to the master queue backlog, the online device number, the second sub-weight corresponding to the online device number, the second sub-threshold corresponding to the online device number, the active device number, the second sub-weight corresponding to the active device number, the second sub-threshold corresponding to the active device number, the data throughput, the second sub-weight corresponding to the data throughput, and the second sub-threshold corresponding to the data throughput; The first load score and the second load score are added to obtain the overload score of the corresponding service node; The second index threshold value includes the second sub-threshold corresponding to the master queue backlog, the online device number, the active device number, and the data throughput.

4. The method of claim 1, wherein the method further comprises: The migration of the device data of the target IoT device in the first target service node to the master queue of the second target service node for consumption, and the caching of new device data uploaded by the target IoT device to the buffer queue of the second target service node for consumption, include: The device data of the target IoT device in the first target service node is migrated to the second target service node, and the device data is cached to the master queue of the second target service node for consumption; In the process of caching and consuming the device data in the master queue, new device data uploaded by the target IoT device is cached to the buffer queue of the second target service node; When the master queue consumes all the device data, the consumption thread of the buffer queue is started to consume the new device data in the buffer queue.

5. The method of claim 1, wherein: The method further includes: When the migration time of the device data exceeds a time threshold value, or the storage amount of the buffer queue exceeds a storage threshold value, the migration is aborted and an alarm is issued.

6. A service node index-based dispatching apparatus, characterized by comprising: It includes: an acquisition module, configured to acquire server resource load indicators of each service node in a Kubernetes cluster for deploying an IoT system and device indicators related to accessed IoT devices in each service node, the server resource load indicators comprising CPU usage, memory usage and node server health value; The node server health value is calculated and obtained by: determining a preset value and a first difference and a second difference between the CPU utilization rate and the memory utilization rate; determining the first difference... The difference between the power and the second value The first product of powers of 1, and calculate the first product. The power of this factor yields the node server's health value; where... As a weighting factor, Sensitivity coefficient; a scoring module, configured to determine an overload score of each service node according to the server resource load indicators of each service node, first indicator threshold values corresponding to the server resource load indicators, each device indicator and second indicator threshold values corresponding to each device indicator; if the CPU usage, the memory usage and the node server health value are all less than or equal to the corresponding first indicator threshold values, or if all the device indicators are less than or equal to the corresponding second indicator threshold values, the overload score comprises a first load score, and the calculation process of the first load score comprises: determining a first ratio of the CPU usage to a first sub-threshold value corresponding to the CPU usage, and determining a second product of the first ratio and a first sub-weight corresponding to the CPU usage; determining a first ratio of the memory usage to a first sub-threshold value corresponding to the memory usage, and determining a third product of the first ratio and a first sub-weight corresponding to the memory usage; determining a first ratio of the node server health value to a first sub-threshold value corresponding to the node server health value, and determining a fourth product of the first ratio and a first sub-weight corresponding to the node server health value; calculating a sum of the second product, the third product and the fourth product to obtain the first load score of the corresponding service node; the first indicator threshold values comprise the first sub-threshold values corresponding to the CPU usage, the memory usage and the node server health value; a determination module, configured to determine a second target service node with the lowest overload score if there is a first target service node with an overload score greater than a score threshold value; a migration module, configured to migrate device data of a target IoT device in the first target service node to a main queue of the second target service node for consumption, and cache new device data uploaded by the target IoT device to a buffer queue of the second target service node for consumption.

7. An electronic device, comprising: comprise: a processor and a memory, the memory storing instructions which are loaded and executed by the processor to implement the method of any one of claims 1-5.

8. A computer readable storage medium, the computer readable storage medium storing a computer program which, when executed, implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Method and device for globally scheduling traffic, and electronic equipment

    CN107872402A

  • A method and system for scheduling physical resources based on kubernets

    CN109167835A