Resource scheduling method and device for distributed message service cluster and electronic equipment
By collecting performance data in the distributed message service cluster, calculating load values, and dynamically adjusting scheduling strategies, the problem of low resource utilization caused by static partition allocation is solved, and dynamic reallocation of resources and optimization of system performance are achieved.
Patent Information
- Application Number
- CN202511147767.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-11
AI Technical Summary
In distributed message service clusters, the use of static partition allocation strategy results in low node resource utilization and high operation and maintenance costs, and a lack of automated mechanisms to respond to traffic changes.
By collecting performance data from each node in the distributed messaging service cluster, calculating partition and node load values, and dynamically adjusting scheduling strategies, including partition migration and node expansion, dynamic resource reallocation is achieved.
It improved node resource utilization, optimized system performance, reduced operation and maintenance costs, and enhanced the system's responsiveness to traffic changes.
Smart Images

Figure CN120935172A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of distributed and financial technology, and more specifically, to a resource scheduling method, apparatus, and electronic device for a distributed messaging service cluster. Background Technology
[0002] During the operation of a distributed messaging service cluster, when the traffic of producers and consumers surges, all read and write requests will flood into the nodes where some partition primary replicas reside, causing a sharp increase in the load on the node's CPU, memory, network, and other resources, and even triggering alarms. Because the existing technology uses a static partition allocation strategy, it can only be done manually and lacks an automated mechanism to respond to traffic changes in real time, resulting in low node resource utilization and high operation and maintenance costs.
[0003] The static partitioning strategy used in distributed message service clusters in related technologies has resulted in low node resource utilization, and no effective solution has yet been proposed. Summary of the Invention
[0004] The main objective of this application is to provide a resource scheduling method, apparatus, and electronic device for a distributed message service cluster, in order to solve the problem of low node resource utilization in related technologies where distributed message service clusters adopt a static partition allocation strategy.
[0005] To achieve the above objectives, according to one aspect of this application, a resource scheduling method for a distributed message service cluster is provided. The method includes: collecting performance data of each node in the distributed message service cluster based on a preset time interval; calculating the partition load value of each partition on each node and the node load value of each node based on the performance data; determining a target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node; and executing the target scheduling strategy on the distributed message service cluster.
[0006] Furthermore, calculating the partition load value of each partition on each node and the node load value of each node based on the performance data includes: calculating the partition load value of each partition based on the number of queries per second and transactions per second in the performance data; and calculating the node load value of each node based on the CPU utilization, memory utilization, disk read / write operations per second, and network throughput in the performance data.
[0007] Furthermore, determining the target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node includes: determining whether a target partition exists based on the partition load value of each partition and a preset threshold, and obtaining the determination result; if the determination result is that a target partition exists, then determining the target node from the nodes other than the node where the target partition is located based on the node load value of each node and preset rules; and determining the target scheduling strategy based on the partition load value of the target partition and the node load value of the target node.
[0008] Furthermore, the preset thresholds include a first threshold and a second threshold, with the second threshold being greater than the first threshold. The existence of a target partition is determined based on the partition load value of each partition and the preset thresholds. The determination results include: comparing the partition load value of each partition with the first threshold; if the partition load value is greater than the first threshold, then comparing the partition load value with the second threshold; if the partition load value is greater than the second threshold, then the existence of a target partition is taken as the determination result.
[0009] Furthermore, determining the target scheduling strategy based on the partition load value of the target partition and the node load value of the target node includes: summing the partition load value of the target partition and the node load value of the target node to obtain a first load value; and determining the target scheduling strategy based on the first load value and a third threshold.
[0010] Further, determining the target scheduling strategy based on the first load value and the third threshold includes: comparing the first load value with the third threshold; if the first load value is greater than the third threshold, then the target scheduling strategy is to expand the node capacity of the distributed message service cluster; if the first load value is less than or equal to the third threshold, then the target scheduling strategy is to migrate the target partition from the node where the target partition is located to the target node.
[0011] Furthermore, if the determination result indicates that a target partition exists, the method further includes: generating alarm information based on the attribute information of the target partition and the node information of the node where the target partition is located, and sending the alarm information to the target object.
[0012] To achieve the above objectives, according to another aspect of this application, a resource scheduling apparatus for a distributed message service cluster is provided. The apparatus includes: a collection unit, configured to collect performance data of each node in the distributed message service cluster based on a preset time interval; a calculation unit, configured to calculate the partition load value of each partition on each node and the node load value of each node based on the performance data; and a determination unit, configured to determine a target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node, and to execute the target scheduling strategy on the distributed message service cluster.
[0013] Furthermore, the computing unit includes: a first computing subunit, used to calculate the partition load value of each partition based on the number of queries per second and the number of transactions per second in the performance data; and a second computing subunit, used to calculate the node load value of each node based on the CPU utilization, memory utilization, disk read / write operations per second, and network throughput in the performance data.
[0014] Further, the determining unit includes: a judging subunit, used to judge whether a target partition exists based on the partition load value of each partition and a preset threshold, and obtain a judging result; a first determining subunit, used to determine the target node from the nodes other than the node where the target partition is located based on the node load value of each node and a preset rule if the judging result is that a target partition exists; and a second determining subunit, used to determine the target scheduling strategy based on the partition load value of the target partition and the node load value of the target node.
[0015] Furthermore, the preset threshold includes a first threshold and a second threshold, wherein the second threshold is greater than the first threshold. The judgment subunit includes: a first processing module, used to compare the partition load value of each partition with the first threshold respectively, and if the partition load value is greater than the first threshold, then compare the partition load value with the second threshold; and a second processing module, used to determine the existence of a target partition as the judgment result if the partition load value is greater than the second threshold.
[0016] Furthermore, the second determining subunit includes: a calculation module, used to sum the partition load value of the target partition and the node load value of the target node to obtain a first load value; and a determining module, used to determine the target scheduling strategy based on the first load value and a third threshold.
[0017] Furthermore, the determining module includes: a comparison submodule, used to compare the first load value with the third threshold; a first determining submodule, used to determine the target scheduling strategy as expanding the nodes of the distributed message service cluster if the first load value is greater than the third threshold; and a second determining submodule, used to determine the target scheduling strategy as migrating the target partition from the node where the target partition is located to the target node if the first load value is less than or equal to the third threshold.
[0018] Furthermore, the device also includes a sending unit, which, if the determination result indicates the existence of a target partition, generates alarm information based on the attribute information of the target partition and the node information of the node where the target partition is located, and sends the alarm information to the target object.
[0019] According to another aspect of the present invention, an electronic device is also provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the program executes the resource scheduling method of the distributed message service cluster described above during runtime.
[0020] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein the storage medium stores a program, wherein the program controls the device where the storage medium is located to execute the resource scheduling method of any of the above-mentioned distributed message service clusters during program execution.
[0021] In this embodiment, the following steps are employed: Based on a preset time interval, performance data of each node in the distributed message service cluster is collected; the partition load value of each partition on each node and the node load value of each node are calculated based on the performance data; a target scheduling strategy for the distributed message service cluster is determined based on the partition load value of each partition and the node load value of each node, and the target scheduling strategy is executed on the distributed message service cluster. This solves the technical problem of low node resource utilization in related technologies where distributed message service clusters employ static partition allocation strategies. In this solution, by collecting performance data from each node, a foundation for real-time monitoring and data analysis is provided, ensuring that the system can promptly perceive dynamic load changes, providing data support for intelligent scheduling decisions, thereby optimizing resource allocation and improving system performance; the scheduling strategy is adjusted in a timely manner based on the load calculation results, realizing dynamic reallocation of resources and improving node resource utilization. Attached Figure Description
[0022] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0023] Figure 1 A hardware structure block diagram of a computer terminal for implementing a resource scheduling method for a distributed message service cluster is shown.
[0024] Figure 2 This is a flowchart of a resource scheduling method for a distributed message service cluster provided according to an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the structure of a distributed message service cluster provided according to an embodiment of this application;
[0026] Figure 4 This is a schematic diagram of the resource scheduling process of a distributed message service cluster provided according to an embodiment of this application;
[0027] Figure 5 This is a schematic diagram of a resource scheduling device for a distributed message service cluster provided according to an embodiment of this application;
[0028] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.
[0032] Example 1
[0033] According to an embodiment of this application, a method embodiment for resource scheduling of a distributed message service cluster is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0034] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1A hardware block diagram of a computer terminal (or mobile device) for implementing a resource scheduling method for a distributed message service cluster is shown. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0035] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the resource scheduling method of the distributed message service cluster in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the resource scheduling method of the distributed message service cluster described above. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0037] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0038] The display may be a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0039] Under the aforementioned operating environment, this application provides the following: Figure 2 The resource scheduling method for the distributed message service cluster is shown. Figure 2 This is a flowchart of a resource scheduling method for a distributed message service cluster according to Embodiment 1 of this application. The resource scheduling method for the distributed message service cluster includes:
[0040] Step S201: Based on a preset time interval, collect performance data for each node of the distributed message service cluster.
[0041] Optionally, a performance data collector can be deployed on each node of the distributed messaging service cluster to collect performance data of each node at a high frequency (e.g., every 5 seconds), such as CPU utilization, network throughput, disk read / write, message production rate, message consumption rate, etc., and store it in the search and data analysis engine.
[0042] Step S202: Calculate the partition load value of each partition on each node and the node load value of each node based on the performance data.
[0043] Optionally, the intelligent partition allocator reads and parses performance data from the search and data analytics engine, calculating the partition load value for each partition on each node and the node load value for each node.
[0044] Step S203: Determine the target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node, and execute the target scheduling strategy for the distributed message service cluster.
[0045] Optionally, the target scheduling strategy is partition primary replica migration or node expansion. Based on the partition load value of each partition and the node load value of each node, it can be determined whether partition primary replica migration or node expansion is needed, and it is executed automatically.
[0046] In summary, by analyzing performance data, calculating load, and determining whether to migrate partition primary replicas or expand nodes, the problem of uneven resource allocation in distributed message service clusters is solved, enabling the system to automatically respond to traffic changes and maintain overall stability and efficiency.
[0047] Optionally, in the resource scheduling method for the distributed message service cluster provided in this application embodiment, calculating the partition load value of each partition on each node and the node load value of each node based on performance data includes: calculating the partition load value of each partition based on the number of queries per second and the number of transactions per second in the performance data; and calculating the node load value of each node based on the CPU utilization, memory utilization, disk read / write operations per second and network throughput in the performance data.
[0048] In an optional embodiment, the partition load value for each partition is calculated based on the queries per second and transactions per second in the performance data, as shown in the following formula:
[0049] Partition load value (days) = Query per second weight * (Average number of queries per second / Peak number of queries per second) + Transaction per second weight * (Average number of transactions per second / Peak number of transactions per second)
[0050] Among them, the weight of queries per second is 0.6, and the weight of transactions per second is 0.4.
[0051] The node load value for each node is calculated based on the CPU utilization, memory utilization, disk read / write operations per second, and network throughput data, using the following formula:
[0052] Node load value = (CPU weight * CPU utilization) + (Memory weight * Memory utilization) + (Disk weight * Disk read / write operations per second) + (Network weight * Network throughput)
[0053] The weighting rules (adjusted according to time period) are shown in the following example:
[0054] Weekday daytime (8:00-20:00): CPU weight = 0.4, memory weight = 0.3, disk weight = 0.1, network weight = 0.2.
[0055] At other times (20:00-8:00): CPU weight = 0.6, memory weight = 0.2, disk weight = 0.1, network weight = 0.1.
[0056] For example, assuming the current time period is a weekday daytime (8:00-20:00), and the node status is: CPU utilization = 0.8, memory utilization = 0.5, disk read / write operations per second = 0.2, network throughput = 0.1, the node load value calculated using the formula is (0.4*0.8) + (0.3*0.5) + (0.1*0.2) + (0.2*0.1) = 0.51.
[0057] By calculating the partition load value for each partition, an accurate assessment of the partition load is ensured, providing a reliable basis for subsequent resource scheduling. By comprehensively considering multiple performance indicators of the nodes, the actual load status of the nodes can be reflected more comprehensively, thereby making more reasonable resource scheduling decisions.
[0058] Optionally, in the resource scheduling method for the distributed message service cluster provided in this application embodiment, determining the target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node includes: determining whether a target partition exists based on the partition load value of each partition and a preset threshold, and obtaining a determination result; if the determination result indicates that a target partition exists, then determining a target node from nodes other than the node where the target partition is located based on the node load value of each node and a preset rule; and determining the target scheduling strategy based on the partition load value of the target partition and the node load value of the target node.
[0059] In an optional embodiment, the existence of a hotspot partition (i.e., a target partition) can be determined based on the partition load value of each partition and a preset threshold. The preset threshold is used to determine whether a partition is overloaded; for example, by comparing the partition load value and the preset threshold. If a hotspot partition exists, a target node is determined from the nodes other than the node containing the hotspot partition, based on the node load value of each node and preset rules. For example, nodes in other cluster nodes that meet the resource conditions (i.e., preset rules) are selected as candidate nodes, and the node with the lowest current load among these nodes is selected as the node to be migrated (i.e., the target node). The resource condition is: node load value + partition load value ≤ 1.0. Then, based on the partition load value of the hotspot partition and the node load value of the target node, it is determined whether to perform partition primary / replica migration or node expansion.
[0060] By intelligently analyzing the load of target partitions and target nodes, decisions can be made on whether to expand node capacity or migrate partitions, effectively avoiding resource waste and system performance degradation.
[0061] Optionally, in the resource scheduling method for the distributed message service cluster provided in this application embodiment, if the determination result is that a target partition exists, the method further includes: generating alarm information based on the attribute information of the target partition and the node information of the node where the target partition is located, and sending the alarm information to the target object.
[0062] In an optional embodiment, if a hotspot partition exists, an alarm message is generated based on the attribute information of the hotspot partition and the node information of the node where the hotspot partition is located, and the alarm message is sent to the target object. For example, the specific node address, message subject, partition name and other information are combined into an email and sent to the operation and maintenance personnel as an alarm.
[0063] The automated alarm mechanism saved labor costs and improved operation and maintenance efficiency.
[0064] Optionally, in the resource scheduling method of the distributed message service cluster provided in the embodiments of this application, the preset threshold includes a first threshold and a second threshold, the second threshold being greater than the first threshold. The determination of whether a target partition exists is based on the partition load value of each partition and the preset threshold, and the determination result includes: comparing the partition load value of each partition with the first threshold respectively; if the partition load value is greater than the first threshold, then comparing the partition load value with the second threshold; if the partition load value is greater than the second threshold, then the existence of a target partition is taken as the determination result.
[0065] In an optional embodiment, the preset threshold includes a first threshold and a second threshold. For example, the first threshold is 0.75 and the second threshold is 0.8. The partition load value of each partition is compared with the first threshold to determine whether the partition load value is greater than the warning value of 0.75. If the partition load value is greater than 0.75, the partition load value is compared with the second threshold to determine whether the partition load value is greater than the alarm value of 0.8. If the partition load value is greater than 0.8, it is determined that there is a hotspot partition, that is, the node traffic is too high.
[0066] By setting tiered thresholds, the system can distinguish different levels of load status and adopt corresponding scheduling strategies, which ensures the system's response speed and avoids unnecessary resource waste.
[0067] Optionally, in the resource scheduling method for the distributed message service cluster provided in this application embodiment, determining the target scheduling strategy based on the partition load value of the target partition and the node load value of the target node includes: summing the partition load value of the target partition and the node load value of the target node to obtain a first load value; and determining the target scheduling strategy based on the first load value and a third threshold.
[0068] In an optional embodiment, the partition load value of the target partition and the node load value of the target node are summed to obtain a first load value, i.e., the first load value = the partition load value of the target partition + the node load value of the target node. Based on the first load value and a third threshold, it can be determined whether to perform partition primary replica migration or node expansion. For example, if the third threshold is 0.8, the partition load value of the hot partition is 0.65, and the node load value of the node to be migrated is 0.1, then the first load value = 0.65 + 0.1 = 0.75. If the total node load of the node to be migrated (0.75) is less than 0.8, then it can be determined whether to perform partition primary replica migration.
[0069] By calculating the total load of the target partition and the target node, a quantitative basis is provided for subsequent decision-making. By comparing the first load value and the third threshold, the system can determine whether the target node can withstand the load of the target partition, thereby deciding whether to expand the node or migrate the partition, ensuring the rationality and effectiveness of resource scheduling.
[0070] Optionally, in the resource scheduling method for the distributed message service cluster provided in this application embodiment, determining the target scheduling strategy based on the first load value and the third threshold includes: comparing the first load value with the third threshold; if the first load value is greater than the third threshold, then determining the target scheduling strategy as expanding the node capacity of the distributed message service cluster; if the first load value is less than or equal to the third threshold, then determining the target scheduling strategy as migrating the target partition from the node where the target partition is located to the target node.
[0071] In an optional embodiment, a first load value is compared with a third threshold. If the first load value is greater than the third threshold, the target scheduling strategy is determined to be to expand the number of nodes in the distributed message service cluster. If the first load value is less than or equal to the third threshold, the target scheduling strategy is determined to be to migrate the hot partition from the node where the hot partition is located to the target node.
[0072] By automating node expansion or migration, dynamic resource reallocation is achieved, improving the adaptability of the distributed message service cluster to different load conditions, optimizing resource utilization, reducing manual operation and maintenance costs, and enhancing cluster stability and performance.
[0073] In an optional embodiment, Figure 3 This is a schematic diagram of the structure of a distributed message service cluster provided according to an embodiment of this application, such as... Figure 3As shown, the system includes producer clients, a distributed message service cluster, consumers, a performance data collector, and an intelligent partition allocator. Producer clients send messages to a specific message topic (e.g., the `service_a` topic) within the distributed message service cluster. This topic has multiple partitions, which serve as the physical carriers of the messages. Each partition uses a leader-follower replication mechanism. The leader replica communicates with clients and handles read / write requests, while the follower replicas pull synchronized data for redundancy and high availability. In other words, client read / write operations connect to the node containing the leader replica (i.e., the partition's leader). The distributed message service cluster stores messages, and consumer clients pull messages from the message topic. A consumer can consume one or more partitions, and the number of consumers in a consumer group is greater than the number of topic partitions. The performance data collector gathers performance data from each node, and the intelligent partition allocator parses this data, calculates partition and node load, and performs intelligent migration of the leader replica or triggers automatic cluster scaling based on the calculation results. This system automatically adjusts the service functionality of partitions when client traffic surges or decreases, meeting demand and providing optimized performance and resource utilization.
[0074] In an optional embodiment, Figure 4 This is a schematic diagram of the resource scheduling process of a distributed message service cluster provided according to an embodiment of this application, such as... Figure 4 As shown, the main steps include:
[0075] Step 401: Deploy a lightweight performance data collector on each node to collect core metrics such as CPU utilization, network throughput, disk read / write, message production rate, and message consumption rate at a frequency of 5 seconds / time, and write the data to the search and data analysis engine for data persistence, then proceed to step 402.
[0076] Step 402: The intelligent partition allocator reads and parses the performance data in the search and data analysis engine, calculates the partition load value of each partition and the node load value of each node, and proceeds to step 403.
[0077] Step 403: Determine if there are any uncalculated load nodes. If so, proceed to step 404; otherwise, the process ends.
[0078] Step 404: Determine whether the partition load value on the node is greater than the warning value of 0.75. If yes, proceed to step 405; otherwise, proceed to step 403.
[0079] Step 405: Determine whether the load value of the hotspot partition is greater than the alarm value of 0.8. If yes, proceed to step 406; otherwise, proceed to step 411.
[0080] Step 406: Node partition traffic overload alarm. The alarm device sends an email containing the specific node address, topic, partition, and other information to the operations and maintenance personnel, proceeding to step 407.
[0081] Step 407: Intelligent partition allocation. For the current hot partition, nodes in the cluster that are not currently located in the partition and meet the resource conditions are selected as candidate nodes. Then, a node is randomly selected from the candidate nodes, and the partition primary replica migration command is assembled. Proceed to step 408. The resource condition is: node load value + partition load value ≤ 1.0. For example, the node with the lowest current load is selected from the candidate nodes as the migration node. For example, if the load value of partition X is 0.65, and the load of node E is 0.1 (satisfying 0.1 + 0.65 ≤ 1.0), the system will migrate partition X from the original node to node E and update the load value of node E.
[0082] Step 408: Determine if the load of the migration node is greater than 0.8. If it is, proceed to step 409; otherwise, proceed to step 410.
[0083] Step 409: Automatically expand the cluster nodes by expanding the container on the node, then proceed to step 407;
[0084] Step 410: Using the migration scheme generated in step 407, the cluster is directly deployed via command to complete the migration of the partition primary replicas, and then proceed to step 404.
[0085] Step 411: Node partition traffic overload warning. Send an email containing the specific node address, topic, partition, and other information to the operations and maintenance personnel. The process ends here.
[0086] The resource scheduling method for a distributed message service cluster provided in this application deploys a lightweight performance data collector to collect core indicators at high frequency and combines this with a search and data analysis engine for real-time analysis, dynamically identifying hotspot partitions (e.g., load > 0.75). Through an intelligent migration algorithm (e.g., migrating a partition with a load of 0.65 to node E with a load of 0.1, satisfying 0.1 + 0.65 ≤ 1.0), dynamic resource reallocation is achieved, optimizing resource utilization, improving the cluster's adaptability to different load conditions, avoiding message backlog or service interruptions caused by single-node overload, and enhancing cluster stability. A two-level threshold mechanism (warning 0.75 / alarm 0.8) is adopted. During the warning stage, partition migration is automatically triggered, prioritizing load balancing to resolve local overload; during the alarm stage (> 0.8), automatic scaling is initiated, and migration strategies are continuously optimized, significantly improving operational efficiency and reducing manual maintenance costs. When making migration decisions, the system prioritizes candidate nodes with the lowest load, effectively avoiding blind scaling. By dynamically allocating resources instead of fixed reservations, hardware costs are reduced, while ensuring the system quickly adapts to traffic fluctuations, balancing scalability and flexibility.
[0087] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0088] Example 2
[0089] This application also provides a resource scheduling device for a distributed message service cluster. It should be noted that the resource scheduling device for a distributed message service cluster in this application can be used to execute the resource scheduling method for a distributed message service cluster provided in this application. The resource scheduling device for a distributed message service cluster provided in this application will be described below.
[0090] According to an embodiment of this application, a resource scheduling apparatus for a distributed message service cluster for implementing the above-described resource scheduling method for the distributed message service cluster is also provided, such as... Figure 5 As shown, the device includes: a data acquisition unit 501, a calculation unit 502, and a determination unit 503.
[0091] The acquisition unit 501 is used to collect performance data of each node in the distributed message service cluster based on a preset time interval.
[0092] The calculation unit 502 is used to calculate the partition load value of each partition on each node and the node load value of each node based on the performance data.
[0093] The determining unit 503 is used to determine the target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node, and to execute the target scheduling strategy for the distributed message service cluster.
[0094] The resource scheduling device for a distributed message service cluster provided in this application embodiment collects performance data of each node of the distributed message service cluster based on a preset time interval by a collection unit 501; a calculation unit 502 calculates the partition load value of each partition on each node and the node load value of each node based on the performance data; and a determination unit 503 determines the target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node, and executes the target scheduling strategy for the distributed message service cluster. This solves the technical problem of low node resource utilization in related technologies where distributed message service clusters use a static partition allocation strategy.
[0095] Optionally, in the resource scheduling device for the distributed message service cluster provided in this application embodiment, the computing unit 502 includes: a first computing subunit, used to calculate the partition load value of each partition based on the number of queries per second and the number of transactions per second in the performance data; and a second computing subunit, used to calculate the node load value of each node based on the CPU utilization, memory utilization, number of disk read / write operations per second and network throughput in the performance data.
[0096] Optionally, in the resource scheduling device for the distributed message service cluster provided in this application embodiment, the determining unit 503 includes: a judging subunit, used to judge whether a target partition exists based on the partition load value of each partition and a preset threshold, and obtain a judging result; a first determining subunit, used to determine a target node from nodes other than the node where the target partition is located based on the node load value of each node and a preset rule if the judging result indicates that a target partition exists; and a second determining subunit, used to determine a target scheduling strategy based on the partition load value of the target partition and the node load value of the target node.
[0097] Optionally, in the resource scheduling device for the distributed message service cluster provided in the embodiments of this application, the preset threshold includes a first threshold and a second threshold, wherein the second threshold is greater than the first threshold, and the judgment subunit includes: a first processing module, used to compare the partition load value of each partition with the first threshold respectively, and if the partition load value is greater than the first threshold, then compare the partition load value with the second threshold; and a second processing module, used to determine the existence of a target partition as the judgment result if the partition load value is greater than the second threshold.
[0098] Optionally, in the resource scheduling device for the distributed message service cluster provided in the embodiments of this application, the second determining subunit includes: a calculation module, used to sum the partition load value of the target partition and the node load value of the target node to obtain a first load value; and a determining module, used to determine a target scheduling strategy based on the first load value and a third threshold.
[0099] Optionally, in the resource scheduling device for the distributed message service cluster provided in the embodiments of this application, the determining module includes: a comparison submodule, used to compare a first load value with a third threshold; a first determining submodule, used to determine the target scheduling strategy as expanding the nodes of the distributed message service cluster if the first load value is greater than the third threshold; and a second determining submodule, used to determine the target scheduling strategy as migrating the target partition from the node where the target partition is located to the target node if the first load value is less than or equal to the third threshold.
[0100] Optionally, in the resource scheduling device for the distributed message service cluster provided in the embodiments of this application, the device further includes: a sending unit, used to generate alarm information based on the attribute information of the target partition and the node information of the node where the target partition is located if the determination result is that a target partition exists, and send the alarm information to the target object.
[0101] It should be noted that the acquisition unit 501, calculation unit 502, and determination unit 503 mentioned above correspond to steps S201 to S203 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal 10 provided in Embodiment 1.
[0102] Example 3
[0103] Embodiments of this application may provide an electronic device. Figure 6 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 (Only one is shown) processor 602, memory 604, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0104] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the above-described methods. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0105] The processor can access information and applications stored in the memory via the transmission device to perform the following steps: collecting performance data of each node in the distributed message service cluster based on a preset time interval; calculating the partition load value of each partition on each node and the node load value of each node based on the performance data; determining the target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node, and executing the target scheduling strategy for the distributed message service cluster.
[0106] The processor can access information and applications stored in memory via the transmission device to perform the following steps: calculate the partition load value for each partition based on the number of queries per second and transactions per second in the performance data; calculate the node load value for each node based on the CPU utilization, memory utilization, disk read / write operations per second, and network throughput in the performance data.
[0107] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: determine whether a target partition exists based on the partition load value of each partition and a preset threshold, and obtain the determination result; if the determination result is that a target partition exists, determine the target node from the nodes other than the node where the target partition is located based on the node load value of each node and preset rules; determine the target scheduling strategy based on the partition load value of the target partition and the node load value of the target node.
[0108] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: compare the partition load value of each partition with a first threshold; if the partition load value is greater than the first threshold, compare the partition load value with a second threshold; if the partition load value is greater than the second threshold, the existence of the target partition is taken as the judgment result.
[0109] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: sum the partition load value of the target partition and the node load value of the target node to obtain a first load value; determine the target scheduling strategy based on the first load value and a third threshold.
[0110] The processor can invoke information and applications stored in the memory through the transmission device to perform the following steps: compare a first load value with a third threshold; if the first load value is greater than the third threshold, determine the target scheduling strategy as expanding the nodes of the distributed message service cluster; if the first load value is less than or equal to the third threshold, determine the target scheduling strategy as migrating the target partition from the node where the target partition is located to the target node.
[0111] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: if the determination result is that the target partition exists, generate alarm information based on the attribute information of the target partition and the node information of the node where the target partition is located, and send the alarm information to the target object.
[0112] Those skilled in the art will understand that Figure 6 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones, tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 6 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 6 The different configurations shown.
[0113] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0114] Example 4
[0115] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the resource scheduling method of the distributed message service cluster provided in Embodiment 1.
[0116] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0117] This application also provides a computer program product, which, when executed on a data processing device, is suitable for performing resource scheduling method steps of a distributed message service cluster.
[0118] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0119] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0120] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0121] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0122] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0123] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0124] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A resource scheduling method for a distributed message service cluster, characterized in that, include: Based on preset time intervals, collect performance data for each node in the distributed message service cluster; Based on the performance data, calculate the partition load value of each partition on each node and the node load value of each node; The target scheduling strategy for the distributed message service cluster is determined based on the partition load value of each partition and the node load value of each node, and the target scheduling strategy is executed on the distributed message service cluster.
2. The method according to claim 1, characterized in that, Calculating the partition load value for each partition on each node and the node load value for each node based on the performance data includes: The partition load value for each partition is calculated based on the number of queries per second and transactions per second in the performance data. The node load value of each node is calculated based on the CPU utilization, memory utilization, disk read / write operations per second, and network throughput in the performance data.
3. The method according to claim 1, characterized in that, Determining the target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node includes: Based on the partition load value and preset threshold of each partition, it is determined whether a target partition exists, and the determination result is obtained; If the determination result indicates that the target partition exists, then based on the node load value of each node and the preset rules, the target node is determined from the nodes other than the node where the target partition is located. The target scheduling strategy is determined based on the partition load value of the target partition and the node load value of the target node.
4. The method according to claim 3, characterized in that, The preset threshold includes a first threshold and a second threshold, where the second threshold is greater than the first threshold. The existence of a target partition is determined based on the partition load value of each partition and the preset threshold, and the determination result includes: The partition load value of each partition is compared with the first threshold. If the partition load value is greater than the first threshold, the partition load value is compared with the second threshold. If the partition load value is greater than the second threshold, then the existence of the target partition is taken as the judgment result.
5. The method according to claim 3, characterized in that, Determining the target scheduling strategy based on the partition load value of the target partition and the node load value of the target node includes: The first load value is obtained by summing the partition load value of the target partition and the node load value of the target node; The target scheduling strategy is determined based on the first load value and the third threshold.
6. The method according to claim 5, characterized in that, Determining the target scheduling strategy based on the first load value and the third threshold includes: Compare the first load value with the third threshold; If the first load value is greater than the third threshold, then the target scheduling strategy is determined to be to expand the number of nodes in the distributed message service cluster. If the first load value is less than or equal to the third threshold, then the target scheduling strategy is determined to be to migrate the target partition from the node where the target partition is located to the target node.
7. The method according to claim 3, characterized in that, If the determination result indicates that the target partition exists, the method further includes: An alarm message is generated based on the attribute information of the target partition and the node information of the node where the target partition is located, and the alarm message is sent to the target object.
8. A resource scheduling device for a distributed message service cluster, characterized in that, include: The data collection unit is used to collect performance data of each node in the distributed message service cluster based on a preset time interval. A calculation unit is used to calculate the partition load value of each partition on each node and the node load value of each node based on the performance data; The determining unit is configured to determine the target scheduling strategy for the distributed message service cluster based on the partition load value of each partition and the node load value of each node, and execute the target scheduling strategy on the distributed message service cluster.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device where the computer-readable storage medium is located to perform the resource scheduling method of the distributed message service cluster according to any one of claims 1 to 7.
10. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program executes the resource scheduling method for a distributed message service cluster according to any one of claims 1 to 7 when it runs.