Intelligent load balancing scheduling system and method based on redundant micro service

Through intelligent load balancing scheduling system and ARIMA model prediction, dynamically adjust the status of the main and spare nodes, solving the problem of low resource utilization of spare nodes in traditional power monitoring systems, and improving the stability and response efficiency of the system.

CN120371548AActive Publication Date: 2025-07-25NANJING NARI IND CONTROL TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510874294.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In traditional power monitoring systems, the resource utilization rate of standby nodes is low, and the response delay of the main nodes at high loads, resulting in insufficient system reliability and availability.

Method used

The intelligent load balancing scheduling system based on redundant microservices is adopted. Through parameter configuration, service communication, data collection, fault diagnosis, load statistics calculation and prediction modules, dynamic load balancing of the main and backup nodes and rapid disaster recovery are realized. The future load prediction is carried out in combination with the ARIMA model, and the main and backup status is dynamically adjusted.

Benefits of technology

Load balancing between the main and spare nodes is realized, system resource utilization and stability is improved, node overload is avoided, response efficiency and adaptability are improved, and system impact resistance is enhanced in high concurrency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371548A_ABST
    Figure CN120371548A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent load balancing scheduling system and method based on redundant micro-services. The system comprises a parameter configuration module, a service communication module, a data acquisition module, a fault diagnosis module, a storage pool, a load statistical calculation module, a prediction module and a main and standby instance switching module. The system monitors node loads in real time by periodically collecting micro-service CPU data, pre-judges future loads in combination with a preset threshold rule and a prediction model, and dynamically adjusts main and standby service states. When a micro-service fault or node overload is detected, main-standby switching is automatically triggered, and resource overload in a service peak period is avoided; according to the invention, intelligent load balancing of the main and standby nodes is realized, the problem of non-uniform resource allocation in a traditional architecture is solved, and the resource utilization efficiency and the operation stability of the system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of rail transit automation, and particularly relates to a redundant microservice-based intelligent load balancing scheduling system and method. Background Art

[0002] In recent years, with the rapid development of the urban rail transit industry, the power monitoring system, as the core business to ensure the power supply safety of rail transit, has increasingly higher requirements for reliability and real-time performance. The traditional power monitoring system is based on the SCADA (Supervisory Control And Data Acquisition) architecture and undertakes key functions such as real-time data acquisition, equipment status monitoring, remote control, and fault warning. It is an important support for the stable operation of the rail transit power supply system. At the same time, the wide application of the urban rail transit cloud platform has promoted the intensive and virtualized transformation of the industry's IT infrastructure, aiming to improve the overall operation and maintenance efficiency and reduce the construction cost through resource sharing and elastic scheduling.

[0003] In the implementation architecture of the traditional SCADA system, the business functions are usually completed by a single service process in a centralized single-process mode, and the business system is deployed in a primary-backup redundant dual-machine hot standby mode. Specifically, two business nodes run simultaneously. The service on one node serves as the primary instance, responsible for real-time data processing, business operations, and external services; the service on the other node serves as the backup instance and is in a standby state. The state synchronization between the primary and backup nodes is achieved through a heartbeat detection mechanism. When the primary node fails, the backup node can quickly switch to the primary role, thus ensuring the high availability of the system. In the urban rail transit cloud environment, this architecture is migrated to run in virtual machines, and the primary node and the backup node are respectively deployed on independent virtual machines, but the resource allocation strategy still follows the traditional mode, and the utilization rate of the computing resources of the backup node is significantly lower than that of the primary node.

[0004] However, the prior art has a prominent problem of unbalanced resource utilization. Since the backup node is idle under normal circumstances, its virtual machine resources (such as CPU and memory) are in a low load level (10 - 15%) for a long time, while the primary node may face an instantaneous load surge (exceeding 90%) during peak business periods (such as power on / off operations or multi-client concurrent accesses), resulting in response delays, control timeouts, and even service anomalies. This unreasonable resource allocation not only fails to fully utilize the redundant capacity of the backup node but also makes the primary node lack elastic expansion support under sudden loads, seriously affecting the reliability and availability of the system. Summary of the Invention

[0005] Objective of the Invention: The objective of the present invention is to provide a redundant microservice intelligent load balancing scheduling system that can dynamically adjust the primary node allocation strategy before the arrival of the business peak period, achieve dynamic load balancing between the primary and standby nodes, and improve the reliability and practicality of the system; on the other hand, provide a redundant microservice intelligent load balancing scheduling method.

[0006] Technical Solution: The intelligent load balancing scheduling system described in the present invention includes: A parameter configuration module, which is used to set the IP address of the service node, communication port, microservice CPU statistics period, heartbeat interval, failure determination rule, CPU threshold switching rule, primary-standby switching rule, prediction model library selection, prediction statistics period, and prediction early switching time; A service communication module, which is used to establish a communication link with other nodes, transmit the microservice data of the current node, and receive the data of the peer node, and store it in the communication buffer; A data acquisition module, which is used to periodically collect the CPU occupancy ratio of each microservice of the current node, and read the data of the peer node from the communication buffer, and store it in the storage pool; A fault diagnosis module, which is used to detect the failure state of the microservice. When the data acquisition module cannot obtain the CPU occupancy ratio of the microservice, mark the microservice as a fault state; A storage pool, which is used to store the real-time state data of the microservice and provide a fast retrieval function; A load statistics calculation module, which is used to calculate the real-time load of the node, compare the load balancing parameters, and generate a primary-standby switching instruction; A prediction module, which is used to regularly read the real-time information of the microservice from the storage pool and store it in the time series database to form historical load data, calculate according to the prediction model parameters set by the parameter configuration module, generate a future business load prediction result, and send an early adjustment instruction to the primary-standby instance switching module based on the prediction result; A primary-standby instance switching module, which is used to execute the dynamic switching of the primary-standby state of the microservice.

[0007] This intelligent load balancing scheduling system realizes flexible strategy customization through the parameter configuration module, ensures real-time data synchronization between nodes through the service communication module, accurately monitors the CPU load of microservices through the data acquisition module, quickly identifies abnormal services through the fault diagnosis module, efficiently manages real-time state data through the storage pool, dynamically evaluates the node load and triggers the balancing strategy through the load statistics calculation module, intelligently predicts the traffic trend based on historical data through the prediction module, and finally executes dynamic adjustment through the primary-standby instance switching module. Overall, it realizes the real-time load balancing of the microservice cluster, fast fault tolerance, and optimization of resource utilization, effectively avoids node overload, and improves the system stability and response efficiency.

[0008] Preferably, the service communication module realizes data synchronization between nodes through multi-threaded Socket communication, including: Allocating an independent thread and a read-write buffer for each target node; Organizing the local microservice data packet LocalPacket of this node according to the heartbeat interval and sending it; Receiving the data packet of the peer node and parsing it into structured data, and updating it to the storage pool.

[0009] This service communication module adopts a multi-threaded Socket communication mechanism. By allocating independent threads and read-write buffers for each target node, it ensures stable data transmission under high concurrency; organizes and sends local microservice data packets (LocalPacket) based on a preset heartbeat interval to achieve real-time status synchronization between nodes; at the same time, receives and parses the peer data packets and quickly updates them to the storage pool, effectively reducing communication latency, improving the consistency of cluster data, and ensuring the real-time performance and reliability of load balancing scheduling.

[0010] Preferably, the data acquisition module polls the list of microservice names by calling the operating system API function, obtains the CPU occupancy information of each microservice, and uses the node name + microservice name as the unique key, and stores the JSON format data containing the CPU occupancy, service status, and timestamp as the value in the two-dimensional structure data table of the storage pool.

[0011] This data acquisition module accurately obtains the real-time CPU occupancy of each microservice through operating system API polling, uses "node name + microservice name" as the unique key, and stores it in a two-dimensional data table in a structured JSON format to ensure the uniqueness and traceability of monitoring data. This design realizes the efficient acquisition and standardized storage of microservice resource occupancy, providing high-timeliness and low-redundancy basic data support for load statistics and predictive analysis.

[0012] Preferably, the load statistics calculation module traverses the microservice status data in the storage pool and performs the following operations: If all microservices are in the standby state, set the highest-priority microservice as the primary according to the priority; if there are multiple primary conflicts, downgrade the redundant instances to the standby state according to the priority.

[0013] This load statistics calculation module intelligently executes the primary / standby switching strategy by dynamically analyzing the microservice status data in the storage pool: when it detects that all microservices are in the standby state, it automatically promotes the highest-priority instance to the primary; if there are multiple primary conflicts, it automatically downgrades the redundant instances according to the priority to ensure that the optimal primary / standby distribution is always maintained. This mechanism effectively avoids resource contention and service conflicts, and realizes intelligent dynamic balancing of the load while ensuring high availability.

[0014] Preferably, the prediction module uses the ARIMA model to perform time series analysis on historical load data, predicts the time points of future business peak periods, and triggers the primary and standby instance switching module to adjust the microservice status based on the predicted advance switching time.

[0015] The prediction module intelligently analyzes historical load data through the ARIMA time series model, accurately predicts future business peak periods, and actively triggers the primary and standby status adjustment based on a preset advance switching time. This design realizes forward-looking load balancing scheduling, effectively avoids performance degradation caused by insufficient resources during business peak periods, and significantly improves the adaptive ability and stability of the system in scenarios with traffic fluctuations.

[0016] The intelligent load balancing scheduling method described in the present invention includes the following steps: (1) The parameter configuration module obtains the microservice startup sequence and deploys multiple independent microservices split by the SCADA system on at least two business nodes; after each microservice starts, it subscribes to the message channel and runs as a standby instance by default; (2) The data collection module obtains the microservice name list, polls the operating system API at a preset statistical period to collect the CPU occupancy information of each microservice; writes the collected CPU occupancy information, the corresponding microservice status flag, and the collection timestamp into the storage pool; if the collection fails, it is marked as a fault state; (3) The service communication module obtains the node IP and communication port, and establishes a multi-threaded Socket connection; obtains the CPU occupancy information and status flag of each microservice from the storage pool at the heartbeat interval, encapsulates them into a structured data packet, and sends them to the peer node through the read-write buffer mechanism, and updates them to the local storage pool after parsing the received peer data into the same structure; (4) The load statistics calculation module periodically traverses the CPU occupancy information and status flag of each microservice in the storage pool: dynamically adjusts the distribution of primary and standby instances according to a preset strategy, writes the status adjustment result back to the storage pool, and provides a query service for the primary and standby status of each microservice through the status query interface; (5) The load statistics calculation module obtains the load threshold rule and calculates the real-time CPU load of each node; if the total CPU load of the primary node exceeds the threshold and the difference from the standby node exceeds the preset range, the microservice with the highest CPU occupancy in this node is selected to be downgraded to the standby state, and the corresponding microservice is promoted from the standby node to the primary state, and the status flag of the corresponding microservice in the storage pool is updated; (6) The prediction module obtains the ARIMA model parameters, based on the historical microservice CPU occupancy information in the time series database, uses a sliding window to statistically analyze the monthly peak pattern, predicts the business peak period of high-load microservices, and triggers the adaptive switching of primary and standby instances in advance.

[0017] The intelligent load balancing and scheduling method proposed by the present invention realizes efficient resource scheduling through multi-step collaboration: in the initialization stage (step 1), a highly available infrastructure is built through multi-node microservice deployment and standby state presetting; in the real-time monitoring stage (steps 2-3), periodic CPU collection and multi-node data synchronization are adopted to ensure the real-time and accuracy of state perception; in the dynamic scheduling stage (steps 4-5), the master-slave distribution is intelligently adjusted based on the load threshold to achieve load balancing among nodes; in the predictive scheduling stage (step 6), the business peak is predicted through time series analysis, and resource allocation is completed in advance. This method forms a complete closed-loop from infrastructure deployment, real-time state monitoring to dynamic load balancing and predictive scheduling, effectively solving problems such as uneven utilization of master-slave resources and delayed response to burst traffic in traditional architectures, and significantly improving the overall stability and resource utilization efficiency of the system.

[0018] Preferably, step 2 includes: The data collection module calls the GetServiceList interface to obtain the microservice name list from the parameter configuration module, and obtains the statistical period of each microservice through the GetCircle interface; Poll the operating system API according to the statistical period to obtain the CPU occupancy ratio information of each microservice; Store the obtained data in a two-dimensional structured data table with the node name + microservice name as the unique key and the JSON format data containing the CPU occupancy ratio, service status, and timestamp as the value; where: When the CPU occupancy ratio information is successfully obtained, if the corresponding key-value exists in the storage pool, update the timestamp field, otherwise add a new record and initialize the service status to standby; When the acquisition of the CPU occupancy ratio information fails, mark the status of the corresponding microservice as faulty and do not update the timestamp field; The storage pool is a memory database, a relational database, or a distributed database; The service status includes three states: primary, standby, and faulty.

[0019] The accurate acquisition of the microservice running status is realized through a dynamic polling mechanism: the data collection module calls the system API according to the configured period to obtain the CPU occupancy ratio of each microservice, and adopts a structured storage scheme to ensure data integrity and timeliness. Through the intelligent status update logic and flexible data storage support, a highly reliable microservice monitoring foundation is built, providing real-time and accurate data support for subsequent load balancing decisions. This design not only ensures the real-time nature of state monitoring but also enhances the fault tolerance of the system through rapid fault identification.

[0020] Preferably, step 3 includes: The service communication module calls the GetNodeInfo interface to obtain the IP address, communication port, and heartbeat interval parameters of each node; Create corresponding communication threads according to the number of nodes, allocate independent read and write buffer areas for each thread, and establish a TCP connection through Socket; Periodically obtain the microservice operation data of this node from the microservice status storage table according to the heartbeat interval, and organize it into LocalPacket data packets with the node name + microservice name as the smallest unit; Store the LocalPacket in the read buffer area, and the communication thread sends it to the target node through Socket, and clears the cached data after receiving the confirmation reply, otherwise trigger the retransmission mechanism; Receive the data packets sent by the peer node and store them in the write buffer area; Call the ParsePacket interface to parse the content of the data packet, and convert it into a two-dimensional key-value structure identical to the microservice status storage table, where the key is the node name + microservice name, and the value is JSON data containing CPU occupancy, service status, and timestamp; Temporarily store the parsed data in the temporary memory mapping area first, and then batch update it to the microservice status storage table: if the corresponding key exists in the table, update the field, otherwise add a new record; after the update is completed, clear the corresponding data in the temporary memory mapping area.

[0021] Implement efficient data synchronization between nodes through a multi-threaded Socket communication architecture: based on dynamic thread allocation and dual buffer design (isolation of read / write buffer areas), combined with the heartbeat mechanism to periodically encapsulate and transmit LocalPacket data packets to ensure the real-time and reliable transmission of microservice status information; through an intelligent parsing and storage mechanism (temporarily store in the temporary mapping area first and then batch update) and an automatic retransmission strategy, while ensuring data consistency, significantly improve the cross-node communication efficiency, provide the system with low-latency and highly available distributed state synchronization capabilities, and effectively support the real-time requirements of load balancing decisions.

[0022] Preferably, step 4 includes: The load statistics calculation module periodically obtains the microservice operation data of each node from the microservice status storage table and stores it in the memory according to the two-dimensional structure of [node identifier, [microservice identifier, operation data]]; Traverse the memory data structure: if all microservice instances are in the standby state, set the microservice instance with the highest priority to the primary state according to the preset priority; if multiple primary instances of the same microservice are detected, retain the instance with the highest priority according to the priority, and the rest are downgraded to the standby state; Write the adjusted status data back to the microservice status storage table, and provide a real-time external master-slave status query service through the status query interface, and the interface returns its current master-slave status according to the input microservice identifier.

[0023] Implement dynamic primary / standby management through the intelligent status arbitration mechanism: The load statistics calculation module periodically scans the microservice operation data, and automatically processes two scenarios, namely the full standby state (promote the highest priority to the primary) and the multi-primary conflict (retain the highest priority and demote redundant instances) using the priority policy, ensuring that each microservice always maintains the optimal primary / standby distribution. By opening the real-time status write-back and query interfaces, it not only guarantees the state consistency within the cluster but also provides the external system with the ability to reliably perceive the service status, effectively enhancing the self-healing ability and service availability of the system.

[0024] Preferably, step 6 includes: The prediction module obtains the ARIMA model parameters and prediction configuration parameters; Collect the CPU occupancy information of each microservice from the microservice status storage table at the intermediate frequency statistical period and store it in the time series database; Based on the CPU occupancy information of each microservice in the historical 24 hours in the time series database, use the sliding window statistics to count the monthly peak pattern of the power outage operation microservices split by the SCADA system; Input the historical peak time points and CPU occupancy information into the ARIMA model to predict the peak time of the power outage operation microservices within the next 24 hours; Thirty minutes before the predicted peak time arrives, switch the power outage operation microservice from the primary node to the standby node, and update the service status and timestamp in the status storage table.

[0025] Significantly enhance the system's ability to handle business peaks through intelligent predictive scheduling: The prediction module analyzes the historical 24-hour CPU data based on the ARIMA model, combines the sliding window statistics to accurately predict the future peak period of the power outage operation microservices, and actively triggers the primary / standby switch 30 minutes before the peak arrives. This mechanism effectively avoids the business delay caused by the traditional passive response mode through forward-looking resource allocation, ensures that the key microservices always have the optimal load-bearing capacity during the peak period, and at the same time guarantees the consistency of the cluster data through real-time status updates.

[0026] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: 1. Through the dynamic primary-backup switching mechanism and predictive scheduling, it effectively prevents the node overload problem caused by the centralized operation of multiple primary services during the business peak period; based on the real-time load monitoring and prediction model, it automatically balances the resource allocation of the primary and backup nodes, significantly improving the system resource utilization efficiency and operation stability, and overcoming the defect of uneven load between the primary and backup nodes in the traditional architecture; 2. Utilize time series analysis to predict the business peak value, and adjust the primary-backup distribution of microservices in advance to ensure that the system can still maintain an even load during the peak period. This mechanism greatly improves the response speed of critical services (such as power outage operations), and at the same time enhances the impact resistance of the system in high-concurrency scenarios; 3. Combine heartbeat detection, CPU monitoring and status marking to achieve rapid perception and automatic recovery of microservice failures. The system can complete the isolation of abnormal instances and the takeover of standby instances in a very short time, ensuring the continuous and stable operation of the service, and significantly reducing the impact of failures on the business; 4. Adopt lightweight communication and efficient data synchronization strategies, enabling the system to dynamically adjust the microservice deployment scale while maintaining low-latency inter-node collaboration. This design provides scalable technical support for the microservice evolution of the SCADA system, meeting the future business growth requirements. Brief Description of the Drawings

[0027] Figure 1 It is a schematic diagram of the system structure of the present invention; Figure 2 It is a schematic diagram of splitting a single SCADA service process of the present invention into multiple microservices; Figure 3 It is the storage content and structure of each microservice in the data table of the present invention. Detailed Embodiments

[0028] The technical solution of the present invention will be further described below with reference to the drawings.

[0029] As Figure 1As shown in the figure, the intelligent load balancing and scheduling system of the present invention includes a parameter configuration module, a service communication module, a data acquisition module, a storage pool, a fault diagnosis module, a load statistics calculation module, a primary and standby instance switching module, and a prediction module. The parameter configuration module sets the IP addresses and communication ports of each node of the service, the CPU statistics period of each microservice, the communication heartbeat interval between service nodes, the microservice failure determination rule, the CPU threshold switching rule, the microservice primary and standby instance switching rule, the selection of the prediction model library, the prediction statistics period, the prediction early switching time, etc.; the service communication module is mainly responsible for establishing a link with other nodes and transmitting the relevant data of each microservice of the current node, receiving the relevant data transmitted by other nodes, and storing them in the communication buffer; the data acquisition module periodically samples the CPU occupancy of each microservice of the current node respectively, reads the relevant data of other nodes from the communication buffer, and stores them in the storage pool according to a predefined method; the storage pool is responsible for data storage and provides a fast data retrieval function; the load statistics calculation module is responsible for retrieving data from the storage pool, statistically calculating the real-time load conditions of the current node and other nodes, comparing them with the predefined load balancing parameters, and sending service instructions to the primary and standby instance switching module; the prediction module regularly reads relevant data from the storage pool, calculates the prediction result according to the relevant model prediction parameters read from the parameter configuration module, and sends relevant service instructions to the primary and standby instance switching module.

[0030] The intelligent load balancing and scheduling method of the present invention includes the following steps: Step 1: Deploy microservices 1.1 As Figure 2 shown in the figure, split the SCADA system function service into a remote control execution microservice (operate_server), a digital quantity processing microservice (di_process), an analog quantity processing microservice (ai_process), a light sign analysis microservice (summary_server), a coloring calculation microservice (topo), a power outage operation microservice (stop_power), a power-on operation microservice (start_power), a formula calculation microservice (scada_calculate), a data acquisition microservice (data_acquire), a protocol conversion microservice (fes_pscada), a real-time library access microservice (rdb_server), a commercial library submission microservice (sql_commit), and an alarm information microservice (alarm_server), etc.

[0031] 1.2 Deploy the above microservices on nodes 1, 2... N (usually 2 nodes for project deployment, and the following will take 2 nodes as an example for illustration), and call the parameter acquisition interface GetStartOrder to obtain the microservice startup sequence from the parameter configuration module, and start the microservices according to predefined requirements.

[0032] 1.3 When all microservices start, they call the parameter acquisition interface GetRegsterChannel to obtain the subscription channel number, subscribe to the message channel, subscribe to external information, and wait for the received message instruction.

[0033] 1.4 Each microservice runs according to the normal logic and periodically calls the parameter acquisition interface IsSuppliService. Since it has just started, the result of this interface call is defaulted to false (this interface will be detailed when it is true later), so each microservice does not provide services externally, that is, it runs its standby instance.

[0034] Step 2: Store the real-time status information of microservices 2.1 The data acquisition module calls the parameter acquisition interface GetServiceList to obtain the name list of microservices started by the current node from the parameter configuration module, and calls GetCircle to obtain the statistical period of each microservice. Then, through this period, it calls the operating system API function to obtain the CPU occupancy ratio information of each microservice by polling the microservice names in the name list; 2.2 The data acquisition module stores the above data in the storage pool in the form of a data table. The data table is named Table-MicroServiceInfo. The stored data adopts a two-dimensional structure and is stored in a key-value mode. The node name + microservice name is used as the unique key, and the value includes a custom data type info. The info data adopts the json format and includes at least 3 fields: CPU occupancy ratio, service status, and timestamp (for specific information, see Figure 3 ) 2.3 The storage pool described in step 2.2 can be an in-memory database, a relational database, or a distributed database; 2.4 Among them, the microservice status has the following 3 types: primary, standby, and faulty; 2.5 For the CPU occupancy ratio information of each microservice obtained by calling the operating system API function in step 2.1, if it can be obtained normally, first query the current microservice in the data table in the storage pool with the node name + microservice as the keyword. If this entry is found, update the timestamp field of the microservice and refresh the time; if this entry is not found, it is considered that the current microservice starts for the first time. At this time, add a record, set the current microservice status to standby, and obtain the current time and write it into the timestamp field; 2.6 If the CPU occupancy ratio information of the microservice cannot be obtained, it means that the current microservice has not started or the microservice has been failing and retrying. Therefore, set the status of the current microservice to faulty and do not update the timestamp information.

[0035] Step 3: Send Information between Nodes 3.1 The service communication module calls the parameter acquisition interface GetNodeInfo to obtain the IP addresses of each service node to be communicated, the communication port list, and the communication heartbeat interval between service nodes from the parameter configuration module. According to the number of nodes to be communicated, the same number of communication threads are started (the current node is node1, so the target node is node2, so the current thread is Tread_NodeName2). The thread applies to the operating system for 2 communication buffer areas, namely the read buffer area and the write buffer area, establishes a TCP connection with the target node through the socket, and sends a test message to verify normal communication of the connection; 3.2 The service communication module obtains all the microservice names and CPU occupancy ratios of the current node from Table-MicroServiceInfo according to the time interval of the communication heartbeat interval between service nodes, and organizes them with the node name + microservice as the smallest unit. Each node is sent as an independent data packet (LocalPacket); 3.3 The service communication module stores the LocalPacket obtained in step 2.2 into the read buffer area of each thread. Each communication thread periodically polls the read buffer area. When a data packet is found in the read buffer area of this thread, it is sent to the target node through the socket. After receiving the reply packet from the target node, the data in the read buffer area is deleted; otherwise, a retransmission mechanism is started; 3.4 When the socket of the communication thread of this node obtains a data packet sent by the peer node, the data packet is immediately retrieved from the socket and stored in the write buffer area of this communication thread; 3.5 The service communication module periodically polls the write buffer area of each communication thread. If data is obtained, it calls the ParsePacket interface to parse the content of the data packet and convert it into a two-dimensional key-value structure (the key is the node name + microservice name, and the value is a custom data type info. The info data is in json format, including CPU occupancy ratio, service status, and timestamp) stored in Table-MicroServiceInfo, and first stores it in the local temporary TempMap memory; 3.6 The business communication module periodically obtains the keys of the temporary TempMap and queries the Table-MicroServiceInfo table using the keys. If the key is found, the CPU occupancy ratio, service status, and timestamp of the microservice are updated; if the entry is not found, it is considered that the target node is sending the microservice information to the current node for the first time. At this time, a new record is added, the key-value pair is written into the Table-MicroServiceInfo table, and the current key in TempMap is deleted. Then, the operation is repeated until TempMap is emptied.

[0036] Step 4: Redundant calculation 4.1 The load statistics calculation module periodically obtains the microservice information of all nodes from the Table-MicroServiceInfo and stores it in the memory CalculateMAP in a two-dimensional key-value mode ([key1, [key2, value2]]). (Where key1 is the node name, key2 is the microservice name, and value2 is a json string containing the CPU occupancy ratio [CPU], service status [state], and timestamp [time]). 4.2 The load statistics calculation module traverses CalculateMAP and obtains the service status (key1-key2-value2-state) in value2 according to each key1-key2. If all key1-key2-value2-states are standby, the highest-priority key1-key2-value2-state is set to primary according to the priority, and other states remain unchanged; if there is a primary state, the state remains unchanged. If there are multiple primary states, the key1-key2-value2-state with a lower priority needs to be set to standby according to the priority requirements, and the result is written back to CalculateMAP; 4.3 The load statistics calculation module rewrites the data in CalculateMAP back to the Table-MicroServiceInfo; 4.4 The load statistics calculation module periodically obtains the primary / standby status of the microservices on the current node and provides services externally through the API interface. The interface name is IsSuppliService (string Microservername), and the input parameter is the microservice name. When the target microservice is in the primary state, this interface returns true, and when the target microservice is in the standby state, this interface returns false. (That is, the source of the interface call mentioned in Step 1.4) Step 5: Load statistics and primary / standby switch 5.1 The load statistics calculation module calls the parameter acquisition interface GetCPUInfo to obtain the CPU statistics period, the CPU threshold switching rule, and the microservice primary / standby instance switching rule from the parameter configuration module, and stores them in the local memory blocks CPUMap, CPUSwitchMap, and MicroServiceSwitchMap respectively.

[0037] 5.2 The load statistics calculation module periodically obtains the microservice CPU occupancy of all nodes from Table-MicroServiceInfo according to the CPU statistics time in CPUMap, and statistically calculates the CPU occupancy of each node in terms of nodes. It also obtains the CPU threshold of each node from CPUSwitchMap, and determines whether load adjustment is required based on the CPU difference between nodes. 5.3 If it is determined that a certain node needs load adjustment (usually by lowering the one with high CPU), traverse the CPU occupancy of each service of the node to be adjusted, select the microservice to be adjusted through the least principle selection algorithm. At this time, change the status of the microservice from the previous primary to standby, refresh the timestamp to the current time, and rewrite the primary / standby status and timestamp of the microservice in Table-MicroServiceInfo by calling the database interface TableModifyByKey.

[0038] Step 6: Model prediction and service switching 6.1 The model prediction module calls the parameter acquisition interface GetPredicteInfo to obtain the prediction model library, the prediction statistics period, and the prediction early switching time from the parameter configuration module. Considering the strong periodic characteristics of power services. In this example, the ARIMA model is used as the prediction model; considering short-term fluctuations and medium- and long-term trend predictions, the prediction statistics period in this example is divided into 3 frequencies, namely high-frequency statistics HighFre (5 seconds), medium-frequency statistics MiddleFre (1 minute), and low statistics LowFre (15 minutes); this example mainly predicts the power outage scenario at night, so the prediction early switching time is 30 minutes.

[0039] 6.2 The model prediction module obtains all microservice information of the current node from Table-MicroServiceInfo every 5 seconds through the middle-frequency statistics MiddleFre, and stores it in the time series database. 6.3 The model prediction module obtains the CPU occupancy of each microservice in the past 24 hours from the time series database, and calls the statistics module to calculate the time when the stop_power microservice appears at the CPU peak. Then, in the way of a sliding time window, count the time points when the stop_power microservice appears at the CPU peak within a month on a daily basis. 6.4 Input the calculated time points and the CPU occupancy ratio data of each microservice in the past 24 hours into the ARIMA model to predict the peak time points of the stop_power microservice in the next 24 hours and output the predicted time points. 6.5 According to the predicted time points and the predicted early switching time, which is 30 minutes in this example, perform business switching 30 minutes in advance, change the status of the microservice stop_power from the previous primary to standby, refresh the timestamp to the current time, and rewrite the primary / standby status and timestamp of the stop_power microservice in Table-MicroServiceInfo by calling the database interface TableModifyByKey.

Claims

1. A redundant microservice-based intelligent load balancing scheduling system, characterized in that, Including: A parameter configuration module, used to set the business node IP address, communication port, microservice CPU statistics period, heartbeat interval, failure determination rule, CPU threshold switching rule, primary / backup switching rule, prediction model library selection, prediction statistics period, and prediction early switching time; A business communication module, used to establish a communication link with other nodes, transmit the microservice data of the current node, and receive the data of the peer node, and store it in the communication buffer; A data acquisition module, used to periodically acquire the CPU occupancy ratio of each microservice of the current node, read the data of the peer node from the communication buffer, and store it in the storage pool; A fault diagnosis module, used to detect the failure status of the microservice. When the data acquisition module fails to obtain the CPU occupancy ratio of the microservice, mark the microservice as a fault state; A storage pool, used to store the real-time status data of the microservice and provide a fast retrieval function; A load statistics calculation module, used to calculate the real-time load of the node, compare the load balancing parameters, and generate a primary / backup switching instruction; A prediction module, used to regularly read the real-time information of the microservice from the storage pool and store it in the time series database to form historical load data, calculate according to the prediction model parameters set by the parameter configuration module, generate a future business load prediction result, and send an early adjustment instruction to the primary / backup instance switching module based on the prediction result; A primary / backup instance switching module, used to perform dynamic switching of the primary / backup status of the microservice.

2. The intelligent load balancing scheduling system according to claim 1, wherein The business communication module realizes data synchronization between nodes through multi-threaded Socket communication, including: Allocating independent threads and read / write buffers for each target node; Organizing the microservice data packet LocalPacket of the current node according to the heartbeat interval and sending it; Receiving the data packet of the peer node and parsing it into structured data, and updating it to the storage pool.

3. The intelligent load balancing scheduling system according to claim 1, wherein The data acquisition module polls the microservice name list by calling the operating system API function to obtain the CPU occupancy ratio information of each microservice, and uses the node name + microservice name as the unique key, and stores the JSON format data including the CPU occupancy ratio, service status, and timestamp as the value in the two-dimensional structure data table of the storage pool.

4. The intelligent load balancing scheduling system according to claim 1, wherein The load statistics calculation module traverses the microservice status data in the storage pool and performs the following operations: If all microservices are in the standby state, set the highest priority microservice as the primary according to the priority; If there are multiple primary conflicts, downgrade the redundant instance to the standby state according to the priority.

5. The intelligent load balancing scheduling system according to claim 1, characterized in that The prediction module uses the ARIMA model to perform time series analysis on the historical load data, predicts the time point of the future business peak period, and triggers the primary / backup instance switching module to adjust the microservice status based on the prediction early switching time.

6. A redundant microservice-based intelligent load balancing and scheduling method, characterized in that, Including the following steps: (1) The parameter configuration module obtains the microservice startup order, and deploys multiple independent microservices split by the SCADA system on at least two business nodes; After each microservice starts, it subscribes to the message channel and runs as a standby instance by default; (2) The data acquisition module obtains the microservice name list, polls the operating system API according to the preset statistics period to acquire the CPU occupancy ratio information of each microservice; Write the acquired CPU occupancy ratio information, the corresponding microservice status flag, and the acquisition timestamp into the storage pool; If the acquisition fails, mark it as a fault state; (3) The business communication module obtains the node IP and communication port, and establishes a multi-threaded Socket connection; it obtains the CPU occupancy ratio information and status flags of each microservice from the storage pool at the heartbeat interval, encapsulates them into a structured data packet, and sends them to the peer node through the read-write buffer mechanism. After receiving the peer data, it parses it into the same structure and updates the local storage pool. (4) The load statistics calculation module periodically traverses the CPU occupancy ratio information and status flags of each microservice in the storage pool: dynamically adjusts the primary and standby instance distribution according to the preset policy, and writes the status adjustment result back to the storage pool. At the same time, it provides the primary and standby status query service of each microservice externally through the status query interface. (5) The load statistics calculation module obtains the load threshold rule and calculates the real-time CPU load of each node; if the total CPU load of the primary node exceeds the threshold and the difference from the standby node exceeds the preset range, it selects the microservice with the highest CPU occupancy ratio in this node to be downgraded to the standby state, and promotes the corresponding microservice from the standby node to the primary state, and updates the status flag of the corresponding microservice in the storage pool. (6) The prediction module obtains the ARIMA model parameters, based on the historical microservice CPU occupancy ratio information in the time series database, uses the sliding window to statistically analyze the monthly peak pattern, predicts the business peak period of the high-load microservice, and triggers the adaptive switching of the primary and standby instances in advance.

7. The intelligent load balancing scheduling method according to claim 6, wherein Step 2 includes: The data collection module calls the GetServiceList interface to obtain the microservice name list from the parameter configuration module, and obtains the statistical period of each microservice through the GetCircle interface. Poll the operating system API according to the statistical period to obtain the CPU occupancy ratio information of each microservice. Store the obtained data with the node name + microservice name as the unique key and the JSON format data containing the CPU occupancy ratio, service status, and timestamp as the value into a two-dimensional structured data table; where: When the CPU occupancy ratio information is successfully obtained, if the corresponding key-value exists in the storage pool, update the timestamp field, otherwise add a new record and initialize the service status to standby. When the CPU occupancy ratio information is obtained fails, mark the status of the corresponding microservice as faulty and do not update the timestamp field. The storage pool is a memory database, a relational database, or a distributed database. The service status includes three states: primary, standby, and faulty.

8. The intelligent load balancing scheduling method according to claim 6, characterized in that Step 3 includes: The business communication module calls the GetNodeInfo interface to obtain the IP address, communication port, and heartbeat interval parameters of each node. Create the corresponding number of communication threads according to the number of nodes. Each thread is assigned an independent read buffer and write buffer, and a TCP connection is established through Socket. Periodically obtain the microservice operation data of this node from the microservice status storage table according to the heartbeat interval, and organize it into a LocalPacket data packet with the node name + microservice name as the smallest unit. Deposit the LocalPacket into the read buffer, and the communication thread sends it to the target node through Socket. After receiving the confirmation reply, clear the buffer data, otherwise trigger the retransmission mechanism. Receive the data packet sent by the peer node and deposit it into the write buffer. Call the ParsePacket interface to parse the content of the data packet and convert it into a two-dimensional key-value structure identical to the microservice status storage table, where the key is the node name + microservice name, and the value is JSON data containing the CPU occupancy ratio, service status, and timestamp; Temporarily store the parsed data in a temporary memory mapping area first, and then batch-update it to the microservice status storage table: if the corresponding key exists in the table, update the fields, otherwise insert a new record; after the update is completed, clear the corresponding data in the temporary memory mapping area.

9. The intelligent load balancing scheduling method according to claim 6, wherein Step 4 includes: The load statistics calculation module periodically obtains the microservice operation data of each node from the microservice status storage table and stores it in memory in a two-dimensional structure of [node identifier, [microservice identifier, operation data]]; Traverse the memory data structure: if all microservice instances are in the standby state, set the microservice instance with the highest priority to the primary state according to the preset priority; if multiple primary instances of the same microservice are detected, retain the instance with the highest priority according to the priority, and the rest are downgraded to the standby state; Write the adjusted status data back to the microservice status storage table, and provide a real-time primary / standby status query service externally through the status query interface, which returns its current primary / standby status according to the input microservice identifier.

10. The intelligent load balancing scheduling method according to claim 6, wherein, Step 6 includes: The prediction module obtains the ARIMA model parameters and prediction configuration parameters; Collect the CPU occupancy ratio information of each microservice from the microservice status storage table at the medium-frequency statistical period and store it in the time series database; Based on the CPU occupancy ratio information of each microservice in the time series database for the past 24 hours, adopt a sliding window statistic to obtain the monthly peak pattern of the power outage operation microservices split by the SCADA system; Input the historical peak time points and CPU occupancy ratio information into the ARIMA model to predict the peak time of the power outage operation microservices within the next 24 hours; 30 minutes before the predicted peak time arrives, switch the power outage operation microservice from the primary node to the standby node, and update the service status and timestamp in the status storage table.

Citation Information

Patent Citations

  • Self-adaptive micro-service resource management method and device

    CN120179395A

  • Automatic load balancing system for cloud servers

    DE202024105913U1

  • Cloud safety computing method, device and storage medium based on cloud fault-tolerant technology

    US20230350709A1