Secondary cluster scheduling system and method

By deploying the first and second scheduling management services in the dual-control equipment, monitoring and switching the main and backup control boards, the business interruption problem caused by the overall failure of the dual-control equipment is solved, and higher stability and reliability are achieved.

CN120276235APending Publication Date: 2025-07-08ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510356458.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing dual controller scheduling system cannot operate normally when the dual-control equipment fails overall, resulting in business interruption.

Method used

The secondary cluster scheduling system is adopted to ensure business continuity by deploying the first scheduling management service and the second scheduling management service in each first-level equipment.

Benefits of technology

It improves the stability and reliability of the scheduling system to ensure that the service can be executed normally when the main equipment is abnormal or the overall failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276235A_ABST
    Figure CN120276235A_ABST
Patent Text Reader

Abstract

The invention provides a secondary cluster scheduling system and method, which is used for improving the stability of the scheduling system, and comprises a primary working group and a standby working group, the main working group and the standby working group respectively comprise at least one primary device, and each primary device comprises a main control panel and a standby control panel; each control panel comprises a main service, a first scheduling management service and a second scheduling management service; the main service is used for executing a preset service; the first scheduling management service is used for monitoring the operation states of the two control boards in each primary device, and initiating switching of the main control board and the standby control board when the main control board of any primary device is abnormal and the standby control board is normal; and the second scheduling management service is used for monitoring the operation states of the two control boards in each primary device, and initiating the switching of the main device and the standby device when the main control board and the standby control board of any main device are abnormal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cluster scheduling, and in particular, to a two-level cluster scheduling system and method. Background Art

[0002] Dual-controller scheduling systems are applied to scenarios with high reliability requirements. Each independent single dual-control device provides services externally, and is suitable for scheduling when a single control board fails. That is, the standby control board monitors in real time whether the peer primary control board fails, and is ready to take over as the primary control board at any time, so as to ensure the continuous and stable operation of the service. However, if the dual-control device fails as a whole, such as a complete power outage, network disconnection, or both control boards are abnormal, the system will not be able to operate normally.

[0003] Therefore, how to improve the stability of the scheduling system is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This application provides a two-level cluster scheduling system and method for realizing the stability of the scheduling system.

[0005] In a first aspect, an embodiment of this application provides a two-level cluster scheduling system. The two-level cluster scheduling system includes a primary working group and a standby working group; wherein, the primary working group includes at least one primary device, and the standby working group includes at least one standby device; the primary device and the standby device are first-level devices, and each first-level device includes a primary control board and a standby control board; each primary control board and each standby control board include a primary service, a first scheduling management service, and a second scheduling management service; The primary service is used to: execute a preset service; The first scheduling management service is used to: monitor the operating states of the primary control board and the standby control board in each first-level device, and initiate a primary-standby control board switch when the primary control board in any first-level device is abnormal and the standby control board is normal; the primary-standby control board switch is used to indicate migrating the service of the primary control board of the any first-level device to the standby control board of the any first-level device; The second scheduling management service is used to: monitor the operating states of the primary control board and the standby control board in each first-level device, and initiate a primary-standby device switch when the primary control board and the standby control board of any primary device are both abnormal; the primary-standby device switch is used to indicate migrating the service of the any primary device to any normal standby device in the standby working group.

[0006] In this embodiment, a secondary cluster is composed of multiple primary devices, and each primary device includes an active control board, a standby control board, a primary service, a first scheduling management service, and a second scheduling management service. Both the first scheduling management service and the second scheduling management service are used to monitor the operating status of the active control board and the standby control board in each primary device. When the first scheduling management service detects that the active control board of any primary device is abnormal but the standby control board is normal, it can perform the active / standby control board switch, that is, migrate the services of the abnormal active control board to the standby control board of the any primary device to ensure that the any primary device can normally execute the preset services. When the second scheduling management service detects that both the active control board and the standby control board of any primary device are abnormal, it indicates that the any primary device as a whole is abnormal, and then performs the active / standby device switch, migrating the services of the any primary device to any normal standby device. In this way, regardless of whether the primary device has an abnormal active control board or is abnormal as a whole, the first scheduling management service and the second scheduling management service can ensure the normal execution of the preset services, thereby improving the stability and reliability of the scheduling system.

[0007] Optionally, before monitoring the operating status of the active control board and the standby control board in each primary device, the first scheduling management service is further used to: initialize each primary device, initiate the negotiation of the active control board according to the historical configuration record; the initialization includes resetting the two control boards in each primary device; the historical configuration record includes the active control board and the standby control board in each primary device configured before initialization; the negotiation of the active control board is used to determine the active control board among the two control boards of each initialized primary device; if the historical configuration record is one active control board and one standby control board, set the active control board in the historical configuration record as the active control board of the initialized primary device; if the historical configuration record is two active control boards, set the active control board in the latest configuration record in the historical configuration record as the active control board of the initialized primary device; if the historical configuration record is two standby control boards, set the control board located in the preset slot as the active control board of the initialized primary device.

[0008] Optionally, the first scheduling management service monitors the operating status of the active control board and the standby control board in each first-level device. When the active control board of any first-level device is abnormal and the standby control board is normal, it initiates the active-standby control board switchover, specifically used for: monitoring the change of the service configuration of the active control board, and updating the service configuration record of the active control board based on the change of the service configuration; monitoring the change of the service working status of the active control board. If it is detected that the service working status of the active control board is abnormal and the operating status of the standby control board is normal, then set the first virtual IP of the standby control board as the first virtual IP of the active control board, configure the standby control board based on the service configuration record of the active control board, and determine that the standby control board is the current active control board of any first-level device.

[0009] Optionally, the first scheduling management service is also used for: reporting the information of the active-standby control board switchover to the second scheduling management service; the second scheduling management service monitors the operating status of the active control board and the standby control board in each first-level device. When the active control board and the standby control board of any active device are both abnormal, it initiates the active-standby device switchover, specifically used for: monitoring the operating status of the active control board and the standby control board in each first-level device; if it is detected that the active control board and the standby control board of any active device are both abnormal, or, it is detected that the active control board of any active device is abnormal and the information of the active-standby control board switchover reported by the first scheduling management service of any active device is not received within the preset time threshold, then migrate the service of any active device to any normal standby device in the standby working group, and set the second virtual IP of any normal standby device as the second virtual IP of any active device, and confirm that any normal standby device is an active device of the current second-level cluster.

[0010] Optionally, after migrating the service of any active device to any normal standby device in the standby working group, the second scheduling management service is also used for: confirming that the active control board of any active device has returned to normal, migrating the service of any normal standby device back to any active device, and setting the second virtual IP of any normal standby device to be empty.

[0011] Optionally, at least one standby device in the standby working group includes a first standby device, and the main control board of the first standby device includes a scheduling management control service working machine; the first standby device is configured to: receive a cluster formation request from a client, where the cluster formation request includes a primary working group and a standby working group specified by the client; obtain the service configurations of all first-level devices included in the primary working group and the standby working group specified by the client according to the cluster formation request; deploy a first scheduling management service and a second scheduling management service in the service configuration of each first-level device among all first-level devices through the scheduling management control service working machine to obtain the updated service configuration of each first-level device; form a secondary cluster based on the updated service configuration of each first-level device; and manage the first scheduling management service and the second scheduling management service on each first-level device through the scheduling management control service working machine.

[0012] Optionally, the standby working group further includes a second standby device, and the second standby device is at least one standby device in the standby working group other than the first standby device; the first standby device is further configured to: create a standby machine for the scheduling management control service based on the scheduling management control service working machine, synchronize the service deployment of the scheduling management control service working machine to the standby machine for the scheduling management control service, and deploy the standby machine for the scheduling management control service to the second standby device in the standby working group; monitor the running status of the standby machine for the scheduling management control service of the second standby device through the scheduling management control service working machine; when it is detected that the standby machine for the scheduling management control service of any second standby device is abnormal and the number of first-level devices whose port services are connected to the standby machine for the scheduling management service is greater than a first threshold, remove the standby machine for the scheduling management control service on the any second standby device; if the second standby device in the standby working group only includes the any second standby device, determine whether there is a third standby device in the standby working group, where the third standby device is a first-level device other than the any second standby device and the first standby device, and if so, deploy the standby machine for the scheduling management control service to the third standby device, and if not, trigger an alarm.

[0013] Optionally, the second standby device is further configured to: monitor the running status of the scheduling management control service working machine of the first standby device through the standby machine for the scheduling management control service; when it is detected that the scheduling management control service working machine of the first standby device is abnormal and the number of first-level devices whose port services are connected to the scheduling management service working machine is greater than a second threshold, manage the first scheduling management service and the second scheduling management service on each first-level device based on the standby machine for the scheduling management control service.

[0014] Second aspect, an embodiment of the present application provides a two - level cluster scheduling method, which is applied to a two - level cluster scheduling system. The two - level cluster scheduling system includes a primary working group and a standby working group; wherein, the primary working group includes at least one primary device, and the standby working group includes at least one standby device; the primary device and the standby device are first - level devices, and each first - level device includes a primary control board and a standby control board; each primary control board and each standby control board include a primary service, a first scheduling management service, and a second scheduling management service; the method includes: Monitor the operating status of the primary control board and the standby control board in each first - level device. When the primary control board in any first - level device is abnormal and the standby control board is normal, initiate the primary - standby control board switch; the primary - standby control board switch is used to indicate migrating the services of the primary control board of the any first - level device to the standby control board of the any first - level device; Monitor the operating status of the primary control board and the standby control board in each first - level device. When the primary control board and the standby control board in any primary device are both abnormal, initiate the primary - standby device switch; the primary - standby device switch is used to indicate migrating the services of the any primary device to any normal standby device in the standby working group.

[0015] Optionally, before monitoring the operating status of the primary control board and the standby control board in each first - level device, the method further includes: initializing each first - level device, initiating the primary control board negotiation according to the historical configuration record; the initialization includes resetting the two control boards in each first - level device; the historical configuration record includes the primary control board and the standby control board in the first - level device configured before initialization; the primary control board negotiation is used to indicate determining the primary control board among the two control boards of each first - level device after initialization; if the historical configuration record is one primary control board and one standby control board, set the primary control board in the historical configuration record as the primary control board of the first - level device after initialization; if the historical configuration record is two primary control boards, set the primary control board in the latest configuration record in the historical configuration record as the primary control board of the first - level device after initialization; if the historical configuration record is two standby control boards, set the control board located in the preset slot as the primary control board of the first - level device after initialization.

[0016] Optionally, monitor the operating status of the primary control board and the standby control board in each first-level device. When the primary control board of any first-level device is abnormal and the standby control board is normal, initiate the primary / standby control board switchover, including: monitoring the change in the service configuration of the primary control board, and updating the service configuration record of the primary control board based on the change in the service configuration; monitoring the change in the service working status of the primary control board. If it is detected that the service working status of the primary control board is abnormal and the operating status of the standby control board is normal, set the first virtual IP of the standby control board as the first virtual IP of the primary control board, configure the standby control board based on the service configuration record of the primary control board, and determine that the standby control board is the current primary control board of any first-level device.

[0017] Optionally, report the information of the primary / standby control board switchover to the second scheduling management service through the first scheduling management service; monitor the operating status of the primary control board and the standby control board in each first-level device. When both the primary control board and the standby control board of any primary device are abnormal, initiate the primary / standby device switchover, including: monitoring the operating status of the primary control board and the standby control board in each first-level device; if it is detected that both the primary control board and the standby control board of any primary device are abnormal, or, if it is detected that the primary control board of any primary device is abnormal and the information of the primary / standby control board switchover reported by the first scheduling management service of any primary device is not received within the preset time threshold, then migrate the service of any primary device to any normal standby device in the standby workgroup, and set the second virtual IP of any normal standby device as the second virtual IP of any primary device, and confirm that any normal standby device is a primary device of the current secondary cluster.

[0018] Optionally, after migrating the service of any primary device to any normal standby device in the standby workgroup, the method further includes: confirming that the primary control board of any primary device has returned to normal, migrating the service of any normal standby device back to any primary device, and setting the second virtual IP of any normal standby device to be empty.

[0019] Optionally, at least one standby device in the standby working group includes a first standby device, and the main control board of the first standby device includes a scheduling management control service working machine; the method further includes: receiving a cluster formation request from a client, where the cluster formation request includes a main working group and a standby working group specified by the client; obtaining the service configurations of all first-level devices included in the main working group and the standby working group specified by the client according to the cluster formation request; deploying a first scheduling management service and a second scheduling management service in the service configuration of each first-level device among all first-level devices through the scheduling management control service working machine to obtain the updated service configuration of each first-level device; forming a second-level cluster based on the updated service configuration of each first-level device; and managing the first scheduling management service and the second scheduling management service on each first-level device through the scheduling management control service working machine.

[0020] Optionally, the standby working group further includes a second standby device, and the second standby device is at least one standby device other than the first standby device in the standby working group; the method further includes: creating a standby machine for the scheduling management control service based on the scheduling management control service working machine, synchronizing the service deployment of the scheduling management control service working machine to the standby machine for the scheduling management control service, and deploying the standby machine for the scheduling management control service to the second standby device in the standby working group; monitoring the running state of the standby machine for the scheduling management control service of the second standby device through the scheduling management control service working machine; when it is detected that the standby machine for the scheduling management control service of any second standby device is abnormal and the number of first-level devices whose port services are connected to the standby machine for the scheduling management service is greater than a first threshold, removing the standby machine for the scheduling management control service on the any second standby device; if the second standby device in the standby working group only includes the any second standby device, determining whether there is a third standby device in the standby working group, where the third standby device is a first-level device other than the any second standby device and the first standby device, if it exists, deploying the standby machine for the scheduling management control service to the third standby device, and if it does not exist, triggering an alarm.

[0021] Optionally, the method further includes: monitoring the running state of the scheduling management control service working machine of the first standby device through the standby machine for the scheduling management control service; when it is detected that the scheduling management control service working machine of the first standby device is abnormal and the number of first-level devices whose port services are connected to the scheduling management service working machine is greater than a second threshold, managing the first scheduling management service and the second scheduling management service on each first-level device based on the standby machine for the scheduling management control service.

[0022] In a third aspect, an embodiment of the present application provides an electronic device, including at least one processor, and when the at least one processor executes a computer program stored in a memory, the method in the first aspect or any optional implementation manner of the first aspect is implemented.

[0023] Fourthly, an embodiment of the present application provides a computer-readable storage medium for storing instructions that, when executed, implement the method in the first aspect or any optional implementation manner of the first aspect.

[0024] Fifthly, an embodiment of the present application provides a computer program product including computer program code that, when running on a computer, implements the method in the first aspect or any optional implementation manner of the first aspect.

[0025] For the technical effects or advantages of one or more technical solutions provided in the second, third, fourth, and fifth aspects in the embodiments of the present application, they can all be correspondingly explained by the technical effects or advantages of the corresponding one or more technical solutions provided in the first aspect. Description of the Drawings

[0026] Figure 1 It is an example diagram of a dual-control device provided by an embodiment of the present application; Figure 2 It is an example diagram of a secondary cluster provided by an embodiment of the present application; Figure 3 It is a structural diagram of a secondary cluster scheduling system provided by an embodiment of the present application; Figure 4 It is an example diagram of the switching of the main and standby control boards provided by an embodiment of the present application; Figure 5 It is an example diagram of the switching of the main and standby devices provided by an embodiment of the present application; Figure 6 It is an example diagram of the switching of the main and standby control boards of a first standby device provided by an embodiment of the present application; Figure 7 It is an example diagram of the switching of a DCS working machine and a DCS standby machine provided by an embodiment of the present application; Figure 8 It is a structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0027] In the technical solution of the present application, the acquisition, dissemination, use, etc. of data all comply with the requirements of relevant national laws and regulations.

[0028] It should be noted that in the embodiments of the present application, some existing industry solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.

[0029] The technical solution of the present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0030] It should be understood that in the description of the embodiments of the present application, "a plurality of" means two or more. The "first", "second", etc. in the embodiments of the present application are used to distinguish different objects, rather than to describe a specific order. The term "and / or" in the embodiments of the present application is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. A module in the embodiments of the present application refers to a part with independent functions in a software system.

[0031] To facilitate the understanding of the embodiments of the present application, the following is an introduction to some professional terms involved in the embodiments of the present application: 1. Dual-control device (dual-controller system): The primary device in the embodiments of the present application is a dual-control device; taking the dual-control device in the field of video storage servers as an example of the dual-control device, as Figure 1 shown, the embodiments of the present application provide an example diagram of a dual-control device.

[0032] The dual-control device is applied to occasions with high reliability requirements. The standby controller monitors in real time whether the peer primary controller fails, so as to take over as the primary controller at any time, thereby ensuring the continuous and stable operation of the service. The two controllers share a batch of disks, but only the primary controller can recognize the hard disks. The dual-control device is presented to the outside in the form of one device and is a 1+1 dual-machine hot standby cluster. Among them, the control board is the board that executes functions on the controller.

[0033] 2. N+M primary and standby cluster (i.e., secondary cluster): Taking the dual-control device in the field of video storage servers as an example of the dual-control device, as Figure 2 shown, the embodiments of the present application provide an example diagram of a secondary cluster.

[0034] In an N+M primary / standby cluster, the N+M service nodes (i.e., first-level devices) of the system jointly form a cluster to provide services externally. Among them, N represents the N primary devices in the cluster workgroup, and M represents the M standby devices in the cluster workgroup. When any number of (not more than M) primary devices in workgroup N fail, the standby devices in standby group M can seamlessly replace the faulty workgroup devices, take over the work of the primary devices, and ensure that the external services continue during the abnormal period.

[0035] 3. The Dispatching Console Service (DCS) is the center of the entire cluster control and is responsible for the management of the entire cluster. The roles played by DCS in the cluster mainly include: initiating negotiation and forming a cluster; monitoring the status of each service node and centrally managing the status; monitoring for abnormalities, initiating primary / standby switching, and migrating services; repairing the primary device and migrating the service cluster back, etc.

[0036] Currently, the dual-control device is a 1+1 first-level cluster that can only provide services externally through a single device, that is, it can only be applied to the abnormal handling scenario when a single control board fails. If there is an overall power outage, network disconnection, or both control boards are abnormal, the entire cluster will not be able to provide services externally.

[0037] In view of this, the embodiments of the present application are provided. Multiple first-level devices (i.e., single dual-control devices) deploy a first dispatching management service and a second dispatching management service in each first-level device to monitor the primary control board and the standby control board in each first-level device. When the first dispatching management service detects that the primary control board of any first-level device is abnormal but the standby control board is normal, it can perform the primary / standby control board switching, that is, migrate the services of the abnormal primary control board to the standby control board of the any first-level device to ensure that the any first-level device can normally execute the preset services. When the second dispatching management service detects that both the primary control board and the standby control board of any primary device are abnormal, it indicates that the any primary device as a whole is abnormal, and then performs the primary / standby device switching, migrating the services of the any primary device to any normal standby device. In this way, regardless of whether the primary device has an abnormal primary control board or is abnormal as a whole, the first dispatching management service and the second dispatching management service can ensure the normal execution of the preset services, thereby improving the stability and reliability of the dispatching system.

[0038] It can be understood that the secondary cluster dispatching system provided by the embodiments of the present application can be applied to any cluster dispatching scenario, including but not limited to the above video storage server field.

[0039] See Figure 3, which is a structural diagram of a secondary cluster scheduling system provided by an embodiment of the present application. The system includes a primary working group 01 and a standby working group 02; among them, the primary working group 01 includes at least one primary device 03, and the standby working group 02 includes at least one standby device 04; the primary device 03 and the standby device 04 are first-level devices; each first-level device includes a primary control board 05 and a standby control board 06; each primary control board 05 and each standby control board 06 include a primary service 07, a first scheduling management service 08, and a second scheduling management service 09; Among them, the primary service 07 on each first device is used to: execute a preset service; The first scheduling management service 08 is used to: monitor the operating status of the primary control board 05 and the standby control board 06 in each first-level device, and initiate a primary-standby control board switch when the primary control board 05 of any first-level device is abnormal and the standby control board 06 is normal; the primary-standby control board switch is used to indicate migrating the service of the primary control board 05 of the any first-level device to the standby control board 06 of the any first-level device; The second scheduling management service 09 is used to: monitor the operating status of the primary control board 05 and the standby control board 06 in each first-level device, and initiate a primary-standby device switch when the primary control board 05 and the standby control board 06 of any primary device 03 are both abnormal; the primary-standby device switch is used to indicate migrating the service of the any primary device 03 to any normal standby device 04 in the standby working group 02.

[0040] It can be understood that Figure 3 The numbers of the primary device 03 and the standby device 04 in [[ ]] are only a possible example. In actual applications, the number of first-level devices in the primary working group 01 and / or the standby working group 02 can be increased according to requirements, and the embodiments of the present application do not limit this. In order to better present the overall service of the secondary cluster, in the following embodiments, the secondary cluster takes the primary working group 01 including two primary devices 03 and the standby working group 02 including two standby devices 04 as an example. In actual applications, the numbers of the primary device 03 and the standby device 04 can be set according to requirements.

[0041] In addition, although the labels and functions of the two control boards deployed on different first-level devices, as well as the primary service 07, the first scheduling management service 08, and the second scheduling management service 09 deployed on the two control boards, are the same, there may still be differences in the two control boards deployed on different first-level devices and the primary service 07, the first scheduling management service 08, and the second scheduling management service 09 deployed on the two control boards in actual deployment, and it can also be understood as being different.

[0042] In a possible embodiment, the first scheduling management service 08 monitors the operating status of the main control board 05 and the standby control board 06 in each first-level device. When the main control board 05 of any first-level device is abnormal and the standby control board 06 is normal, the specific implementation method of initiating the switching of the main and standby control boards is as follows: Monitoring the service configuration changes of the primary control board 05, and updating the service configuration record of the primary control board 05 based on the service configuration changes; Monitor changes in the business working status of the main control board 05. If it is monitored that the business working status of the main control board 05 is abnormal and the operating status of the standby control board 06 is normal, set the first virtual IP of the standby control board 06 as the first virtual IP of the main control board 05, and configure the standby control board 06 based on the business configuration record of the main control board 05, and determine that the standby control board 06 is the main control board 05 of any of the current first-level devices.

[0043] The first virtual IP is the IP that the main control board 05 provides services to the outside.

[0044] For example, Figure 4 The following is an example of the switchover between the main and standby control boards: When the first scheduling management service 08 on the main control board 05 of the main device 03 on the left side of the main workgroup 01 detects that there is an abnormality in the main service 07 of the main control board 05, and the first scheduling management service 08 of the main control board 05 confirms that the main service 07 of the standby control board 06 is normal through information exchange with the first scheduling service 08 of the standby control board 06, the first scheduling service management 08 of the main control board 05 actively initiates the main-standby control board switching, migrates the business on the main service 07 of the main control board 05 to the main service 07 of the standby control board 06, and sets the first virtual IP of the standby control board 06 to the first virtual IP of the main control board 05, so that the standby control board 06 can take over the abnormal main control board 05 to provide services to the outside, thereby ensuring that the external business is not interrupted during the abnormal period of the main control board 05.

[0045] It can be understood that since the active control board 05 and the standby control board 06 located in the same primary device share the same storage resources, if the original active control board 05 is restored to normal, the service migration operation may not be performed.

[0046] In a possible embodiment, the first scheduling management service 08 is further used to: report the information of the switching of the main and standby control boards to the second scheduling management service 09; The second scheduling management service 09 monitors the operating status of the active control board 05 and the standby control board 06 in each primary device. When both the active control board 05 and the standby control board 06 of any primary device 03 are abnormal, it initiates the primary-standby device switchover, specifically for: Monitoring the operating status of the active control board 05 and the standby control board 06 in each primary device; If it is monitored that both the active control board 05 and the standby control board 06 of any primary device 03 are abnormal, or, it is monitored that the active control board 05 of any primary device 03 is abnormal and the information on the primary-standby control board switchover reported by the first scheduling management service 08 of this any primary device 03 has not been received within the preset time threshold, then the service of this any primary device 03 is migrated to any normal standby device 04 in the standby working group 02, and the second virtual IP of this any normal standby device 04 is set to the second virtual IP of this any primary device 03, and it is confirmed that this any normal standby device 04 is a primary device 03 of the current secondary cluster.

[0047] Among them, the second virtual IP is the IP through which the primary device 03 provides services externally.

[0048] Exemplarily, taking Figure 5 an example of a primary-standby device switchover example diagram shown as follows: The second scheduling management service 09 on the active control board 05 of the left primary device 03 in the primary working group 01 monitors that the primary service 07 of the active control board 05 is abnormal and, through the second scheduling management service 09 on the standby control board 06, monitors that the primary service 07 of the standby control board 06 is also abnormal, that is, it is determined that both the active control board 05 and the standby control board 06 of this primary device 03 are abnormal, and this primary device 03 can no longer provide services externally; Or, the second scheduling management service 09 on the active control board 05 of this primary device 03 monitors that the primary service 07 of the active control board 05 is abnormal and the information on the primary-standby control board switchover reported by the first scheduling management service 08 of the active control board 05 has not been received within the preset duration threshold, that is, it is determined that the primary service 07 of the active control board 05 is abnormal but the problem that the active control board 05 cannot provide services externally cannot be solved through the primary-standby control board switchover; Therefore, the second scheduling management service 09 of the active control board 05 actively initiates a primary-standby device switchover to the left standby device 04 in the standby working group 02, migrates the service of the primary service 07 of the active control board 05 of this primary device 03 to the active control board 05 of this standby device 04, and sets the second virtual IP of this standby device 04 to the second virtual IP of this primary device 03 to achieve normal external services of the secondary cluster.

[0049] It can be understood that Figure 5Taking the migration of the services of the abnormal primary device 03 to the standby device 04 on the left side in the standby working group 02 as an example, in practice, the services of the abnormal primary device 03 can be migrated to any normal standby device 04 in the standby working group 02, and the embodiments of the present application do not limit this.

[0050] Optionally, since the storage resources of the primary device 03 and the standby device 04 are different, after the second scheduling management service 09 of the primary control board 05 of the primary device 03 confirms that the primary service 07 on the primary control board 05 has returned to normal, it migrates the services migrated to the standby device 04 back to the primary device 03, and sets the second virtual IP of the standby device 04 to null or other default values.

[0051] It can be understood that in the above embodiments, determining that the primary control board 05 or the primary device 03 is abnormal by monitoring the abnormality of the primary control board 05 or the primary service 07 is only an optional method for confirming abnormalities provided by the embodiments of the present application. In practical applications, abnormalities also include situations such as control board short circuits and power outages, and the embodiments of the present application do not limit the manner of confirming abnormalities.

[0052] In a possible embodiment, before the first scheduling management service 08 monitors the operating states of the primary control board 05 and the standby control board 06 in each first-level device, it is also necessary to configure the primary control board 05 and the standby control board 06 in each first-level device. The specific configuration method is as follows: Initialize each first-level device, initiate primary control board negotiation according to the historical configuration record; the initialization includes resetting the two control boards in each first-level device; the historical configuration record includes the primary control board 05 and the standby control board 06 in the first-level device configured before initialization; the primary control board negotiation is used to indicate determining the primary control board 05 among the two control boards of each initialized first-level device; If the historical configuration record is one primary control board 05 and one standby control board 06, set the primary control board 05 in the historical configuration record as the primary control board 05 of the initialized first-level device; If the historical configuration record is two primary control boards 05, set the primary control board 05 in the latest configuration record in the historical configuration record as the primary control board 05 of the initialized first-level device; If the historical configuration record is two standby control boards 06, set the control board located in the preset slot as the primary control board 05 of the initialized first-level device.

[0053] It can be understood that each configuration situation of the two control boards in each primary device will be stored in the local historical configuration record. After each power-off and restart of each primary device, an initialization will be performed, that is, the two control boards will be reset, and the historical configuration record will be read to determine the primary control board 05 in the two reset control boards, and the current configuration situation will be recorded in the historical configuration record.

[0054] In a possible design, in order to ensure the stable operation of each second scheduling management service 09 and the first scheduling management service 08, it is also necessary to manage each second scheduling management service 09 and the first scheduling management service 08. Therefore, in the embodiment of the present application, a scheduling management control service working machine 11, that is, a DCS working machine, is also deployed in the secondary cluster scheduling system to manage the second scheduling management service 09 and the first scheduling management service 08 in each primary device.

[0055] The scheduling management control service working machine 11 can directly initiate the primary and standby device switching when it monitors that the second scheduling management service 09 and the first scheduling management service 08 cannot operate normally, such as when the primary device is powered off as a whole. In this way, the stability of the scheduling system can be further improved.

[0056] In addition, the scheduling management control service working machine 11 can also directly monitor the primary service 07 of the primary controller 05 of each primary device to further ensure that abnormal situations of each primary device can be monitored in a timely manner.

[0057] Specifically, at least one standby device 04 in the standby workgroup 02 includes a first standby device 10, and the primary control board 05 of the first standby device 10 includes a scheduling management control service working machine 11.

[0058] Optionally, since the scheduling management control service working machine 11 is used to manage the second scheduling management service 09 and the first scheduling management service 08 in each primary device, it is necessary to ensure the normal operation of the scheduling management control service working machine 11, that is, it is necessary to ensure the normal operation status of the control board where the scheduling management control service working machine 11 is located and the primary device. If the primary control board 05 of the first standby device 10 has an abnormality, the primary and standby control board switching is also required. As Figure 6 shown, since the first standby device 10 is currently a standby device 04 and does not need to provide services externally, when migrating, it is not necessary to migrate the services on the primary control board 05 to the standby control board 06, and only the scheduling management control service working machine 11 (that is, the DCS working machine 11) on the primary control board 05 needs to be migrated to the standby control board 06 (that is, Figure 6 the state migration shown).

[0059] In a possible design, the present application also provides a method for constructing a secondary cluster scheduling system, and the specific implementation manner of this method is as follows: The first standby device 10 is used to: receive a cluster construction request from a client, where the cluster construction request includes a primary working group 01 and a standby working group 02 specified by the client; Obtain the service configurations of all first-level devices included in the primary working group 01 and the standby working group 02 specified by the client according to the cluster construction request; Deploy the first scheduling management service 08 and the second scheduling management service 09 in the service configuration of each first-level device among all first-level devices through the scheduling management control service working machine 11 to obtain the updated service configuration of each first-level device; Based on the updated service configuration of each first-level device, construct a secondary cluster; Manage the first scheduling management service 08 and the second scheduling management service 09 on each first-level device through the scheduling management control service working machine 11.

[0060] In this embodiment, in addition to deploying the scheduling management control service working machine 11, it is also necessary to create a scheduling management control service standby machine 12. The service configuration of the scheduling management control service standby machine 12 is the same as that of the scheduling management control service working machine 11 to ensure that when the scheduling management control service working machine 11 fails, the second scheduling management service 09 and the first scheduling management service 08 can also be managed through the scheduling management control service standby machine 12.

[0061] Specifically, the standby working group 02 further includes a second standby device 13, and the second standby device 13 is at least one standby device 04 other than the first standby device 10 in the standby working group 02; the first standby device 10 is further used to: Create a scheduling management control service standby machine 12 through the scheduling management control service working machine 11, synchronize the service configuration of the scheduling management control service working machine 11 to the scheduling management control service standby machine 12, and deploy the scheduling management control service standby machine 12 on the second standby device 13 in the standby working group; Monitor the running status of the scheduling management control service standby machine 12 of the second standby device 13 through the scheduling management control service working machine 11; when it is detected that the scheduling management control service standby machine 12 of any second standby device 13 is abnormal and the number of first-level devices connected to the port service of the scheduling management service standby machine 12 is greater than the first threshold, remove the scheduling management control service standby machine 12 on the any second standby device 13; If the second standby device 13 in the standby working group 02 only includes any one of the second standby devices 13, it is determined whether there is a third standby device in the standby working group 02. The third standby device is a primary device other than any one of the second standby devices 13 and the first standby device 10. If it exists, the scheduling management control service standby machine is deployed on the third standby device. If it does not exist, an alarm is triggered.

[0062] It can be understood that whether the port service of the scheduling management service standby machine 12 is connected can be determined by determining whether the port of the scheduling management service standby machine 12 can normally receive and send messages of the primary device.

[0063] In addition, since the second standby device 13 is the standby device 04 and does not need to provide services externally, its main function is to replace the function of the abnormal scheduling management control service working machine 11 through the scheduling management control service standby machine 12. Therefore, the second standby device 13 can also only include the scheduling management control service standby machine 12, that is, the scheduling management control service standby machine 12 itself is the second standby device 13 (it can be understood that the second standby device 13 is the standby device of the scheduling management control service working machine 11); of course, the second standby device 13 can also include the basic modules of the standby device 04: the primary control board 05, the standby control board 06, etc. That is, the second standby device 13 can be used as the standby device of the scheduling management control service working machine 11 and also as the standby device of the primary device 03; the embodiments of the present application do not limit this.

[0064] In this embodiment, the scheduling management control service working machine 11 can monitor the scheduling management control service standby machine 12 in real time. When it is determined that the scheduling management control service standby machine 12 is abnormal, the scheduling management control service standby machine 12 is reset to ensure that there is an available scheduling management control service standby machine 12 in the standby working group 02, further improving the stability of the scheduling system.

[0065] Optionally, the scheduling management control service standby machine 12 can also monitor the scheduling management control service working machine 11 in real time. If it is detected that the scheduling management control service working machine 11 is abnormal, but the number of primary devices connected to the port service of the scheduling management service working machine 11 is not greater than the first threshold, the scheduling management control service standby machine 12 considers itself abnormal and triggers an alarm to prompt the scheduling management control service working machine 11 of the first standby device 10 to replace the scheduling management control service standby machine 12 of the second standby device 13, or prompt the staff to perform abnormal handling.

[0066] It can be understood that the specific value of the first threshold can be set according to actual needs, and the embodiments of the present application do not limit this.

[0067] In a possible embodiment, the second standby device 13 is further configured to: schedule the management control service standby machine 12 to monitor the operating status of the scheduling management control service working machine 11 of the first standby device 10; when it is detected that the scheduling management control service working machine 11 of the first standby device 10 is abnormal and the number of primary devices connected to the port service of the scheduling management service working machine 11 is greater than a second threshold, manage the first scheduling management service 08 and the second scheduling management service 09 on each primary device based on the management control service standby machine 12 of the scheduling management control service.

[0068] Exemplarily, taking Figure 7 the example of the switching of the DCS working machine 11 (i.e., the scheduling management service working machine 11) and the DCS standby machine 12 (i.e., the management control service standby machine 12 of the scheduling management control service) shown, when the DCS standby machine 12 detects that the DCS working machine is abnormal and the number of primary devices connected to the port service of the DCS working machine is greater than a second threshold, the DCS standby machine 12 actively initiates negotiation to become the DCS working machine and performs the management work of the first scheduling management service 08 and the second scheduling management service 09 on each primary device.

[0069] In this embodiment, the operating status of the DCS working machine 11 is monitored by the DCS standby machine 12. When the DCS working machine 11 is abnormal, the DCS standby machine 12 can promptly become the new DCS working machine and replace the DCS working machine 11 to manage the first scheduling management service 08 and the second scheduling management service 09, so as to improve the stability of the scheduling system.

[0070] Optionally, if the scheduling management control service working machine 11 detects that the management control service standby machine 12 of any second standby device 13 is abnormal and the number of primary devices connected to the port service of the any second standby device 13 is not greater than a first threshold, the scheduling management control service working machine 11 confirms that it is abnormal and triggers an alarm to prompt the management control service standby machine 12 of the second standby device 13 to replace the scheduling management control service working machine 11 of the first standby device 10, or prompt the staff to perform abnormal handling.

[0071] It can be understood that the specific value of the second threshold can be set according to actual needs, and the embodiments of the present application do not limit this.

[0072] In this way, through the combination of the self-test of the scheduling management control service working machine 11 and the monitoring of the management control service standby machine 12 of the scheduling management control service, the abnormality of the scheduling management control service working machine 11 can be identified in a timely and accurate manner, and then the standby machine can take over the work of the working machine, improving the stability of the scheduling system.

[0073] In this embodiment, a secondary cluster is composed of multiple primary devices, and each primary device includes a primary control board 05, a standby control board 06, a primary service 07, a first scheduling management service 08, and a second scheduling management service 09. Both the first scheduling management service 08 and the second scheduling management service 09 are used to monitor the operating status of the primary control board 05 and the standby control board 06 in each primary device. When the first scheduling management service 08 detects that the primary control board 05 of any one primary device is abnormal but the standby control board 06 is normal, it can perform the primary-standby control board switchover, that is, migrate the services of the abnormal primary control board 05 to the standby control board 06 of the any one primary device to ensure that the any one primary device can normally execute the preset services. When the second scheduling management service 09 detects that both the primary control board 05 and the standby control board 06 of any one primary device 03 are abnormal, it indicates that the any one primary device 03 as a whole is abnormal, and then performs the primary-standby device switchover, migrating the services of the any one primary device 03 to any normal standby device 04. In this way, whether the primary device 03 has an abnormal primary control board 05 or is abnormal as a whole, the first scheduling management service 08 and the second scheduling management service 09 can ensure the normal execution of the preset services, thereby improving the stability and reliability of the scheduling system.

[0074] Based on the same technical concept, an embodiment of the present application also provides a secondary cluster scheduling method, which is applied to the above secondary cluster scheduling system. The method includes: Monitor the operating status of the primary control board 05 and the standby control board 06 in each primary device. When the primary control board 05 of any one primary device is abnormal and the standby control board is normal, initiate the primary-standby control board switchover. The primary-standby control board switchover is used to indicate migrating the services of the primary control board 05 of the any one primary device to the standby control board 06 of the any one primary device. Monitor the operating status of the primary control board 05 and the standby control board 06 in each primary device. When both the primary control board 05 and the standby control board 06 of any one primary device 03 are abnormal, initiate the primary-standby device switchover. The primary-standby device switchover is used to indicate migrating the services of the any one primary device 03 to any normal standby device 04 in the standby working group 02.

[0075] Optionally, before monitoring the operating status of the active control board 05 and the standby control board 06 in each first-level device, the method further includes: initializing each first-level device, initiating negotiation of the active control board according to the historical configuration record; the initialization includes resetting the two control boards in each first-level device; the historical configuration record includes the active control board 05 and the standby control board 06 in the first-level device configured before initialization; the negotiation of the active control board is used to indicate determining the active control board 05 among the two control boards of each initialized first-level device; if the historical configuration record is one active control board 05 and one standby control board 06, set the active control board 05 in the historical configuration record as the active control board 05 of the initialized first-level device; if the historical configuration record is two active control boards 05, set the active control board 05 in the latest configuration record in the historical configuration record as the active control board 05 of the initialized first-level device; if the historical configuration record is two standby control boards 06, set the control board in the preset slot as the active control board 05 of the initialized first-level device.

[0076] Optionally, monitor the operating status of the active control board 05 and the standby control board 06 in each first-level device. When the active control board 05 of any first-level device is abnormal and the standby control board 06 is normal, initiate the active / standby control board switch, including: monitoring the change of the service configuration of the active control board 05, and updating the service configuration record of the active control board 05 based on the change of the service configuration; monitoring the change of the service working status of the active control board 05. If it is monitored that the service working status of the active control board 05 is abnormal and the operating status of the standby control board 06 is normal, set the first virtual IP of the standby control board 06 as the first virtual IP of the active control board 05, and configure the standby control board 06 based on the service configuration record of the active control board 05, and determine that the standby control board 06 is the active control board 05 of the current any first-level device.

[0077] Optionally, the information about the primary and standby control board switching is reported to the second scheduling management service 09 by the first scheduling management service 08; the operating states of the primary control board 05 and the standby control board 06 in each first-level device are monitored, and when both the primary control board 05 and the standby control board 06 of any primary device 03 are abnormal, a primary and standby device switching is initiated, including: monitoring the operating states of the primary control board 05 and the standby control board 06 in each first-level device; if it is monitored that both the primary control board 05 and the standby control board 06 of any primary device 03 are abnormal, or, it is monitored that the primary control board 05 of any primary device 03 is abnormal and the information about the primary and standby control board switching reported by the first scheduling management service 08 of this any primary device 03 is not received within a preset time threshold, then the service of this any primary device 03 is migrated to any normal standby device 04 in the standby working group 02, and the second virtual IP of this any normal standby device 04 is set to the second virtual IP of this any primary device 03, and it is confirmed that this any normal standby device 04 is a primary device 03 of the current second cluster.

[0078] Optionally, after the service of this any primary device 03 is migrated to any normal standby device 04 in the standby working group 02, the method further includes: confirming that the primary control board 05 of this any primary device 03 has returned to normal, migrating the service of this any normal standby device 04 back to this any primary device 03, and setting the second virtual IP of this any normal standby device 04 to be empty.

[0079] Optionally, at least one standby device 04 in the standby working group 02 includes a first standby device 10, and the primary control board 05 of the first standby device 10 includes a scheduling management control service working machine 11; the method further includes: receiving a cluster formation request from a client, where the cluster formation request includes the primary working group 01 and the standby working group 02 specified by the client; obtaining the service configurations of all first-level devices included in the primary working group 01 and the standby working group 02 specified by the client according to the cluster formation request; deploying the first scheduling management service 08 and the second scheduling management service 09 in the service configurations of each first-level device among all first-level devices through the scheduling management control service working machine 11 to obtain the updated service configurations of each first-level device; forming a secondary cluster based on the updated service configurations of each first-level device; and managing the first scheduling management service 08 and the second scheduling management service 09 on each first-level device through the scheduling management control service working machine 11.

[0080] Optionally, the standby working group 02 further includes a second standby device 13, and the second standby device 13 is at least one standby device 04 in the standby working group 02 other than the first standby device 10; the method further includes: creating a scheduling management control service standby machine 12 through the scheduling management control service working machine 11, synchronizing the service configuration of the scheduling management control service working machine 11 to the scheduling management control service standby machine 12, and deploying the scheduling management control service standby machine 12 to the second standby device 13 in the standby working group 02; monitoring the running status of the scheduling management control service standby machine 12 of the second standby device 13 through the scheduling management control service working machine 11; when it is monitored that the scheduling management control service standby machine 12 of any second standby device 13 is abnormal and the number of first-level devices connected to the port service of the scheduling management service standby machine 12 is greater than the first threshold, removing the scheduling management control service standby machine 12 on the any second standby device 13; if the second standby device 13 in the standby working group 02 only includes the any second standby device 13, determining whether there is a third standby device in the standby working group 02, and the third standby device is a first-level device other than the any second standby device 13 and the first standby device 10, if it exists, deploying the scheduling management control service standby machine 12 to the third standby device, if it does not exist, triggering an alarm.

[0081] Optionally, the method further includes: monitoring the running status of the scheduling management control service working machine 11 of the first standby device 10 through the scheduling management control service standby machine 12; when it is monitored that the scheduling management control service working machine 11 of the first standby device 10 is abnormal and the number of first-level devices connected to the port service of the scheduling management service working machine 11 is greater than the second threshold, managing the first scheduling management service 08 and the second scheduling management service 09 on each first-level device based on the scheduling management control service standby machine 13.

[0082] It should be understood that all relevant contents of the steps involved in the above method embodiments can be cited in the function descriptions of the corresponding functional modules of the system, and will not be repeated here.

[0083] Based on the same technical concept, see Figure 8 , the embodiments of the present application further provide an electronic device 800, including: At least one processor 801; and a communication interface 803 communicatively connected to the at least one processor 801; the at least one processor 801 executes instructions stored in the memory 802, so that the electronic device 800 executes the method steps performed by the two-dimensional cluster scheduling system in the above method embodiments through the communication interface 803.

[0084] Optionally, the memory 802 is located outside the electronic device 800.

[0085] Optionally, the electronic device 800 includes the memory 802, which is connected to the at least one processor 801, and the memory 802 has instructions executable by the at least one processor 801. Attached Figure 8 The memory 802 is shown as optional for the electronic device 800 with a dashed line.

[0086] Wherein, the at least one processor 801 and the memory 802 may be coupled through an interface circuit or integrated together, which is not limited herein.

[0087] In the embodiments of the present application, the specific connection medium between the at least one processor 801, the memory 802, and the communication interface 803 is not limited. In the embodiments of the present application Figure 8 it is shown that the at least one processor 801, the memory 802, and the communication interface 803 are connected through a bus 804, and the bus is shown as a thick line in Figure 8 The connection manners between other components are only for illustrative purposes and are not limiting. This bus part may be an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 8 it is only shown as a thick line, but it does not mean that there is only one bus or one type of bus.

[0088] It should be understood that the processor mentioned in the embodiments of the present application may be implemented by hardware or by software. When implemented by hardware, the processor may be a logic circuit, an integrated circuit, etc. When implemented by software, the processor may be a general-purpose processor that implements by reading software code stored in the memory.

[0089] Exemplarily, the processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0090] It should be understood that the memory mentioned in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which serves as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0091] It should be noted that when the processor is a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, the memory (storage module) may be integrated in the processor.

[0092] It should be noted that the memory described herein is intended to include but not be limited to these and any other suitable types of memory.

[0093] Based on the same technical concept, the embodiments of the present application also provide a computer-readable storage medium, which is used to store instructions. When the instructions are executed, the computer executes the method steps executed by any of the devices in the above method embodiments.

[0094] Based on the same technical concept, the embodiments of the present application also provide a computer program product, including computer program code. When the computer program code runs on the computer, the method steps executed by any of the devices in the above method embodiments are implemented.

[0095] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0096] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0097] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0099] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A two-level cluster scheduling system, characterized in that, The secondary cluster scheduling system includes an active working group and a standby working group; Among them, the active working group includes at least one active device, and the standby working group includes at least one standby device; the active device and the standby device are first-level devices, and each first-level device includes an active control board and a standby control board; each active control board and each standby control board include a main service, a first scheduling management service, and a second scheduling management service; The main service is used to: execute a preset service; The first scheduling management service is used to: monitor the operating status of the active control board and the standby control board in each first-level device, and initiate an active-standby control board switch when the active control board of any first-level device is abnormal and the standby control board is normal; the active-standby control board switch is used to indicate migrating the services of the active control board of any first-level device to the standby control board of any first-level device; The second scheduling management service is used to: monitor the operating status of the active control board and the standby control board in each first-level device, and initiate an active-standby device switch when the active control board and the standby control board of any active device are both abnormal; the active-standby device switch is used to indicate migrating the services of any active device to any normal standby device in the standby working group.

2. The system according to claim 1, wherein Before monitoring the operating status of the active control board and the standby control board in each first-level device, the first scheduling management service is also used to: Initialize each first-level device, and initiate active control board negotiation according to the historical configuration record; the initialization includes resetting the two control boards in each first-level device; the historical configuration record includes the active control board and the standby control board of each first-level device configured before initialization; the active control board negotiation is used to indicate determining the active control board among the two control boards of each first-level device after initialization; If the historical configuration record is an active control board and a standby control board, set the active control board in the historical configuration record as the active control board of the first-level device after initialization; If the historical configuration record is two active control boards, set the active control board in the latest configuration record in the historical configuration record as the active control board of the first-level device after initialization; If the historical configuration record is two standby control boards or empty, set the control board in the preset slot as the active control board of the first-level device after initialization.

3. The system according to claim 1, wherein When the first scheduling management service monitors the operating status of the active control board and the standby control board of each first-level device and initiates an active-standby control board switch when the active control board of any first-level device is abnormal and the standby control board is normal, it is specifically used to: Monitor the service configuration change of the active control board, and update the service configuration record of the active control board based on the service configuration change; Monitor changes in the business working status of the main control board. If it is monitored that the business working status of the main control board is abnormal and the operating status of the standby control board is normal, set the first virtual IP of the standby control board as the first virtual IP of the main control board, and configure the standby control board based on the business configuration record of the main control board to determine that the standby control board is the main control board of any of the current first-level devices.

4. The system according to claim 1, wherein The first dispatching management service is further used to: report the information of the switching of the main and standby control boards to the second dispatching management service; The second scheduling management service monitors the operating status of the main control board and the standby control board in each primary device, and initiates the switching of the main and standby devices when both the main control board and the standby control board of any main device are abnormal, specifically for: Monitoring the operating status of the main control board and the standby control board in each of the primary devices; If it is monitored that both the main control board and the backup control board of any of the main devices are abnormal, or, it is monitored that the main control board of any of the main devices is abnormal, and the information on the switching of the main and backup control boards reported by the first scheduling management service of any of the main devices is not received within the preset time threshold, then the business of any of the main devices is migrated to any normal backup device in the backup workgroup, and the second virtual IP of any of the normal backup devices is set as the second virtual IP of any of the main devices, confirming that any of the normal backup devices is a main device of the current secondary cluster.

5. The system according to claim 4, wherein After migrating the service of any active device to any normal standby device in the standby working group, the second scheduling management service is further used to: Confirm that the active control board of any active device has recovered to normal, migrate the services of any normal standby device back to any active device, and set the second virtual IP of any normal standby device to empty.

6. The system according to claim 1, wherein At least one standby device in the standby working group includes a first standby device, and a main control board of the first standby device includes a scheduling management control service working machine; The first backup device is used for: Receiving a cluster establishment request from a client, wherein the cluster establishment request includes a primary workgroup and a backup workgroup specified by the client; Obtaining, according to the cluster establishment request, service configurations of all first-level devices included in the active working group and the standby working group specified by the client; Deploy the first scheduling management service and the second scheduling management service in the service configuration of each first-level device among all the first-level devices by the scheduling management control service worker to obtain the updated service configuration of each first-level device; Based on the updated service configuration of each primary device, forming the secondary cluster; The first scheduling management service and the second scheduling management service on each primary device are managed by the scheduling management control service worker.

7. The system according to claim 6, wherein The standby working group further includes a second standby device, where the second standby device is at least one standby device in the standby working group except the first standby device; The first standby device is further configured to: create a scheduling management control service standby machine based on the scheduling management control service working machine, synchronize the service configuration of the scheduling management control service working machine to the scheduling management control service standby machine, and deploy the scheduling management control service standby machine to a second standby device in the standby working group; Monitor the running status of the scheduling management control service standby machine of the second standby device through the scheduling management control service working machine; when it is detected that the scheduling management control service standby machine of any second standby device is abnormal and the number of first-level devices whose port services are connected to the scheduling management service standby machine is greater than a first threshold, remove the scheduling management control service standby machine on the any second standby device; If the second standby device in the standby working group only includes the any second standby device, determine whether there is a third standby device in the standby working group, where the third standby device is a first-level device other than the any second standby device and the first standby device. If there is, deploy the scheduling management control service standby machine to the third device. If not, trigger an alarm.

8. The system according to claim 7, wherein The second standby device is configured to: Monitor the running status of the scheduling management control service working machine of the first standby device through the scheduling management control service standby machine; when it is detected that the scheduling management control service working machine of the first standby device is abnormal and the number of first-level devices whose port services are connected to the scheduling management service working machine is greater than a second threshold, manage the first scheduling management service and the second scheduling management service on each first-level device based on the scheduling management control service standby machine.

9. A two - level cluster scheduling method, applied to a two - level cluster scheduling system, is characterized in that, The secondary cluster scheduling system includes a primary working group and a standby working group; wherein, the primary working group includes at least one primary device, and the standby working group includes at least one standby device; the primary device and the standby device are first-level devices, and each first-level device includes a primary control board and a standby control board; each primary control board and each standby control board include a primary service, a first scheduling management service, and a second scheduling management service. The method includes: Monitor the running status of the primary control board and the standby control board in each first-level device. When the primary control board of any first-level device is abnormal and the standby control board is normal, initiate a primary-standby control board switch; the primary-standby control board switch is used to indicate migrating the services of the primary control board of the any first-level device to the standby control board of the any first-level device; Monitor the running status of the primary control board and the standby control board in each first-level device. When the primary control board and the standby control board of any primary device are both abnormal, initiate a primary-standby device switch; the primary-standby device switch is used to indicate migrating the services of the any primary device to any normal standby device in the standby working group.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a computer, the method according to any one of claims 9 is implemented.