Task Processing Method, Device, Computer-Readable Medium, and Electronic Device
By detecting the working status of the task processing node in the task processing system and using the message queue to reassign tasks, the task cannot be handled due to the downtime of the task processing node is solved, and the system's disaster recovery capabilities and availability are improved.
Patent Information
- Application Number
- CN202110492295.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-06
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-05-06
AI Technical Summary
In a task processing system, the downtime of the task processing node causes the assigned tasks to be unable to be processed, reducing the availability of the system.
By detecting the working status of the task processing node, if the node is detected to be down, other task processing nodes can obtain unprocessed task identification information from the message queue and process it to ensure timely processing of the task.
The system's disaster recovery capability and availability are improved, ensuring that the downtime of task processing nodes does not lead to task loss, realizing the decoupling of task allocation and processing, and improving the reliability of the system through message queues.
Smart Images

Figure CN113064744B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of task processing. Specifically, it relates to a task processing method, apparatus, computer-readable medium, and electronic device applied to a multi-node system. Background Art
[0002] In a task processing system, usually, a task dispatcher assigns tasks to task processing nodes, and the task processing nodes process the assigned tasks. However, in the related art, once a task processing node fails or crashes, then this task processing node will be unable to process the tasks assigned to it, resulting in the tasks not being processed, thereby making the availability of the entire task processing system relatively low. Summary of the Invention
[0003] Embodiments of this application provide a task processing method, apparatus, computer-readable medium, and electronic device applied to a multi-node system, which can, at least to a certain extent, avoid tasks assigned not being processed due to a task processing node crashing, and can improve the disaster tolerance and availability of the system.
[0004] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0005] According to one aspect of the embodiments of this application, there is provided a task processing method applied to a multi-node system. The multi-node system includes multiple task processing nodes. The method includes: detecting the working status of the multiple task processing nodes; obtaining the identification information of tasks to be consumed from a message queue, where the message queue contains the identification information of first tasks that have not been processed completely by a first task processing node that has crashed, and the identification information of the first tasks is obtained by a second task processing node after detecting that the first task processing node has crashed and is put into the message queue; and processing the tasks to be consumed according to the obtained identification information of the tasks to be consumed.
[0006] According to one aspect of the embodiments of this application, there is provided a task processing apparatus applied to a multi-node system. The multi-node system includes multiple task processing nodes. The apparatus includes: a detection unit configured to detect the working status of the multiple task processing nodes; an obtaining unit configured to obtain the identification information of tasks to be consumed from a message queue, where the message queue contains the identification information of first tasks that have not been processed completely by a first task processing node that has crashed, and the identification information of the first tasks is obtained by a second task processing node after detecting that the first task processing node has crashed and is put into the message queue; and a processing unit configured to process the tasks to be consumed according to the obtained identification information of the tasks to be consumed.
[0007] In some embodiments of the present application, based on the foregoing solution, the obtaining unit is further configured to: if it is detected that the first task processing node fails, obtain the identification information of the first task that the first task processing node has not completed processing from other task processing nodes among the multiple task processing nodes in a competitive manner; wherein, the second task processing node includes the task processing node that succeeds in the competition.
[0008] In some embodiments of the present application, based on the foregoing solution, the task processing device applied to the multi-node system is located in a designated task processing node among the multiple task processing nodes, and the obtaining unit is further configured to: if the designated task processing node is the task processing node with the longest continuous normal working time in the multi-node system, obtain the identification information of the second task that has not been completed in the multi-node system; and put the identification information of the second task into the message queue.
[0009] In some embodiments of the present application, based on the foregoing solution, the obtaining unit is configured to: if the designated task processing node is the task processing node with the longest continuous normal working time in the multi-node system, obtain the identification information of the second task that has not been completed in the multi-node system after waiting for a predetermined time.
[0010] In some embodiments of the present application, based on the foregoing solution, before detecting the working states of the multiple task processing nodes, the detecting unit is further configured to: send a registration request to the registration center and establish a heartbeat connection with the registration center after successful registration; the detecting unit is configured to: detect the state of the heartbeat connection maintained in the registration center to detect the working states of the multiple task processing nodes.
[0011] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: load the task entity binary data corresponding to the identification information of the task to be consumed from the cache corresponding to the multi-node system; perform deserialization processing on the task entity binary data to obtain the task to be consumed; and process the task to be consumed.
[0012] In some embodiments of the present application, based on the foregoing solution, the task to be consumed includes at least one network instruction, and the task entity binary data includes the original binary data corresponding to the network instruction; the processing unit is configured to: sequentially execute the network instructions included in the task to be consumed, and wherein, after each network instruction is successfully executed, perform serialization processing on each network instruction, and replace the original binary data corresponding to each network instruction in the cache with the binary data obtained after serialization processing.
[0013] In some embodiments of the present application, based on the foregoing solution, the processing unit is further configured to: if the execution of the network instruction included in the task to be consumed fails, roll back the network instructions that have been executed in the task to be consumed, so as to reprocess the task to be consumed.
[0014] In some embodiments of the present application, based on the foregoing solution, the processing unit is further configured to: after loading the task entity binary data corresponding to the identification information of the task to be consumed from the cache corresponding to the multi-node system, persist the task entity binary data in the cache to the disk.
[0015] In some embodiments of the present application, based on the foregoing solution, a network controller is deployed on each task processing node in the multi-node system, and the network controller includes a task scheduling and management framework, and the task scheduling and management framework is used to sequentially execute the network instructions included in the task to be consumed.
[0016] In some embodiments of the present application, based on the foregoing solution, the network controller further includes a remote call client, and the processing unit is configured to: sequentially send the network instructions in the task to be consumed to the remote call client, so that the remote call client sequentially sends the network instructions to the network devices corresponding to the network instructions.
[0017] In some embodiments of the present application, based on the foregoing solution, the network controller further includes a task placement client, and the identification information of a third task is further included in the message queue, and the identification information of the third task is placed in the message queue by the task placement client.
[0018] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the task processing method applied to a multi-node system as described in the foregoing embodiments is implemented.
[0019] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the task processing method applied to a multi-node system as described in the foregoing embodiments.
[0020] In the technical solutions provided by some embodiments of the present application, when a task processing node fails, other task processing nodes can redeliver the identification information of the tasks that the task processing node has not completed to the message queue. Then, other task processing nodes can obtain the identification information of the tasks to be consumed from the message queue, so as to process the tasks that the task processing node has not completed. Therefore, the embodiments of the present application allow task processing nodes in a multi-node system to fail. Even when a task processing node fails, the tasks processed by it will not be lost, and the tasks that the task processing node has not completed can be processed in a timely manner, improving the disaster tolerance of the system and greatly improving the availability of the system. In addition, the task processing node processes the tasks to be consumed according to the obtained identification information of the tasks to be consumed. Therefore, the decoupling of task allocation and task processing is realized, and the message queue can also smooth the task peaks, thereby further improving the reliability of the system and reducing the complexity of the system.
[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0023] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied is shown;
[0024] Figure 2 A flowchart of a task processing method applied to a multi-node system according to an embodiment of the present application is shown;
[0025] Figure 3 A schematic diagram of the organizational form of a network controller and a task orchestration management framework on a node according to an embodiment of the present application is shown;
[0026] Figure 4 Shows according to an embodiment of the present application Figure 2 The flowchart of the steps before step 220 and the details of step 220;
[0027] Figure 5 A schematic diagram of the processing flow of a multi-node system when a certain node fails according to an embodiment of the present application is shown;
[0028] Figure 6 Shows a flowchart of processing a task to be consumed according to the identification information of the task to be consumed obtained according to an embodiment of the present application;
[0029] Figure 7 Shows a schematic diagram of the processing flow of a multi-node system when all nodes are down according to an embodiment of the present application;
[0030] Figure 8 Shows a schematic diagram of a task execution process based on a task orchestration management framework according to an embodiment of the present application;
[0031] Figure 9 Shows a block diagram of a task processing device applied to a multi-node system according to an embodiment of the present application;
[0032] Figure 10 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0033] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0034] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0035] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0036] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0037] In the field of network control, orchestrating instructions for network controllers is an important task. The inventors of this application have found that for the orchestration of instructions to proceed smoothly, the following characteristics need to be met:
[0038] First, since there are interdependent relationships between instructions, it is crucial to ensure the atomicity of the orchestration for instruction orchestration. Second, if the orchestrated instructions or tasks are lost, it will directly affect the success rate of instruction orchestration. Therefore, the network controller also needs to have strong disaster tolerance capabilities to ensure that the instructions and tasks orchestrated after the controller crashes are not lost. Third, for the current development of controller instruction orchestration, since each set of network controllers has its own logic, it is necessary to improve the overall development efficiency.
[0039] However, in the traditional process of network controller instruction orchestration, it is difficult to ensure the atomicity of a set of instructions, and often only error correction can be completed through later configuration reviews. Moreover, the network controller itself often does not have disaster tolerance capabilities. Once a software error occurs, the instructions cannot be restored, greatly affecting the effect of instruction orchestration. In addition, currently, corresponding instruction orchestration strategies need to be developed separately for each set of network controllers, resulting in low development efficiency and high development costs for the entire system.
[0040] In related technologies, some open-source task management frameworks can be used to implement task distribution and execution. Such task management frameworks can include, for example, AirFlow. These task management frameworks achieve the orchestration of network instructions by embedding network instructions into tasks.
[0041] However, there are still the following defects in using task management frameworks:
[0042] First, almost all of these open-source task management frameworks are independently deployed and cannot be embedded into the controller as components, which will greatly delay the issuance of network instructions.
[0043] Second, these task management frameworks also often do not have rollback configurations and it is difficult to achieve the atomicity of a set of instruction orchestrations, or it is difficult to conveniently achieve the atomicity of a set of instruction orchestrations.
[0044] Third, the existing open-source task management frameworks have insufficient support for disaster tolerance capabilities, and tasks are likely to be lost when nodes go offline.
[0045] To this end, the present application first provides a task processing method applicable to a multi-node system. The task here can be any task that can be represented by program code and used to process a certain process, such as a task for implementing network instruction orchestration. The task is represented by a certain entity in the program code. For example, in a Java program, the task can be represented in the form of an object. Therefore, a task processing method applicable to a multi-node system provided by an embodiment of the present application can be applied to various task processing scenarios.
[0046] When the task processing method applicable to a multi-node system provided by an embodiment of the present application is used in a scenario for orchestrating network instructions, it can overcome the above defects and can implement the characteristics required for successfully completing instruction orchestration as described above.
[0047] Figure 1 The schematic diagram of an exemplary system architecture to which the technical solution of the embodiment of the present application can be applied is shown. Below, Figure 1 the shown system architecture is used for network instruction orchestration to introduce it.
[0048] As Figure 1 shown, the system architecture 100 may include network devices (such as Figure 1 one or more of the switch 101, router 102, and gateway device 103 shown in
[0049] It should be understood that Figure 1 the number of terminal devices, networks, and task processing nodes in the multi-node system in
[0050] The task processing nodes in the multi-node system 110 orchestrate a set of network instructions included in a task by processing the task, and send the network instructions to the network devices corresponding to the network instructions. When the task processing method applied to the multi-node system provided in the embodiments of the present application is used in a scenario for orchestrating network instructions, if a task processing node experiences a downtime event during the task processing, then, through the task processing method applied to the multi-node system provided in the embodiments of the present application, each task processing node can obtain the identification information of the tasks to be consumed from the message queue, and the message queue contains the identification information of the tasks that the task processing node that has experienced downtime has not processed completely. Therefore, the task processing nodes that have not experienced downtime can obtain the identification information of the tasks that the task processing node that has experienced downtime has not processed completely, so that the tasks that have not been processed completely by the task processing node can be processed by other task processing nodes that have not experienced downtime events, thus ensuring that all the assigned tasks can be effectively processed.
[0051] It should be noted that although the embodiments of the present application are used in a scenario for orchestrating network instructions, in fact, it can be applied to any task processing process, such as in the control process of a workflow; although the task processing nodes in the embodiments of the present application are servers, in other embodiments of the present application, the task processing nodes can be any type of terminal device, the task processing nodes can be not only physical nodes, but also virtualized nodes. In addition, the task processing nodes can also be a cluster of servers; and although the multi-node system in the embodiments of the present application only includes task processing nodes, the multi-node system can also include other various types of entities such as databases and message queues. The embodiments of the present application do not make any limitations in this regard, and the protection scope of the present application should not be restricted accordingly.
[0052] Moreover, it is easy to understand that the task processing method applied to the multi-node system provided in the embodiments of the present application is generally executed by a server. Correspondingly, the task processing device applied to the multi-node system is generally arranged in the server. However, in other embodiments of the present application, the terminal device can also have a similar function to the server, so as to execute the task processing solution applied to the multi-node system provided in the embodiments of the present application.
[0053] The task processing method applied to the multi-node system provided in the embodiments of the present application can be applied not only in the fields of cloud computing or cloud storage, but also in the Blockchain network.
[0054] Cloud computing is a computing model that distributes computing tasks across a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" appear to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.
[0055] As a basic capability provider of cloud computing, a cloud computing resource pool (referred to as the cloud platform, generally called the IaaS (Infrastructure as a Service) platform) will be established, and various types of virtual resources will be deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtual machines, including operating systems), storage devices, and network devices.
[0056] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. The SaaS layer can also be directly deployed on the IaaS. PaaS is a platform for software operation, such as databases, web containers, etc. SaaS is various business software, such as web portals, SMS mass senders, etc. Generally speaking, SaaS and PaaS are the upper layers relative to IaaS.
[0057] Cloud storage is a new concept extended and developed from the cloud computing concept. A distributed cloud storage system (hereinafter referred to as the storage system) refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through cluster applications, grid technology, and distributed storage file systems, etc., and works together through application software or application interfaces to jointly provide data storage and business access functions to the outside world.
[0058] Currently, the storage method of the storage system is as follows: Create a logical volume. When creating a logical volume, physical storage space is allocated for each logical volume. This physical storage space may be a certain storage device or the disks of several storage devices. The client stores data on a certain logical volume, that is, stores the data on the file system. The file system divides the data into many parts, and each part is an object. The object not only contains the data but also contains additional information such as data identification (ID, ID entity). The file system writes each object into the physical storage space of the logical volume respectively, and the file system will record the storage location information of each object. Thus, when the client requests to access the data, the file system can enable the client to access the data according to the storage location information of each object.
[0059] The process of the storage system allocating physical storage space for a logical volume is specifically as follows: According to the capacity estimation of the objects stored in the logical volume (this estimation often has a large margin relative to the capacity of the objects to be actually stored) and the group of the redundant array of independent disks (RAID, Redundant Array of Independent Disk), the physical storage space is pre-divided into stripes. A logical volume can be understood as a stripe, thereby allocating physical storage space for the logical volume.
[0060] Blockchain is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0061] The underlying blockchain platform may include processing modules such as user management, basic services, smart contracts, and operation monitoring. Among them, the user management module is responsible for managing the identity information of all blockchain participants, including maintaining the generation of public and private keys (account management), key management, and maintaining the correspondence between the real identity of users and blockchain addresses (permission management). And under authorization, it supervises and audits the transaction situations of certain real identities, and provides rule configuration for risk control (risk control audit); the basic service module is deployed on all blockchain node devices to verify the validity of business requests, and after reaching a consensus on valid requests, records them in storage. For a new business request, the basic service first performs interface adaptation parsing and authentication processing (interface adaptation), then encrypts the business information through a consensus algorithm (consensus management), and after encryption, transmits it to the shared ledger completely and consistently (network communication), and records and stores it; the smart contract module is responsible for the registration and issuance of contracts, contract triggering, and contract execution. Developers can define contract logic through a certain programming language, publish it to the blockchain (contract registration), and according to the logic of the contract terms, call keys or other events to trigger execution, complete the contract logic, and at the same time also provide functions for contract upgrade and cancellation; the operation monitoring module is mainly responsible for the deployment, configuration modification, contract setting, cloud adaptation during the product release process, and visual output of the real-time status during product operation, such as: alarm, monitoring network conditions, monitoring the health status of node devices, etc.
[0062] The platform product service layer provides the basic capabilities and implementation frameworks of typical applications. Based on these basic capabilities, developers can overlay the characteristics of the business to complete the blockchain implementation of the business logic. The application service layer provides application services based on the blockchain solution for business participants to use.
[0063] When the embodiment of the present application is applied to the scenario of orchestrating network instructions, it can be specifically applied to the field of edge computing and the network architecture of software-defined network (SDN, Software Defined Network). The orchestrated network instructions can be used to control the edge gateway. Software-defined network is a new open network architecture. By separating control from forwarding, centralizing the control plane, and opening the interface for programmability, it realizes flexible scheduling of network resources from a global perspective and rapid deployment of new services, can simplify operation and maintenance, improve network resource utilization, and enhance customer perception.
[0064] The following elaborates in detail on the implementation details of the technical solution of the embodiment of the present application:
[0065] Figure 2 The flowchart of the task processing method applied to a multi-node system according to an embodiment of the present application is shown. The task processing method applied to a multi-node system can be executed by a device with computing and communication functions, such as Figure 1The first server 111 shown in [reference], the multi-node system includes a plurality of task processing nodes. Refer to Figure 2 As shown, the task processing method applied to the multi-node system at least includes the following steps:
[0066] In step 220, detect the working status of a plurality of task processing nodes.
[0067] When used in the scenario of network instruction orchestration in the embodiments of the present application, the task processing node can adopt Figure 3 The organizational form shown, that is, the node includes a network controller, and the network controller further includes a northbound control API, a task orchestration management framework, an RPCClient (Remote Procedure Call Client), and a function module for generating network instructions. The function module for generating network instructions specifically includes a network orchestration and computing module, a network model object management module, and a global quality monitoring traffic scheduling module. Among them, the northbound control API is a user management application interface provided by the network controller northbound to provide an entry for network instruction orchestration; each function module for generating network instructions can generate network configuration instructions through operations such as network model object management and network orchestration and computing; the task orchestration management framework can be used to execute the task processing method applied to the multi-node system provided in the embodiments of the present application; the RPCClient is used to send the network instructions from the task orchestration management framework to the network device corresponding to the network instruction. Hereinafter, unless otherwise specified, the node and the task processing node in the multi-node system are equivalent.
[0068] Each task processing node in the multi-node system can detect the working status of other task processing nodes. Figure 2 The steps shown in the embodiment can be executed by any one of the nodes in the multi-node system. The working status of the task processing node can include a normal state and a down state.
[0069] In an embodiment of the present application, the working status of the task processing node in the multi-node system can be detected in multiple ways.
[0070] Figure 4 Shows a Figure 2 Flowchart of the steps before step 220 and the details of step 220 according to an embodiment of the present application.
[0071] Please refer to Figure 4 , specifically including the following steps:
[0072] In step 210, send a registration request to the registration center and establish a heartbeat connection with the registration center after successful registration.
[0073] The registration center is used to register each task processing node in the multi-node system and perform heartbeat detection on each successfully registered task processing node. The task processing node registers by sending a registration request to the registration center, and the registration request can carry information related to the task processing node, such as the identifier of the task processing node.
[0074] A heartbeat connection between the task processing node and the registration center is established by sending heartbeat packets. Specifically, it can be that the task processing node sends a heartbeat packet to the registration center, or it can be that the registration center sends a heartbeat packet to the task processing node. The registration center completes the heartbeat detection based on the result of the received or sent heartbeat packet.
[0075] In step 220', the status of the heartbeat connection maintained in the registration center is detected to detect the working status of the multiple task processing nodes.
[0076] The status of the heartbeat connection maintained in the registration center can be represented by various data.
[0077] For example, the registration information or identifier information of the task processing nodes with normal heartbeat connections can be maintained in the registration center; when a task processing node detects the registration center, it will obtain all the registration information or identifier information of the nodes that maintain normal heartbeat connections with the registration center and store them locally; when a task processing node crashes, the heartbeat of this task processing node in the registration center stops, and the registration center will remove the registration information or identifier information of this task processing node. Other task processing nodes can determine that this task processing node has crashed based on storing the registration information or identifier information of this task processing node locally but finding that the registration information or identifier information of this task processing node has disappeared in the registration center.
[0078] For another example, the identifier information and status information of the task processing nodes can also be maintained in the registration center, where the status information represents that the heartbeat of the corresponding task processing node in the registration center has stopped or is normal. For example, the status information can be represented by 0 or 1, where 0 can represent that the heartbeat of the corresponding task processing node in the registration center has stopped, and 1 can represent that the heartbeat of the corresponding task processing node in the registration center is normal; when a task processing node detects the status of the heartbeat connection maintained in the registration center, it will obtain the identifier information and the corresponding status information of each task processing node. If the obtained status information represents that the heartbeat has stopped, then it can be determined that the working status of the corresponding task processing node is crashed.
[0079] It can be seen that the methods for detecting the working status of task processing nodes and determining whether a task processing node has crashed can be various and are not limited to those described above.
[0080] Please continue to refer to Figure 2 In step 230, the identification information of the task to be consumed is obtained from the message queue. The message queue contains the identification information of the first task that was not completed by the first task processing node that crashed. The identification information of the first task was obtained by the second task processing node after detecting that the first task processing node crashed and was put into the message queue.
[0081] Identification information is information that can be used to uniquely identify a task. Identification information usually adopts the data format of a string. Identification information can be composed of various symbols such as letters and numbers. For example, the identification information can be a string of numbers.
[0082] Optionally, the identification information of the task can be stored in the cache.
[0083] In an embodiment of the present application, the identification information of the first task is put into the position in the message queue that is consumed first.
[0084] In the embodiment of the present application, since the first task is a task that was not completed by the first task processing node, therefore, by putting the identification information of the first task into the position in the message queue that is consumed first, the task that was not completed can be processed preferentially.
[0085] The first task processing node and the second task processing node are both nodes in a multi-node system. After detecting that the first task processing node has crashed, the second task processing node can obtain the identification information of the first task that was not completed by the first task processing node and put it into the message queue.
[0086] In an embodiment of the present application, the method further includes: if it is detected that the first task processing node has crashed, then obtain the identification information of the first task that was not completed by the first task processing node from other task processing nodes among the multiple task processing nodes in a competitive manner; wherein, the second task processing node includes the task processing node that has succeeded in the competition.
[0087] Specifically, the second task processing node can be the task processing node that has succeeded in the competition.
[0088] That each task processing node obtains the identification information of the first task in a competitive manner refers to the process of screening each task processing node according to certain criteria to obtain the only task processing node that is legally entitled to obtain the identification information of the first task.
[0089] The competition rules adopted in the competition method can be diverse. For example, the competition rule can be that the task processing node that first obtains the identification information of the first task is the task processing node that succeeds in the competition, and this task processing node broadcasts to other task processing nodes to make other task processing nodes stop obtaining the identification information of the first task; the competition rule can also be that the task processing node that first obtains the identification information of the first task is the task processing node that succeeds in the competition, and this task processing node broadcasts to other task processing nodes to make other task processing nodes have no right to deliver it to the message queue even if they obtain the identification information of the first task; in addition, the competition rule can also be that the task processing node with the lowest CPU utilization rate or memory usage rate is the task processing node that succeeds in the competition. The advantage of adopting this competition rule is to ensure that the identification information of the first task can be delivered to the message queue in a timely and accurate manner.
[0090] In an embodiment of the present application, the second task processing node is a task processing node preset in the multi-node system.
[0091] For example, a task processing node with high performance, high configuration, high availability, and a standby power supply can be preset in the multi-node system, and this task processing node is responsible for obtaining the identification information of the tasks that other task processing nodes have not completed. Since the availability of this preset task processing node itself is very high, if it is specified that this task processing node obtains the identification information of the tasks, the high availability of the entire system can also be ensured.
[0092] Figure 5 The schematic diagram of the processing flow of the multi-node system when a certain node crashes according to an embodiment of the present application is shown.
[0093] Please refer to Figure 5 , when a certain node crashes, the processing flow of the multi-node system can be as follows:
[0094] 1. Nodes 01, 02, and 03 are all registered in the registration center, and after successful registration, a heartbeat connection with the registration center is established.
[0095] 2. Nodes 01 and 02 detect the status of the heartbeat connection maintained in the registration center. When the heartbeat of node 03 stops in the registration center, nodes 01 and 02 will detect that node 03 has crashed.
[0096] 3. Nodes 01 and 02 obtain the task list of node 03 from the cache in a competitive manner. The cache stores the task lists of each node, and the task list contains one or more task identification information.
[0097] 4. After Node 02 succeeds in the competition, Node 02 successfully obtains the task list of Node 03 and republishes the identification information of the tasks in the task list that have not been completed by Node 03 to the message queue.
[0098] 5. Since Node 01 and Node 02 have not crashed, Node 01 and Node 02 can obtain the identification information of the tasks to be consumed from the message queue, and thus obtain the tasks to be consumed. Among them, when the identification information of the tasks that have not been completed by Node 03 is obtained, the tasks that have not been completed by Node 03 can be processed.
[0099] Please continue to refer to Figure 2 , in step 240, according to the obtained identification information of the tasks to be consumed, process the tasks to be consumed.
[0100] After obtaining the identification information of the tasks to be consumed, it can be consumed, thereby realizing the processing of the tasks to be consumed.
[0101] Any task processing node that has not crashed in the multi-node system can obtain the identification information of the tasks to be consumed from the message queue.
[0102] When the current task processing node obtains the identification information of the first task, the current task processing node can implement the processing of the first task; when other task processing nodes obtain the identification information of the first task, other task processing nodes can also implement the processing of the first task.
[0103] Therefore, after the identification information of the first task is put into the message queue, whether the identification information of the first task is obtained by the current task node or by other task processing nodes, the recovery processing of the first task can be realized, thereby greatly improving the availability of the system.
[0104] Figure 6 The flowchart shows the processing of the tasks to be consumed according to the obtained identification information of the tasks to be consumed according to an embodiment of the present application. Please refer to Figure 6 , step 240 may specifically include the following steps:
[0105] Step 610, load the task entity binary data corresponding to the identification information of the tasks to be consumed from the cache corresponding to the multi-node system.
[0106] The task entity binary data and the identification information of the tasks to be consumed can be stored in the same cache or in different caches respectively.
[0107] The cache can belong to a multi-node system or be outside the multi-node system. The cache is a high-speed storage space jointly used by task processing nodes in the multi-node system. The cache can be located on a physical node or a virtual node, and the physical essence of the cache can be memory. The advantage of using the cache is that it can improve the loading efficiency of the binary data of the task entity, thereby improving the recovery efficiency of the task after the task processing node fails, and thus improving the availability of the system.
[0108] In other embodiments of the present application, the binary data of the task entity can also be stored on a disk.
[0109] The binary data of the task entity is a byte sequence obtained by serializing the task. The binary data of the task entity is flushed into the cache when the task is allocated. By storing the task in the form of the binary data of the task entity, it is convenient for task storage and transmission.
[0110] In one embodiment of the present application, the method further includes:
[0111] After loading the binary data of the task entity corresponding to the identification information of the task to be consumed from the cache corresponding to the multi-node system, persist the binary data of the task entity in the cache to the disk.
[0112] Step 620, perform deserialization processing on the binary data of the task entity to obtain the task to be consumed.
[0113] Performing deserialization processing on the binary data of the task entity can obtain an object of the task to be consumed, and the object can be represented in the form of a Java object.
[0114] Step 630, process the task to be consumed.
[0115] In one embodiment of the present application, the task to be consumed includes at least one network instruction, and the binary data of the task entity includes the original binary data corresponding to the network instruction.
[0116] The step of processing the task to be consumed may include:
[0117] Sequentially execute the network instructions included in the task to be consumed. Among them, after each network instruction is successfully executed, serialize each network instruction, and replace the original binary data corresponding to each network instruction in the cache with the binary data obtained after serialization processing.
[0118] After all network instructions in a task to be consumed are successfully executed, the original binary data corresponding to each network instruction in the cache will be overwritten with binary data.
[0119] The binary data of the network instruction replaces the original binary data corresponding to the network instruction in the cache, which can also be referred to as a snapshot of the binary data of the task entity.
[0120] Obviously, in the scenario where the embodiments of the present application are used to orchestrate network instructions, after a network instruction is successfully executed, the network instruction is serialized, and the binary data obtained through the serialization process replaces the original binary data corresponding to the network instruction in the cache. The original binary data and the binary data corresponding to a network instruction are different, representing different states of the network instruction. Therefore, it can be determined that if there is original binary data in the cache, it means that the network instruction corresponding to the original binary data has not been successfully executed. If there is binary data in the cache, it means that the network instruction corresponding to the binary data has been successfully executed.
[0121] Based on this, after the identification information of the first task is put into the message queue, if some network instructions in the first task are successfully executed and some are not, then the cache includes the binary data corresponding to some network instructions and the original binary data corresponding to some other network instructions. At this time, when the specified task processing node in the multi-node system obtains the identification information of the first task, it can only load the original binary data corresponding to the first task, so as to only process the network instructions in the first task that have not been successfully executed, thus avoiding the repeated execution of the network instructions that have been successfully executed in the first task and ensuring the accuracy of the instruction orchestration task recovery after the node crashes.
[0122] On this basis, each task processing node can distinguish which tasks and which network instructions in the tasks have been executed. Based on these characteristics, the solutions in the above embodiments can be adaptively adjusted.
[0123] Specifically, please continue to refer to Figure 5 , although the nodes that did not crash in Figure 5 obtain the task list of the crashed node, and the task list may include the identification information of the tasks that have been processed and the identification information of the tasks that have not been processed at the same time; but in other embodiments, it is also possible to only obtain the identification information of the tasks in the task list that have not been processed by the crashed node.
[0124] Moreover, when a node obtains a task list, the identification information of the tasks pushed by the node to the message queue can be either the identification information of the tasks that have not been processed in the task list or the identification information of all tasks in the task list. The reason for being able to push the identification information of all tasks to the message queue is that the node that obtains the identification information of the task to be consumed from the message queue can distinguish the network instructions that have been successfully executed according to the data in the cache, thus avoiding repeated execution.
[0125] In one embodiment of the present application, the method further includes: after the to-be-consumed task is successfully executed, persisting the binary data corresponding to the to-be-consumed task in the cache to the disk.
[0126] In the embodiment of the present application, the binary data is persistently processed, and the task processing records are written to the disk, so that the processed tasks can be traced back and audited afterwards.
[0127] The purpose of the persistent processing is to retain more data for traceability and auditing. Therefore, the manner of implementing persistence in the embodiment of the present application can be arbitrary. Persistent processing can be performed at any stage of task processing, can be performed regularly, or can be performed according to user instructions. Persistent processing can be performed only on binary data, or can be performed on the binary data of the task entity including the original binary data.
[0128] In one embodiment of the present application, the step of processing the to-be-consumed task further includes: if the execution of the network instruction included in the to-be-consumed task fails, rolling back the network instruction that has been executed in the to-be-consumed task to reprocess the to-be-consumed task.
[0129] Specifically, the to-be-consumed task includes a plurality of network instructions that need to be executed in sequence. When the execution of a network instruction included in the to-be-consumed task fails, the network instructions executed before the execution of this network instruction need to be re-executed, which ensures the atomicity of the execution of the network instruction, avoids errors in the execution of the network instruction, and thus improves the reliability of the execution of the network instruction.
[0130] In one embodiment of the present application, a network controller is deployed on each task processing node in the multi-node system. The network controller includes a task scheduling and management framework, and the task scheduling and management framework is used to sequentially execute the network instructions included in the to-be-consumed task.
[0131] Please continue to refer to Figure 3 , which shows the organizational form of a node in the multi-node system, and each node in the multi-node system can adopt this organizational form.
[0132] In the embodiment of the present application, the task scheduling and management framework is embedded in the network controller as a component, which greatly reduces the latency of network instruction distribution. At the same time, since a set of task scheduling and management frameworks can be respectively embedded in the network controllers of different task processing nodes, the reusability of the program code is also improved; and secondary development can be carried out based on the task scheduling and management framework to enable different network controllers to implement different instruction scheduling strategies, so the development efficiency can also be improved; in addition, each task scheduling and management framework can be maintained and tested separately, which also reduces the coupling of the system.
[0133] In one embodiment of the present application, sequentially executing the network instructions included in the task to be consumed includes: sequentially sending the network instructions included in the task to be consumed to the network device corresponding to the network instruction.
[0134] The network device here can be various types of network devices such as an edge gateway.
[0135] The successful execution of the network instruction can mean that the network instruction is successfully sent to the network device corresponding to the network instruction, and it can also mean that after the network instruction is sent to the corresponding network device, a response message representing the successful execution of the network instruction is received.
[0136] In one embodiment of the present application, the network controller further includes a remote call client. The step of sequentially executing the network instructions included in the task to be consumed may include: sequentially sending the network instructions in the task to be consumed to the remote call client, so that the remote call client sequentially sends the network instructions to the network device corresponding to the network instruction.
[0137] Please refer to Figure 3 , the RPCClient in the network controller is the remote call client. Therefore, the embodiment of the present application also realizes the orchestration of network instructions based on the remote call protocol.
[0138] In one embodiment of the present application, the method is executed by a designated task processing node among multiple task processing nodes, and the method further includes:
[0139] If the designated task processing node is the task processing node with the longest continuous normal working time in the multi-node system, obtain the identification information of the second task that has not been processed and completed in the multi-node system;
[0140] Put the identification information of the second task into the message queue.
[0141] The task processing nodes in the multi-node system can form a cluster. The task processing node with the longest continuous normal working time, that is, the task processing node with the longest duration of maintaining a normal heartbeat connection with the registration center in the multi-node system, can also be called the oldest node in the cluster.
[0142] In one embodiment of the present application, if the designated task processing node is the task processing node with the longest continuous normal working time in the multi-node system, the step of obtaining the identification information of the second task that has not been processed and completed in the multi-node system may specifically include:
[0143] If the specified task processing node is the task processing node with the longest continuous normal working time in the multi-node system, the identification information of the second task that has not been processed in the multi-node system is obtained after waiting for a predetermined time.
[0144] By obtaining the identification information of the second task after waiting for a predetermined time, the tasks queued up at the specified task processing node can be reduced, and the load on the specified task processing node can be reduced.
[0145] The predetermined time can be set arbitrarily as needed. For example, it can be set to 2 minutes.
[0146] Figure 7 The schematic diagram of the processing flow of the multi-node system when all nodes are down according to an embodiment of the present application is shown. Among them, Figure 7 The schematic diagram of the relationship between the node status and time is shown in the upper right corner. Please refer to Figure 7 When all nodes are down, the processing flow of the multi-node system can be as follows:
[0147] 1. Node 01 fails and exits at time point T0.
[0148] 2. Node 02 fails at time point T0 and exits at time point T2. The reason why the failure time point and the exit time point of node 02 are different is due to heartbeat delay, and node 02 is not discovered to have failed by the registration center until time point T2.
[0149] 3. Node 03 goes online at time point T1. According to the solution in the foregoing embodiment, node 03 can detect that node 02 has failed, can obtain the task list of node 02 from the cache, and put the tasks in the task list into the message queue, so as to resume the execution of the tasks not completed by node 02. Node 03 can take over these tasks, while the tasks not completed by node 01 cannot be taken over.
[0150] 4. When node 03 discovers through the registration center that it is the oldest node in the current cluster, that is, when it discovers that it is the node with the longest continuous normal working time, it will execute a cluster recovery event after 2 minutes (only for example).
[0151] 5. The cluster recovery event will obtain the identification information of the tasks in the cache whose execution time of all tasks is earlier than the birth time of the oldest node (node 03) in the current cluster according to the task list of each node in the cache and the set of task execution times, and re-deliver the identification information of these tasks to the message queue.
[0152] 6. Node 03 obtains the identification information of the tasks from the message queue to consume the tasks, so that the tasks not completed by node 01 can be resumed.
[0153] In the embodiments of the present application, even if all task processing nodes in a multi-node system crash, as long as one task processing node goes online, the tasks that other task processing nodes have not completed can be continued to be executed, thereby further improving the availability of the system.
[0154] In the example mentioned above, the message queue only contains the identification information of the first task that the node has not completed, but the message queue can also include the identification information of the tasks assigned to the task processing nodes under normal circumstances.
[0155] In an embodiment of the present application, the network controller further includes a task delivery client, and the message queue also contains the identification information of the third task, which is delivered to the message queue by the task delivery client.
[0156] The task delivery client is responsible for sending the identification information of the tasks that need to be processed under normal circumstances to the task orchestration management framework. The task orchestration management framework can not only process the uncompleted tasks of other nodes when other nodes crash, but also process the tasks that need to be executed under normal circumstances. For the tasks that have not been completed by other task processing nodes, the current task processing node can process them in the same way as the tasks that need to be processed under normal circumstances.
[0157] Please refer to Figure 3 , the task delivery client can include various functional modules for generating network instructions, such as a network orchestration and computing module, a network model object management module, and a global quality monitoring traffic scheduling module, etc. The position of the task delivery client in the network controller is similar to the positions of these modules in the network controller. Therefore, each task processing node can include a task delivery client and a task orchestration management framework, and both the task delivery client and the task orchestration management framework are embedded on the network controller of the task processing node.
[0158] Figure 8 Fig. shows a schematic diagram of a task execution process based on a task orchestration management framework according to an embodiment of the present application.
[0159] Next, first introduce the components and corresponding functions in the system based on the task orchestration management framework.
[0160] (1) Client: Responsible for generating network instructions, combining a set of network instructions into tasks for submission. Here, the client is equivalent to the task delivery client in the foregoing embodiments.
[0161] (2) Message queue: The bridge between the client and the task orchestration management framework, responsible for peak shaving of the task volume.
[0162] (3) Storage / Cache: To complete the persistence and caching of tasks and instructions, ensuring that tasks are not lost, recoverable, traceable, and auditable.
[0163] (4) Registration Center: Responsible for service registration and discovery, to ensure the high availability and disaster tolerance of the task framework.
[0164] (5) Task Orchestration and Management Framework: The specific executor of network instructions, responsible for executing a set of instructions in the completed task, and constantly refreshing the snapshot to ensure that tasks are not lost and can be taken over by the task orchestration framework of other nodes when the node fails.
[0165] Among the above components, a client and a task orchestration and management framework can be located on the same node. Each node can contain a client and a task orchestration and management framework. The registration center and the message queue can be located on any node, and the storage / cache can also be located on one or more nodes. Additionally, multiple message queues can be set up to improve the availability of the system.
[0166] Figure 8 The task execution process based on the task orchestration and management framework in the embodiment is as follows:
[0167] a) At the beginning of system startup, each task orchestration and management framework needs to complete registration in the registration center and detect the status of other task orchestration and management frameworks, so as to ensure that when other frameworks / nodes fail, it can be detected in time and perform fault migration. Here, the system refers to the system including the task orchestration and management framework and the above components.
[0168] b) The client sends the task ID to the message queue according to the user's instruction, and the client also puts the task ID and the corresponding task entity binary into the cache.
[0169] c) The task orchestration and management framework obtains the task ID from the message queue.
[0170] In an embodiment of the present application, the task orchestration and management framework obtains the task ID sent by the client corresponding to the task orchestration and management framework from the message queue.
[0171] In this case, the task list of the node in the cache of the foregoing embodiment can be written by the client, that is, the client specifies the tasks to be processed by the node, and for the task IDs of the tasks not completed by other nodes in the task list, the node that obtains the task ID can write them into the cache immediately after obtaining the task ID.
[0172] In other embodiments of the present application, the task orchestration and management framework can also obtain the task IDs sent by other clients.
[0173] d) Load the corresponding task entity binary according to the task ID.
[0174] e) Deserialize the task entity in binary to obtain the task.
[0175] f) Sequentially execute the network instructions in the task. When the execution of a network instruction fails, the framework will automatically roll back the network instructions that have been executed. This step is a key step to ensure atomicity. Executing a network instruction means sending the network instruction to the network device corresponding to the network instruction.
[0176] g) Serialize the task to obtain the task binary.
[0177] h) Take a snapshot of the task binary and flush it into the cache. The snapshot here is the task binary located in the cache. This step is a key step for high availability and disaster tolerance. Each time the task snapshot is refreshed, it is ensured that when the framework / node crashes, other frameworks / nodes can resume the execution of the task based on the snapshot information, ensuring the completion of the task.
[0178] i) If the task is not completed, continue to execute step e) and subsequent steps.
[0179] j) Persist the task to the disk for storage and perform disk writing.
[0180] In summary, according to the task processing method applied to a multi-node system in the embodiments of the present application, the atomicity of instruction orchestration can be guaranteed, it can ensure that the instruction tasks orchestrated by the controller are not lost after a crash, and it can also improve the overall development efficiency.
[0181] The following introduces the device embodiments of the present application, which can be used to execute the task processing method applied to a multi-node system in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the embodiments of the task processing method applied to a multi-node system above.
[0182] Figure 9 The block diagram of a task processing device applied to a multi-node system according to an embodiment of the present application is shown.
[0183] Refer to Figure 9 As shown, a task processing device 900 applied to a multi-node system according to an embodiment of the present application includes: a detection unit 910, an acquisition unit 920, and a processing unit 930.
[0184] Among them, the detection unit 910 is used to detect the working status of the multiple task processing nodes; the acquisition unit 920 is used to obtain the identification information of the task to be consumed from the message queue, and the message queue contains the identification information of the first task that has not been processed completely by the first task processing node that has crashed. The identification information of the first task is obtained by the second task processing node after detecting the crash of the first task processing node and then put into the message queue; the processing unit 930 is used to process the task to be consumed according to the obtained identification information of the task to be consumed.
[0185] In some embodiments of the present application, based on the foregoing solution, the acquisition unit 920 is further configured to: if it is detected that the first task processing node has crashed, obtain the identification information of the first task that has not been processed completely by the first task processing node through competition with other task processing nodes among the multiple task processing nodes; wherein, the second task processing node includes the task processing node that has succeeded in the competition.
[0186] In some embodiments of the present application, based on the foregoing solution, the task processing device 900 applied to the multi-node system is located in a specified task processing node among the multiple task processing nodes, and the acquisition unit 920 is further configured to: if the specified task processing node is the task processing node with the longest continuous normal working time in the multi-node system, obtain the identification information of the second task that has not been processed completely in the multi-node system; put the identification information of the second task into the message queue.
[0187] In some embodiments of the present application, based on the foregoing solution, the acquisition unit 920 is configured to: if the specified task processing node is the task processing node with the longest continuous normal working time in the multi-node system, obtain the identification information of the second task that has not been processed completely in the multi-node system after waiting for a predetermined time.
[0188] In some embodiments of the present application, based on the foregoing solution, before detecting the working status of the multiple task processing nodes, the detection unit 910 is further configured to: send a registration request to the registration center and establish a heartbeat connection with the registration center after successful registration; the detection unit 910 is configured to: detect the status of the heartbeat connection maintained in the registration center to detect the working status of the multiple task processing nodes.
[0189] In some embodiments of the present application, based on the foregoing solution, the processing unit 930 is configured to: load the task entity binary data corresponding to the identification information of the task to be consumed from the cache corresponding to the multi-node system; perform deserialization processing on the task entity binary data to obtain the task to be consumed; process the task to be consumed.
[0190] In some embodiments of the present application, based on the foregoing solution, the task to be consumed includes at least one network instruction, and the task entity binary data includes the original binary data corresponding to the network instruction; the processing unit 930 is configured to: sequentially execute the network instructions included in the task to be consumed, wherein after each network instruction is successfully executed, serialize the each network instruction, and replace the original binary data corresponding to the each network instruction in the cache with the binary data obtained after serialization processing.
[0191] In some embodiments of the present application, based on the foregoing solution, the processing unit 930 is further configured to: if the execution of the network instruction included in the task to be consumed fails, roll back the network instructions that have been executed in the task to be consumed to reprocess the task to be consumed.
[0192] In some embodiments of the present application, based on the foregoing solution, the processing unit 930 is further configured to: after loading the task entity binary data corresponding to the identification information of the task to be consumed from the cache corresponding to the multi-node system, persist the task entity binary data in the cache to the disk.
[0193] In some embodiments of the present application, based on the foregoing solution, a network controller is deployed on each task processing node in the multi-node system, and the network controller includes a task orchestration management framework for sequentially executing the network instructions included in the task to be consumed.
[0194] In some embodiments of the present application, based on the foregoing solution, the network controller further includes a remote call client, and the processing unit 930 is configured to: sequentially send the network instructions in the task to be consumed to the remote call client, so that the remote call client sequentially sends the network instructions to the network devices corresponding to the network instructions.
[0195] In some embodiments of the present application, based on the foregoing solution, the network controller further includes a task placement client, and the identification information of a third task is further included in the message queue, and the identification information of the third task is placed in the message queue by the task placement client.
[0196] Figure 10 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0197] It should be noted that Figure 10 The computer system 1000 of the electronic device shown is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.
[0198] Such asFigure 10 As shown in Figure 10 , computer system 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a Read-Only Memory (ROM) 1002 or a program loaded from a storage section 1008 into a Random Access Memory (RAM) 1003, such as executing the methods described in the above embodiments. In the RAM 1003, various programs and data required for system operation are also stored. The CPU 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.
[0199] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including, for example, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed.
[0200] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from the removable medium 1011. When the computer program is executed by a Central Processing Unit (CPU) 1001, various functions defined in the system of the present application are executed.
[0201] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0202] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and the above module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0203] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the unit itself.
[0204] As one aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the above embodiments.
[0205] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0206] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0207] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
[0208] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A task processing method applied to a multi-node system, characterized in that The multi-node system includes multiple task processing nodes, and network controllers are deployed on each task processing node in the multi-node system. The network controller includes a task orchestration management framework and a remote call client. The method is executed by any one of the multiple task processing nodes, and the method includes: Detect the working status of the multiple task processing nodes; Obtain the identification information of the tasks to be consumed from the message queue. The message queue contains the identification information of the first tasks that were not completed by the first task processing node that has crashed. The identification information of the first tasks was obtained by the second task processing node after detecting that the first task processing node has crashed and was put into the message queue; Process the tasks to be consumed according to the obtained identification information of the tasks to be consumed. The tasks to be consumed include at least one network instruction, and the task orchestration management framework is used to sequentially execute the network instructions included in the tasks to be consumed; Wherein, the processing of the tasks to be consumed according to the obtained identification information of the tasks to be consumed includes: Process the tasks to be consumed in the following manner: Execute the network instructions included in the tasks to be consumed in the following manner: sequentially send the network instructions in the tasks to be consumed to the remote call client, so that the remote call client sequentially sends the network instructions to the network devices corresponding to the network instructions; If the execution of the network instructions included in the tasks to be consumed fails, roll back the network instructions that have been executed in the tasks to be consumed to reprocess the tasks to be consumed.
2. The task processing method applied to a multi-node system according to claim 1, wherein The method further includes: If it is detected that the first task processing node has crashed, obtain the identification information of the first tasks that were not completed by the first task processing node through competition with other task processing nodes in the multiple task processing nodes; wherein, the second task processing node includes the task processing node that has succeeded in the competition.
3. The task processing method applied to a multi-node system according to claim 1, characterized in that The method is executed by a designated task processing node among the multiple task processing nodes, and the method further includes: If the designated task processing node is the task processing node with the longest continuous normal working time in the multi-node system, obtain the identification information of the second tasks that were not completed in the multi-node system; Put the identification information of the second tasks into the message queue.
4. The task processing method applied to a multi-node system according to claim 3, characterized in that The obtaining the identification information of the second tasks that were not completed in the multi-node system if the designated task processing node is the task processing node with the longest continuous normal working time in the multi-node system includes: If the designated task processing node is the task processing node with the longest continuous normal working time in the multi-node system, obtain the identification information of the second tasks that were not completed in the multi-node system after waiting for a predetermined time.
5. The task processing method applied to a multi-node system according to claim 1, wherein Before detecting the working status of the multiple task processing nodes, the method further includes: Send a registration request to the registration center and establish a heartbeat connection with the registration center after successful registration; The detecting the working status of the multiple task processing nodes includes: Detect the status of the heartbeat connection maintained in the registry to detect the working status of the multiple task processing nodes.
6. The task processing method applied to a multi-node system according to claim 1, characterized in that Before processing the to-be-consumed task in the following manner, the processing of the to-be-consumed task according to the obtained identification information of the to-be-consumed task further includes: Load the task entity binary data corresponding to the identification information of the to-be-consumed task from the cache corresponding to the multi-node system; Perform deserialization processing on the task entity binary data to obtain the to-be-consumed task.
7. The task processing method applied to a multi-node system according to claim 6, wherein The task entity binary data includes the original binary data corresponding to the network instruction; after each network instruction is successfully executed, serialize the each network instruction, and replace the original binary data corresponding to the each network instruction in the cache with the binary data obtained after serialization processing.
8. The task processing method applied to a multi-node system according to claim 6, wherein The method further includes: After loading the task entity binary data corresponding to the identification information of the to-be-consumed task from the cache corresponding to the multi-node system, persist the task entity binary data in the cache to the disk.
9. The task processing method applied to a multi-node system according to claim 7, wherein The network controller further includes a task placement client, and the message queue further contains the identification information of the third task, and the identification information of the third task is placed in the message queue by the task placement client.
10. A task processing device applied to a multi-node system, characterized in that, The multi-node system includes multiple task processing nodes, and a network controller is deployed on each task processing node in the multi-node system. The network controller includes a task orchestration management framework and a remote call client. The device includes: A detection unit, configured to detect the working status of the multiple task processing nodes; An acquisition unit, configured to acquire the identification information of the to-be-consumed task from the message queue. The message queue contains the identification information of the first task that was not completed by the first task processing node that had a downtime. The identification information of the first task was acquired and placed in the message queue by the second task processing node after detecting the downtime of the first task processing node; A processing unit, configured to process the to-be-consumed task according to the obtained identification information of the to-be-consumed task. The to-be-consumed task includes at least one network instruction, and the task orchestration management framework is configured to sequentially execute the network instructions included in the to-be-consumed task; The processing unit is configured as: Sequentially execute the network instructions included in the to-be-consumed task in the following manner: sequentially send the network instructions in the to-be-consumed task to the remote call client, so that the remote call client sequentially sends the network instructions to the network devices corresponding to the network instructions; If the execution of the network instructions included in the to-be-consumed task fails, perform rollback processing on the network instructions that have been executed in the to-be-consumed task to re-process the to-be-consumed task.
11. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the task processing method applied to a multi-node system according to any one of claims 1 to 9.
12. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the task processing method applied to a multi-node system as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Distributed scheduling system capable of transmitting data between nodes
CN110035103A
Task management method, device and system, computer storage medium and electronic equipment
CN111290854A