Services application unified data exchange mechanism and dynamic scheduling fault tolerance method and system
By applying a unified data exchange mechanism and dynamic scheduling fault tolerance method in service, real-time monitoring and response to service status is achieved, fault tolerance processing is triggered, and the stability and availability of the task scheduling system is solved, and the system's fault tolerance capabilities and task execution are improved.
Patent Information
- Application Number
- CN202510535437.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-25
AI Technical Summary
The existing task scheduling systems lack real-time monitoring and response mechanisms for service status, resulting in reduced system stability and insufficient fault tolerance. They lack rapid migration and recovery mechanisms in case of service failures, and low system availability.
The service-oriented application of unified data exchange mechanism and dynamic scheduling fault tolerance method is adopted. By starting the service scheduling module, task management module, configuration center module, status monitoring module and exception management module, real-time monitoring and response to service status is achieved, fault-tolerant processing or fault migration is triggered, the number of service instances is adjusted using dynamic scaling strategies, the circuit breaker mechanism is triggered to reduce the risk of service crash, and task events are processed asynchronously through the event bus module.
It improves the stability and fault tolerance of the system, ensures the stable operation of tasks, enhances the flexibility and availability of the system, and reduces the risk of service crashes.
Smart Images

Figure CN120371478A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of service scheduling, and in particular, to a unified data exchange mechanism for service-oriented applications, a dynamic scheduling fault tolerance method, and a system. Background Art
[0002] With the increasing complexity of computing tasks and the widespread application of distributed systems, traditional task scheduling and service management methods are difficult to meet the requirements of high availability, flexibility, and scalability. In the prior art, task scheduling systems usually lack a real-time monitoring and response mechanism for service status, making it difficult to handle service exceptions in a timely manner, resulting in a decline in system stability. In addition, the system has insufficient fault tolerance, and lacks a fast migration and recovery mechanism in case of service failures, resulting in low system availability. Summary of the Invention
[0003] Object of the Invention: To provide a unified data exchange mechanism for service-oriented applications, a dynamic scheduling fault tolerance method, and a system, so as to at least solve one of the problems existing in the above prior art.
[0004] Technical Solution: A unified data exchange mechanism for service-oriented applications and a dynamic scheduling fault tolerance method, including:
[0005] Start a service scheduling module, a task management module, a configuration center module, a communication management module, a status monitoring module, and an exception management module. The service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module, and receive external task creation, deletion, and configuration modification requests through the interface management module;
[0006] The task management module receives a task configuration file and parses the service dependencies therein, generates execution nodes of the task, and advances the execution of the task in a data flow manner;
[0007] The task management module distributes binary executable files and runtime configuration files required for the task to target device nodes according to the parsed task configuration, and starts the target service through the service scheduling module;
[0008] Through the event bus module of the service scheduling module, node call events, scheduling events, and exception events generated during the task execution process are submitted to an event processing queue, and are asynchronously processed by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, the task management module is called back to update the task status.
[0009] Preferably, the service scheduling module includes:
[0010] During the task execution, the service scheduling module determines the service status based on the heartbeat, resource utilization, and throughput data provided by the status monitoring module. When a service anomaly is detected, it triggers a fault tolerance process or a failover event, and redeploys the task through backup nodes or idle resources to ensure the stable operation of the task.
[0011] Preferably, the service scheduling module further includes:
[0012] During the task execution, according to the real-time load situation, it adjusts the number of service instances through a dynamic scaling policy, triggers a circuit breaker mechanism when the service has an anomaly or its performance degrades, pauses task scheduling, reduces the risk of service crashes, and updates the load balancing policy.
[0013] Preferably, the communication management module includes:
[0014] Inter-service communication adopts a name addressing mechanism, and the actual address is resolved by the local nameserver;
[0015] When communicating across nodes, it synchronizes the information of all network nodes through the hostserver to achieve service discovery and direct connection transmission;
[0016] Serializes data based on the Protobuf protocol, and supports the extraction of non-continuous fields according to the mapping table.
[0017] Preferably, the field mapping table includes:
[0018] Defines the correspondence between the input fields of the lower-level service and the output fields of the upper-level service;
[0019] Extracts specified fields from the full-scale data through Protobuf deserialization;
[0020] Dynamically updates the mapping relationship to adapt to configuration changes without restarting the service.
[0021] Preferably, the configuration center module includes:
[0022] Stores the runtime parameters of the service and supports hot updates;
[0023] When the parameters change, it triggers the restart of the service instance or the reconstruction of the call chain DAG;
[0024] Persistently stores the task execution history and event processing logs.
[0025] Preferably, the interface management module includes:
[0026] Receives external task creation, deletion, and parameter modification requests through the HTTP interface;
[0027] After verifying the requested permissions, distribute the instructions to the task management module or the configuration center for execution;
[0028] Return the task execution status and the resource occupancy report.
[0029] Preferably, the service scheduling module further includes: a service high availability module, and the service high availability module includes: a service fusing module; the service fusing module includes:
[0030] When the service error rate exceeds 5% or the latency is higher than 500 ms, trigger a fusing event;
[0031] Suspend distributing requests to the faulty service, release associated resources and record logs;
[0032] Probe the service status every 30 seconds. After recovery, re-enable it and update the call chain.
[0033] Preferably, the service high availability module includes: a load balancing module; the load balancing module includes:
[0034] Adopt the weighted round-robin algorithm and allocate weights according to the node computing power;
[0035] Generate a call sequence and distribute requests according to the sequence;
[0036] When the nodes are expanded or contracted, dynamically adjust the weights and regenerate the sequence.
[0037] Preferably, the task management module further includes: a service distribution module, and the service distribution module includes:
[0038] The master node queries the available service list from the daemon process of the target device node;
[0039] If the target node lacks the required service, transmit the program package and verify its integrity, and start an instance through the platform management service;
[0040] Bind the service instance to the specified device type according to the virtual mapping table, and preferably select a device with a load lower than 50%.
[0041] Preferably, the exception management module includes:
[0042] When the event bus receives a submitted exception event, obtain all subscribers of this event through the exception event type, and obtain the processing function corresponding to this event from the subscribers, and package it into a task form and submit it to the thread pool for processing
[0043] To achieve the above object, according to another aspect of the present application, a unified data exchange mechanism and dynamic scheduling and fault tolerance system for service-oriented applications are provided.
[0044] The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance system according to the present application includes:
[0045] A start receiving module, which is used to start a service scheduling module, a task management module, a configuration center module, a communication management module, a status monitoring module, and an exception management module. The service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module, and receive external task creation, deletion, and configuration modification requests through the interface management module;
[0046] A parsing and execution module, which is used for the task management module to receive a task configuration file and parse the service dependencies therein, generate execution nodes of the task, and advance the execution of the task in a data flow manner;
[0047] A task distribution and start module, which is used for the task management module to distribute binary executable files and runtime configuration files required for the task to target device nodes according to the parsed task configuration, and start the target service through the service scheduling module;
[0048] An event scheduling and management module, which is used to submit node call events, scheduling events, and exception events generated during the task execution process to an event processing queue through the event bus module of the service scheduling module, and perform asynchronous processing by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, the task management module is called back to update the task status.
[0049] To achieve the above object, according to another aspect of the present application, there is provided an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to any one of the present invention.
[0050] To achieve the above object, according to another aspect of the present application, there is provided a computer-readable storage medium, in which a computer instruction is stored, and the computer instruction is used to implement the service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to any one of the present invention when executed by a processor.
[0051] Beneficial effects: In the embodiments of the present application, by adopting the method of unified data exchange and dynamic scheduling fault tolerance, through starting the service scheduling module, task management module, configuration center module, communication management module, status monitoring module, and exception management module, the service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module, and receive external task creation, deletion, and configuration modification requests through the interface management module; the task management module receives the task configuration file and parses the service dependency relationships therein, generates the execution nodes of the task, and advances the execution of the task in a data flow manner; the task management module distributes the binary executable files and runtime configuration files required for the task to the target device nodes according to the parsed task configuration, and starts the target service through the service scheduling module; through the event bus module of the service scheduling module, the node call events, scheduling events, and exception events generated during the task execution process are submitted to the event processing queue, and are asynchronously processed by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, the task management module is called back to update the task status, achieving the purpose of ensuring the stable operation of the task, thereby realizing the technical effect of improving the system stability and fault tolerance ability, and further solving the technical problems that the current task scheduling system usually lacks a real-time monitoring and response mechanism for service status, is difficult to handle service exceptions in a timely manner, resulting in a decline in system stability; in addition, the system has insufficient fault tolerance ability, lacks a fast migration and recovery mechanism when a service fails, and the system availability is low. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 FIG. is a service model structure diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method of service-oriented applications according to the embodiments of the present application;
[0053] Figure 2 FIG. is a schematic diagram of the image classification task processing flow of the unified data exchange mechanism and dynamic scheduling fault tolerance method of service-oriented applications according to the embodiments of the present application;
[0054] Figure 3 FIG. is a schematic diagram of the service splitting definition of the unified data exchange mechanism and dynamic scheduling fault tolerance method of service-oriented applications according to the embodiments of the present application;
[0055] Figure 4 FIG. is a service scheduling timing diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method of service-oriented applications according to the embodiments of the present application;
[0056] Figure 5 FIG. is a service scheduling module architecture diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method of service-oriented applications according to the embodiments of the present application;
[0057] Figure 4It is a service scheduling time sequence diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0058] Figure 6 It is a general workflow diagram of task scheduling for the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0059] Figure 7 It is a schematic diagram of the service distribution and deployment process for the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0060] Figure 8 It is a schematic diagram of the composition of the event bus for the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0061] Figure 9 It is a message publishing and subscribing time sequence diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0062] Figure 10 It is a service fault tolerance function time sequence diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0063] Figure 11 It is a service expansion function time sequence diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0064] Figure 12 Schematic diagram of the service load balancing function of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0065] Figure 13 It is a service circuit breaker function time sequence diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0066] Figure 14 It is a schematic diagram of the service scheduling interface function of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0067] Figure 15 It is a schematic diagram of the detailed design of the function of the communication management module for the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0068] Figure 16 It is a configuration center processing time sequence diagram of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0069] Figure 17It is a flowchart of the task management function of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0070] Figure 18 It is a schematic diagram of the service instantiation process of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0071] Figure 19 It is a flowchart of the status monitoring service processing of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0072] Figure 20 It is a timing diagram of middleware exception handling of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0073] Figure 21 It is a flowchart of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application;
[0074] Figure 22 It is a schematic diagram of the structure of the unified data exchange mechanism and dynamic scheduling fault tolerance system for service-oriented applications according to an embodiment of the present application; and
[0075] Figure 23 It is a schematic diagram of the structure of an electronic device of the unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to an embodiment of the present application. Detailed implementation manners
[0076] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0077] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0078] In addition, the terms "installed", "set up", "provided with", "connected", "linked", and "socketed" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral structure; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or there can be internal communication between two devices, components, or constituent parts. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0079] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will describe this application in detail with reference to the drawings and in combination with the embodiments.
[0080] As Figure 21 shown, according to an embodiment of the present invention, a unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications are provided. The method includes the following steps S101 to step S104:
[0081] Step S101, start the service scheduling module, task management module, configuration center module, communication management module, status monitoring module, and exception management module. The service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module and receive external task creation, deletion, and configuration modification requests through the interface management module;
[0082] By adopting a modular structure, the system can be more flexible, can start and manage different functional modules according to requirements, and is convenient for later expansion and maintenance; efficient communication, through the communication management module, ensures efficient communication between modules and can quickly process external requests; flexible external interfaces, the system can accept external dynamic configuration and task management requests, enhancing the flexibility and scalability of the system.
[0083] Step S102, the task management module receives the task configuration file and parses the service dependencies therein, generates the execution nodes of the task, and advances the execution of the task in a data flow manner;
[0084] The task management module automatically generates task execution nodes according to the configuration, reducing manual intervention and improving the automation degree of task scheduling; service dependency analysis, parsing service dependencies can help the system better understand the order and priority of task execution and avoid dependency problems or task conflicts; data flow-driven, advancing task execution in a data flow manner makes the progress of the task clear and easy to track and manage.
[0085] Step S103: The task management module distributes the binary executable file and the runtime configuration file required for the task to the target device node according to the parsed task configuration, and starts the target service through the service scheduling module;
[0086] The task management module automatically processes the distribution of task files, avoiding manual operations and improving efficiency; unified control, the target service is uniformly started through the service scheduling module, making the system management more centralized and standardized; resource optimization, accurately distributing files according to the needs of the task to ensure the efficient use of resources and the smooth execution of tasks.
[0087] Step S104: Through the event bus module of the service scheduling module, submit the node call events, scheduling events, and exception events generated during the task execution process to the event processing queue, and perform asynchronous processing by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, call back the task management module to update the task status.
[0088] Through asynchronous processing, the system can process events without blocking the execution of the main task, improving the efficiency and responsiveness of task execution; event-driven mechanism, adopting the event bus mechanism can conveniently manage and track various events during the task execution process, enhancing the traceability of the system; flexible status update, updating the task status through callback to ensure real-time feedback on the task progress and status, improving the transparency and controllability of task management.
[0089] According to an embodiment of the present invention, preferably, the service scheduling module includes:
[0090] During the task execution process, the service scheduling module judges the service status according to the heartbeat, resource utilization rate, and throughput data provided by the status monitoring module, and when detecting service anomalies, triggers fault tolerance processing or fault migration events, and redeploys the task through backup nodes or idle resources to ensure the stable operation of the task.
[0091] According to an embodiment of the present invention, preferably, the service scheduling module further includes:
[0092] During the task execution process, according to the real-time load situation, adjust the number of service instances through the dynamic scaling policy, and trigger the circuit breaker mechanism when the service has anomalies or performance degradation, suspend task scheduling, reduce the risk of service crashes, and update the load balancing policy.
[0093] According to an embodiment of the present invention, preferably, the communication management module includes:
[0094] Name addressing mechanism is adopted for communication between services, and the actual address is resolved by the local nameserver;
[0095] During cross-node communication, the hostserver is used to synchronize the information of all nodes in the network to achieve service discovery and direct connection transmission;
[0096] Serialize data based on the Protobuf protocol, and support the extraction of discontinuous fields according to the mapping table.
[0097] According to an embodiment of the present invention, preferably, the field mapping table includes:
[0098] Define the correspondence between the input fields of the lower-level service and the output fields of the upper-level service;
[0099] Extract the specified fields from the full-scale data through Protobuf deserialization;
[0100] Dynamically update the mapping relationship to adapt to configuration changes without restarting the service.
[0101] According to an embodiment of the present invention, preferably, the configuration center module includes:
[0102] Store the runtime parameters of the service and support hot update;
[0103] When the parameters change, trigger the restart of the service instance or the reconstruction of the call chain DAG;
[0104] Persistently store the task execution history and event processing logs.
[0105] According to an embodiment of the present invention, preferably, the interface management module includes:
[0106] Receive external task creation, deletion, and parameter modification requests through the HTTP interface;
[0107] After verifying the request permissions, distribute the instructions to the task management module or the configuration center for execution;
[0108] Return the task execution status and resource occupancy report.
[0109] According to an embodiment of the present invention, preferably, the service scheduling module further includes: a service high availability module, and the service high availability module includes: a service fusing module; the service fusing module includes:
[0110] When the service error rate exceeds 5% or the latency is higher than 500ms, trigger a fusing event;
[0111] Suspend distributing requests to the faulty service, release the associated resources and record the logs;
[0112] Detect the service status every 30 seconds, and re-enable and update the call chain after recovery.
[0113] According to an embodiment of the present invention, preferably, the service high availability module includes: a load balancing module; the load balancing module includes:
[0114] The weighted round-robin algorithm is adopted to allocate weights according to the computing power of nodes;
[0115] Generate a call sequence and distribute requests according to the sequence;
[0116] When nodes are expanded or contracted, the weights are dynamically adjusted and the sequence is regenerated.
[0117] According to an embodiment of the present invention, preferably, the task management module further includes: a service distribution module, and the service distribution module includes:
[0118] The master node queries the available service list from the daemon process of the target device node;
[0119] If the required service is missing from the target node, transmit the program package and verify its integrity, and start an instance through the platform management service;
[0120] Bind the service instance to the specified device type according to the virtual mapping table, and preferably select a device with a load lower than 50%.
[0121] According to an embodiment of the present invention, preferably, the exception management module includes:
[0122] When the event bus receives a submitted exception event, obtain all subscribers of this event through the exception event type, obtain the processing function corresponding to this event from the subscribers, and package it into a task form and submit it to the thread pool for processing.
[0123] The working process of the present invention is as follows:
[0124] This application is based on a task abstraction model of services. The main function of service scheduling is to manage computing resources and task scheduling through scheduling between services. Among them, services and data streams are two basic abstractions in this abstraction model. The data stream is an abstraction of the task execution process, and its flow direction is guided by the service manager. A service is a collection composed of code and data, and is pushed by the data stream to execute.
[0125] As Figure 1 shown, it shows the service scheduling function structure based on the service model. The core service is the key component of the entire computing task abstraction. It mainly provides basic services such as communication mechanisms between services, management of execution data streams / services, interrupts and exceptions, and maintains key abstract data structures. A service is a static concept, has certain computing resources, and can complete a certain computing task, such as decoding and reasoning of intelligent computing tasks. Services communicate through a unified encapsulation interface and process messages under the push of the data stream.
[0126] Data flow is the dynamic part of the system and represents the flow of data through each process of a computing task. It is a continuous concept that does not depend on a fixed association with hardware resources. The role of data flow is to drive the service to complete task execution. The same data flow can be reused to achieve parallelism, and multiple data flows can flow simultaneously to achieve concurrency. Data flow provides flexibility and efficiency for task execution.
[0127] In addition, to achieve chaining between services, RPC can be used to implement data communication. RPC (Remote Procedure Call) is a protocol for inter-process communication. In RPC, the client application sends requests to the server application over the network, and the server application processes the requests and returns the results to the client.
[0128] Taking the common intelligent computing task of graphic classification as an example, as Figure 2 shown, the preprocessing part mainly includes operations such as "decoding, scaling, and cropping", and the backbone network is ResNet-50. The preprocessing part receives the image binary stream, generates the tensor data required by the backbone network, and then inputs it into the backbone network to predict the classification result of the output picture. When the inference volume is large, the CPU will become the bottleneck of the overall system processing, and the inference service cannot fully utilize the computing performance of the relevant hardware. In this scenario, this module classifies and divides services according to the hardware resources used, that is, the preprocessing service is one service, and the inference service is one service. The flowchart of the task processing after splitting is as Figure 3 shown.
[0129] After splitting, the preprocessing service is separately deployed to the CPU node, and the inference service is deployed to the NPU node, so that the CPU preprocessing service can be horizontally scaled infinitely, meet the data supply for NPU processing, and fully utilize the performance of the NPU. At the same time, by decoupling the CPU and NPU computing tasks, the waiting time for data exchange between the CPU and NPU is reduced, thereby reducing the overall task processing time.
[0130] Considering the operation efficiency and flexibility during the operation of computing tasks, the scheduling module needs to have the function of task scheduling strategy. Compared with the traditional resource-sharing task parallelism, first of all, this method defines a task through the DAG (Directed Acyclic Graph) method, where the nodes in the DAG represent services, and the upper and lower relationships of the nodes in the task DAG are the scheduling order of each service. These services are scheduled in a serial manner according to the established order during the scheduling process to ensure the correctness of the data processing flow. The service scheduling time sequence is as Figure 4 shown.
[0131] Furthermore, the scheduler can create multiple task DAGs to achieve parallel computing tasks. Different services among these tasks work in parallel in the form of task pipelines. The scheduler balances the execution time of each service through certain strategies, minimizes the drain time during the task pipeline process as much as possible, achieves the optimal efficiency of the task pipeline, and thereby improves the utilization rate of overall computing resources.
[0132] As Figure 5 shown, the service scheduling module architecture is to execute tasks and manage tasks by parsing the configuration file. The scheduling module includes several functional modules such as service scheduling, task management, configuration center, communication bus, and exception management.
[0133] The role of the service scheduling module is to ensure the stability of task operation through scheduling methods such as service fault tolerance, fault migration, dynamic scaling, and service circuit breaking.
[0134] The role of the status monitoring module is to receive data such as heartbeats, resource utilization rates, and throughputs reported by services and use this as a basis to judge the status of services and trigger corresponding scheduling events.
[0135] The role of the communication management module is to provide support for data communication between services with the help of communication middleware.
[0136] The role of the task management module is to parse task configurations and advance task execution in the form of data streams. During the task execution process, corresponding node call events will be generated for each frame of data and submitted to the event bus.
[0137] The task management module is also responsible for service distribution, distributing the binary executable file corresponding to the task and the runtime configuration file to the specified device nodes and starting them.
[0138] The role of the configuration center is to store service configurations in the form of key-value pairs or files, and at the same time has the function of updating service parameters during the task process.
[0139] The event bus processes events generated during the entire scheduling process through a dynamically scalable thread pool. The types of events included are exception / error events, node call events, service scheduling events, etc. At the same time, for node call events with a very high generation rate, the event bus can also perform traffic control on them through the thread pool to prevent service crashes caused by the backlog of call requests.
[0140] The interface management module receives external task creation / deletion and parameter modification requests through the http method, enabling the framework to have the functions of runtime configuration modification and task control.
[0141] The exception management module processes exceptions by subscribing to exception events.
[0142] Specifically, the computing task scheduling module includes the following details:
[0143] I. Function description
[0144] As Figure 6 shown, at the operating system level, after the system starts, the task management module completes the startup of each sub-module and the initialization of exception management via the initialization interface. At the same time, the middleware completes the initialization operations of service scheduling and task call chain by loading the runtime configuration file, and performs service distribution, and then starts the task running. After completing the initialization work, the middleware can receive requests such as external task creation, deletion, scheduling configuration modification, and service parameter modification through the API of the task management module.
[0145] As Figure 7 shown, for the service distribution and deployment process, all device nodes in the system except the master node contain a service deployment daemon process that runs when the system starts. When the task management module performs service distribution, the master node first obtains the list of available services already distributed on the target node from this daemon process and compares it with the services required by the current task. For services that the task needs but are not available in the device node, the master node will distribute the complete service program package. After the service distribution is completed, the task management module will continue to distribute all service configurations and complete the startup of the service through the device node daemon process.
[0146] II. Event bus
[0147] The event bus is the core function of service scheduling and an important tool for promoting data processing between services, service high availability, and service governance. As Figure 8 shown, it is a schematic diagram of the event bus composition. The event bus is used to handle exceptions / errors / scheduling / node call events generated during task execution and service scheduling. At the same time, the event bus maintains a dynamically scalable thread pool to process various events and performs traffic control on node call events to prevent service crashes caused by excessive request backlogs.
[0148] As Figure 9As shown in the figure, for the event bus publish-subscribe process, the event bus provides an event handling base class. Users can build event bus subscribers by inheriting the event handling base class and passing in the subscribed event type through a template. Subscribers will generate virtual function interfaces for event handling through a generic template. Subscribers must write corresponding event handling functions for all subscribed events. Event subscription can be completed by adding subscribers to the event bus. After successful subscription, the subscription information is stored in the relevant container of the event bus in the form of key-value pairs. When the publisher publishes an event through the event bus event submission interface, the event bus will first query and find the corresponding subscriber from the internal container, and wrap the event handling function corresponding to the subscriber for the event into a task and submit it to the thread pool attached to the event bus for asynchronous processing. The thread pool contains several event handling threads, which loop to take out the tasks to be processed from the queue and process them. When the thread pool takes out all the processing tasks generated by the event and finishes executing, it means that the current event processing is completed.
[0149] III. Service High Availability
[0150] The purpose of service high availability is to ensure the stability of task operation. The functions of the service high availability module include service fault tolerance, load balancing, fault migration, dynamic scaling, service circuit breaker, etc.
[0151] Service Fault Tolerance
[0152] As Figure 10 shown in the figure, for the service fault tolerance processing process, when the status monitoring module detects that the heartbeat of a certain service times out, it will submit a heartbeat timeout event to the event bus, and the event bus will notify the corresponding subscriber in the scheduler to trigger the execution of the service fault tolerance logic. If there is a fault-tolerant service on the node corresponding to the service and the fault tolerance is successful, the scheduling system will continue to trigger the load balancing event, execute the load balancing algorithm to generate a new load balancing call sequence, and update it to the call chain of the task management module, thus completing a complete service fault tolerance scheduling.
[0153] When performing service fault tolerance, the system will give priority to using the computing resources of the current node. When there is no node available for fault tolerance on the node corresponding to the abnormal service, it will further trigger the fault migration event and execute the fault migration scheduling logic. After successful fault migration, it will also further trigger the load balancing event and update the load balancing call sequence to the call chain of the task management module.
[0154] Dynamic Scaling
[0155] Dynamic scaling can respond to changes in system load by automatically increasing or decreasing service nodes, thereby achieving better performance and scalability. When the dynamic scaling function is turned off, the number of parallel services for a single node is a fixed value, specified by the user or automatically generated at compile time; when the dynamic scaling function is turned on, the number of parallel services for a single node is a numerical range, and the dynamic scaling trigger condition is specified by the user. The dynamic scaling process is as Figure 11 shown.
[0156] Load balancing
[0157] The processing capacity of a service is closely related to the hardware performance of the node where it is located. To evenly distribute requests to each service as much as possible, the weighted round-robin algorithm is used to achieve service load balancing. The principle of the weighted round-robin algorithm is: according to the different processing capabilities of nodes, different weights are assigned to each service so that it can accept service requests corresponding to the number of weights.
[0158] For example, there are three services, a, b, and c, with weights of 1, 2, and 4 respectively. The result of the weighted round-robin algorithm is to generate a service call sequence. Whenever a request arrives, the next service in the sequence is taken out in turn to process the request. For the above three services, the weighted round-robin algorithm will generate the sequence {c, c, b, c, a, b, c}. In this way, for every 7 client requests received, 1 of them will be forwarded to service a, 2 of them will be forwarded to service b, and 4 of them will be forwarded to service c. For the 8th request received, polling starts again from the head of the sequence.
[0159] As Figure 12 shown, the service scheduling module generates a service weight sequence according to the task configuration information, then uses the weighted round-robin algorithm to calculate the corresponding service call sequence, and then distributes requests according to this sequence. When the service changes, the service scheduler repeats this process to generate a new service call sequence and update it to the service executor.
[0160] Service fusing
[0161] As Figure 13 shown, the service fusing business logic diagram. First, when the status monitoring module detects that the service fusing trigger condition is met, it will trigger a service fusing event. At the same time as the event is triggered, the enable of the fusing trigger corresponding to this task will be turned off to prevent the event from being triggered repeatedly. After the event bus receives the fusing event, it will notify the service scheduling module to execute the corresponding service fusing logic. The service scheduling module stops this task through the execution management module. After the task stops successfully, the service scheduling module will re-enable the fusing trigger, thus completing a complete service fusing event trigger.
[0162] Interface management
[0163] AsFigure 14 As shown in Figure 14 , the interface management module serves as the entry point of the framework, providing external control request reception and execution distribution functions. For example, task start and stop control instructions are distributed to the task management module, and service parameter update requests are distributed to the configuration center.
[0164] Communication Management
[0165] The communication management module solves the problem of communication between services through name addressing. Each service can have its own name; a service called nameserver runs on each node, responsible for assigning addresses to services, managing the mapping between service names and addresses, resolving service names, and publishing service addresses. The name server is similar to DNS on the Internet. To support service name addressing, the communication management module adds a format in addition to the two types of URLs as the name address, with the format svc: / / service name.
[0166] The name address is a virtual address. No matter where the service is located, as long as its name address remains unchanged, a connection can be established with it through this address. If a service calls bind() to bind the name address (the address starting with svc: / / ), the nameserver will assign an actual address (the address starting with tcp: / / or file: / / ) to it and register the name and address in the mapping table. If another service connects to the name address, the nameserver will look up the actual address of the service based on the name and select the most suitable actual address to publish to the connection initiator. Other services establish a point-to-point direct connection with the target service through this address to conduct data transmission. The process of establishing a connection between services using the name address with the assistance of the nameserver.
[0167] Since the address of the nameserver on a single node is fixed, after the service starts, it will automatically connect to the nameserver to register (server) or resolve (client) names. In the case of multiple nodes, each node runs its own nameserver, responsible for its own name service. To support cross-node name resolution, a service is needed to manage all nodes in the system, synchronize node information to all nameservers, and then these nameservers can establish connections and work together to complete the name service within the entire network. This service is the hostserver.
[0168] The working principle of the host server is as follows: There is one host server running throughout the network, which is located in the communication management module. The name servers of all nodes are connected to the host server and register the nodes they are located in with it. The host server maintains a host list containing the IP addresses of each node and synchronizes this list to all the name servers in the network. The name servers establish connections with the name servers on all nodes based on this list.
[0169] Once the name servers on all nodes have established pairwise connections, cross-node service name resolution and service online notification can be completed through an internal protocol. For example, when a client on a host requests the address corresponding to the service name resolution from the local name server, the local name server can broadcast this request to all connected name servers to search for the service within the entire network.
[0170] For data that is not directly transmitted between services, the communication management module designs a full-data address caching function. As Figure 15 shown, the communication management module holds the full-data address of each frame in the task process. When the data types of the input of the lower-level service and the output of the upper-level service are the same, the corresponding input data can be directly obtained through the data address. When the input of the lower-level service is composed of some field combinations of all the data of the upper-level services, a field mapping table needs to be generated during the task configuration stage to assist in obtaining the input data of this service. The field mapping table stores the mapping relationship between the input Message fields of the current service and the existing Message fields of the upper-level service. All services internally hold general serialization and deserialization methods implemented based on the protobuf reflection principle. By combining this method with the field mapping table, the required data can be requested from the upper-level service and filled into the input Message of the current service, thus completing the acquisition of the input data.
[0171] Configuration Center
[0172] As Figure 16 shown, the configuration center loads and stores service configurations in the form of key-value pairs and registers configuration change events with the communication middleware. After the service starts, it subscribes to configuration change events by default. When the configuration parameters corresponding to this service change, the corresponding service will receive the changed parameters and trigger the parameter update function.
[0173] Task Management
[0174] As Figure 17The figure shows the workflow of the task management module. As a middleware entry module, the task management module has the ability to schedule tasks and reconstruct application functions according to task requirements, and can realize unified scheduling of tasks on all nodes.
[0175] The task management module initialization interface has two functions: one is to start other independently running sub-modules, such as the event bus, configuration center, status monitoring, etc.; the other is to read the task configuration to be run from the local and start it according to a fixed process.
[0176] The task startup process of the task management module is as follows: First, the task management module reads the task information from the task storage directory, then reads the parameter configuration of the service corresponding to the task and stores it in the configuration center. After completing the storage of the service configuration parameters, the task management module will read the scheduling parameters from the scheduling parameter file corresponding to the task and initialize the scheduling configuration. For the service scheduling configuration process, please refer to the service scheduling module section. After completing the service scheduling configuration initialization, the task management module will create a task call chain according to the service call sequence contained in the task, and store the task ID and its corresponding task call chain DAG graph in the key-value pair container that comes with the task management module. Finally, when the above preparations are completed, the task management module will start the relevant services and start the task running. If there are still tasks that have not been started in the task configuration file, the task management module will repeat the above steps until all tasks are started.
[0177] In addition, the task management module also has a series of task management interfaces, which are exposed to the outside through the interface management module, making it convenient for external execution of task creation, deletion, service parameter modification, scheduling parameter modification and other operations, greatly improving the flexibility of task operation.
[0178] Compute resource binding
[0179] like Figure 18 As shown in the service instantiation flow chart, when the task construction tool builds a task, it will generate a virtual mapping table for each service. This virtual mapping table only represents the device type and architecture to which the current service can be distributed, and is not bound to the actual device. Therefore, before the actual service distribution, it is necessary to generate a specified number of instances for each service and bind them to the real devices of the current heterogeneous computing system.
[0180] When the task scheduling module adds and runs this task, it will first perform device binding with the help of the device node list. The device node list is a static list, which is manually written and added by the user when the task scheduling module is deployed to a certain heterogeneous intelligent computing platform. Each device node will default to start with the daemon process and start the platform management plane service composed of fixed fields of the node name, such as svc: / / host1_service_deployment. The task scheduling module will perform instance binding according to the established binding policy and the parallelism of the current service. After the instance is bound to the device, the task scheduling module will deploy and start the corresponding instance on the specified device node with the help of the platform management plane service. When all instances are started successfully, it means that the service deployment is successful.
[0181] Status Monitoring
[0182] The status monitoring module has the ability to monitor the status of computing nodes and the status of services. It can detect whether there are abnormal working nodes by collecting information such as the working status and running status reported by computing nodes, and record the log information of the corresponding nodes.
[0183] As Figure 19 shown in the working flow chart of the status monitoring module, first the task management module will start the reporting data listening of the status monitoring module to receive information such as heartbeat data, working status, running status, resource utilization rate, and power consumption reported by computing nodes; then the service scheduling module will add triggers associated with the reported data to the status monitoring module to trigger corresponding scheduling events when the service status meets the conditions; after the trigger settings are completed, the status monitoring module will use the trigger to detect the status information every time it receives the reported data. If the trigger condition is met, it will submit the corresponding event to the event bus.
[0184] The status monitoring module receives the monitoring data pushed by the service through OTel (OpenTelemetry), which can be used to collect data from the application. OTel (OpenTelemetry) is a set of tools, APIs, and SDKs that can be used to instrument, generate, collect, and export telemetry data (metrics, logs, and traces) to help analyze the performance and behavior of the application.
[0185] The status monitoring module receives the metric data pushed by the service through http. When the metric data of a certain service is updated, the status monitoring module will update the values of all these metrics in the metric storage container and check the metric values with its corresponding trigger function. If the condition is met, the trigger will submit the corresponding scheduling event to the event bus, thus triggering the scheduler to schedule the service.
[0186] As shown in Table 1, the Metrics of OpenTelemetry have four basic metrics. The status monitoring module mainly collects status through the metrics of a single numerical type in OTel.
[0187] Table 1 OpenTelemetry Basic Metric Types
[0188] Index type OpenTelemetry field name Single value Gauge Counter Sum Histogram Histogram Sampling Summary
[0189] The descriptions of the four metrics are as follows:
[0190] Single Value: A measurement of a simple numerical value that can be simply incremented or decremented.
[0191] Counter: An accumulated measurement that can distinguish values under different labels by adding several labels.
[0192] Histogram: Accumulates the intervals of the interface obtained by observing observer(). Therefore, it is necessary to pass in buckets to indicate the intervals to be divided. For example, if the bucket values passed in are the three numbers {1, 10, 100}, then we can obtain the cumulative values of the data distributed in the four intervals {0 - 1, 0 - 10, 0 - 100, 0 - +Inf}, and at the same time, we can also obtain the overall total and the number of data.
[0193] Sampling: Similar to the histogram, but the passed-in values are quantiles, such as {0.5, 0.9} and their precisions, which are mainly used in scenarios where the specific numerical values of the data distribution are not known in advance (in contrast, the histogram requires passing in specific numerical values).
[0194] Exception Management
[0195] As Figure 20 shown, each module in the framework submits the exceptions that occur during the running process to the event bus in the form of exception events, and triggers the exception management module to handle the exception events. After receiving the submitted exception events, the event bus obtains all subscribers of this event through the exception event type, and obtains the processing function corresponding to this event from the subscribers, and packages it into a task form and submits it to the thread pool for processing.
[0196] From the above description, it can be seen that the present application achieves the following technical effects:
[0197] In the embodiments of the present application, a unified data exchange and dynamic scheduling fault tolerance method is adopted. By starting the service scheduling module, task management module, configuration center module, communication management module, status monitoring module, and exception management module, the service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module, and receive external task creation, deletion, and configuration modification requests through the interface management module; the task management module receives the task configuration file and parses the service dependency relationship therein, generates the execution nodes of the task, and promotes the execution of the task in a data flow manner; the task management module distributes the binary executable file and runtime configuration file required for the task to the target device nodes according to the parsed task configuration, and starts the target service through the service scheduling module; through the event bus module of the service scheduling module, the node call events, scheduling events, and exception events generated during the task execution process are submitted to the event processing queue, and are asynchronously processed by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, the task management module is called back to update the task status, achieving the purpose of ensuring the stable operation of the task, thereby realizing the technical effect of improving the system stability and fault tolerance ability, and further solving the technical problems that the current task scheduling system usually lacks a real-time monitoring and response mechanism for service status, is difficult to handle service exceptions in a timely manner, resulting in a decrease in system stability; in addition, the system has insufficient fault tolerance ability, lacks a fast migration and recovery mechanism when a service fails, and the system availability is low.
[0198] As Figure 22 shown, to achieve the above object, according to another aspect of the present application, a unified data exchange mechanism and dynamic scheduling fault tolerance system for service-oriented applications is provided. The unified data exchange mechanism and dynamic scheduling fault tolerance system for service-oriented applications includes:
[0199] A start receiving module 2201, configured to start the service scheduling module, task management module, configuration center module, communication management module, status monitoring module, and exception management module. The service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module, and receive external task creation, deletion, and configuration modification requests through the interface management module;
[0200] By adopting a modular structure, the system can be more flexible, can start and manage different functional modules according to requirements, and is convenient for later expansion and maintenance; efficient communication, through the communication management module, ensures efficient communication between modules and can quickly process external requests; flexible external interfaces, the system can accept external dynamic configuration and task management requests, enhancing the flexibility and scalability of the system.
[0201] The parsing and execution module 2202 is configured to receive a task configuration file by the task management module, parse the service dependencies therein, generate execution nodes for the task, and advance the execution of the task in a data flow manner;
[0202] The task management module automatically generates task execution nodes according to the configuration, reducing manual intervention and improving the automation level of task scheduling; service dependency analysis, parsing service dependencies can help the system better understand the order and priority of task execution, avoiding dependency issues or task conflicts; data flow driven, advancing task execution in a data flow manner makes the progress of the task clear and easy to track and manage.
[0203] The task distribution and startup module 2203 is configured to distribute binary executable files and runtime configuration files required for the task to target device nodes by the task management module according to the parsed task configuration, and start the target service through the service scheduling module;
[0204] The task management module automatically processes the distribution of task files, avoiding manual operations and improving efficiency; unified control, starting the target service uniformly through the service scheduling module makes the management of the system more centralized and standardized; resource optimization, precisely distributing files according to the needs of the task ensures the efficient use of resources and the smooth execution of tasks.
[0205] The event scheduling and management module 2204 is configured to submit node call events, scheduling events, and exception events generated during the task execution process to an event processing queue through the event bus module of the service scheduling module, and perform asynchronous processing by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, the task management module is called back to update the task status.
[0206] Through asynchronous processing, the system can process events without blocking the execution of the main task, improving the efficiency and responsiveness of task execution; event-driven mechanism, adopting the event bus mechanism can conveniently manage and track various events during the task execution process, enhancing the traceability of the system; flexible status update, updating the task status by callback ensures real-time feedback on the task progress and status, improving the transparency and controllability of task management.
[0207] From the above description, it can be seen that the present application achieves the following technical effects:
[0208] In the embodiment of the present application, a unified data exchange and dynamic scheduling fault tolerance method is adopted. By starting a service scheduling module, a task management module, a configuration center module, a communication management module, a status monitoring module, and an exception management module, the service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module, and receive external task creation, deletion, and configuration modification requests through the interface management module; the task management module receives a task configuration file and parses the service dependency relationship therein, generates execution nodes of the task, and promotes the execution of the task in a data flow manner; the task management module distributes the binary executable file and the runtime configuration file required for the task to the target device nodes according to the parsed task configuration, and starts the target service through the service scheduling module; through the event bus module of the service scheduling module, node call events, scheduling events, and exception events generated during the task execution process are submitted to the event processing queue, and are asynchronously processed by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, the task management module is called back to update the task status, achieving the purpose of ensuring the stable operation of the task, thereby realizing the technical effect of improving the system stability and fault tolerance ability, and further solving the technical problems that the current task scheduling system usually lacks a real-time monitoring and response mechanism for service status, is difficult to handle service exceptions in a timely manner, resulting in a decline in system stability; in addition, the system has insufficient fault tolerance ability, lacks a fast migration and recovery mechanism when a service fails, and the system availability is low.
[0209] As Figure 23 shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0210] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0211] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the service-oriented application unified data exchange mechanism and the dynamic scheduling fault tolerance method.
[0212] In some embodiments, the service-oriented application unified data exchange mechanism and the dynamic scheduling fault tolerance method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the service-oriented application unified data exchange mechanism and the dynamic scheduling fault tolerance method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the service-oriented application unified data exchange mechanism and the dynamic scheduling fault tolerance method by any other suitable means (e.g., by means of firmware).
[0213] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0214] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0215] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0216] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0217] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0218] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0219] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0220] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. Unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications, characterized in that, Including: A service scheduling module, a task management module, a configuration center module, a communication management module, a status monitoring module, and an exception management module. The service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module, and receive external task creation, deletion, and configuration modification requests through the interface management module; The task management module receives a task configuration file and parses the service dependencies therein, generates execution nodes of the task, and advances the execution of the task in a data flow manner; The task management module distributes binary executable files and runtime configuration files required for the task to target device nodes according to the parsed task configuration, and starts the target service through the service scheduling module; Through the event bus module of the service scheduling module, node call events, scheduling events, and exception events generated during the task execution process are submitted to an event processing queue, and are asynchronously processed by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, the task management module is called back to update the task status.
2. The unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to claim 1, characterized in that, The service scheduling module includes: During the task execution process, the service scheduling module judges the service status according to the heartbeat, resource utilization rate, and throughput data provided by the status monitoring module. When a service exception is detected, it triggers a fault tolerance processing or a failover event, and redeploys the task through backup nodes or idle resources to ensure the stable operation of the task.
3. The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to claim 1, characterized in that The service scheduling module further includes: During the task execution process, according to the real-time load situation, adjusts the number of service instances through a dynamic scaling policy, and triggers a circuit breaker mechanism when the service has an exception or its performance degrades, pauses task scheduling, reduces the risk of service crashes, and updates the load balancing policy.
4. The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to claim 1, wherein The communication management module includes: Service-to-service communication adopts a name addressing mechanism, and the actual address is resolved by the local nameserver; During cross-node communication, synchronize the information of all network nodes through the hostserver to achieve service discovery and direct connection transmission; Serialize data based on the Protobuf protocol, and support the extraction of discontinuous fields according to the mapping table.
5. The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to claim 4, characterized in that, The field mapping table includes: Define the correspondence between the input fields of the lower-level service and the output fields of the upper-level service; Extract specified fields from the full amount of data through Protobuf deserialization; Dynamically update the mapping relationship to adapt to configuration changes without restarting the service.
6. The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to claim 1, wherein The configuration center module includes: Store the runtime parameters of the service and support hot updates; When the parameters change, trigger the restart of the service instance or the reconstruction of the call chain DAG; Persistently store the task execution history and event processing logs.
7. The unified data exchange mechanism and dynamic scheduling fault tolerance method for service-oriented applications according to claim 1, characterized in that The interface management module includes: Receive external task creation, deletion, and parameter modification requests through the HTTP interface; After verifying the request permissions, distribute the instructions to the task management module or the configuration center for execution; Return the task execution status and resource occupancy report.
8. The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to claim 1, characterized in that The service scheduling module further includes: a service high availability module, and the service high availability module includes: a service circuit breaker module; the service circuit breaker module includes: When the service error rate exceeds 5% or the latency is higher than 500 ms, a circuit breaker event is triggered; Suspend distributing requests to the faulty service, release associated resources, and record logs; Probe the service status every 30 seconds. After recovery, re-enable and update the call chain.
9. The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to claim 8, characterized in that The service high-availability module includes: a load balancing module; the load balancing module includes: Adopt the weighted round-robin algorithm and allocate weights according to node computing power; Generate a call sequence and distribute requests according to the sequence; When nodes are scaled out or in, dynamically adjust the weights and regenerate the sequence.
10. The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to claim 1, characterized in that, The task management module further includes: a service distribution module, and the service distribution module includes: The master node queries the available service list from the daemon process of the target device node; If the required service is missing from the target node, transmit the program package, verify its integrity, and start an instance through the platform management service; Bind the service instance to the specified device type according to the virtual mapping table, and preferentially select a device with a load lower than 50%.
11. The service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to claim 1, characterized in that The exception management module includes: When the event bus receives a submitted exception event, obtain all subscribers of this event through the exception event type, obtain the processing function corresponding to this event from the subscribers, and package it into a task form and submit it to the thread pool for processing.
12. Service-oriented application unified data exchange mechanism and dynamic scheduling fault-tolerant system, characterized in that, Includes: A startup receiving module, used to start the service scheduling module, task management module, configuration center module, communication management module, status monitoring module, and exception management module. The service scheduling module, the task management module, the configuration center module, the status monitoring module, the exception management module, and the interface management module are communicatively connected through the communication management module, and receive external task creation, deletion, and configuration modification requests through the interface management module; A parsing and execution module, used for the task management module to receive a task configuration file, parse the service dependencies therein, generate the execution nodes of the task, and promote the execution of the task in a data flow manner; A task distribution and startup module, used for the task management module to distribute the binary executable file and runtime configuration file required for the task to the target device node according to the parsed task configuration, and start the target service through the service scheduling module; An event scheduling and management module, used to submit node call events, scheduling events, and exception events generated during the task execution process to the event processing queue through the event bus module of the service scheduling module, and perform asynchronous processing by the dynamic thread pool of the event bus module. After the asynchronous processing is completed, call back the task management module to update the task status.
13. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the service-oriented application unified data exchange mechanism and dynamic scheduling fault tolerance method according to any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the unified data exchange mechanism and dynamic scheduling fault tolerance method of the service-oriented application according to any one of claims 1 to 11 when executed by a processor.
Citation Information
Patent Citations
Unitized distributed scheduling system and method based on DAG
CN112379995A
Task scheduling method and device, computer equipment, storage medium and program product
CN115016915A
Method for dynamically calling multiple models for deep learning algorithm and service architecture
CN119645607A
A dynamic, event-driven API orchestration system for cloud-native solutions
DE202025101075U1
Cloud safety computing method, device and storage medium based on cloud fault-tolerant technology
US20230350709A1
Cited By
Multi-card cluster heterogeneous scheduling method and system based on event bus
CN121858233A