A distributed cluster, multi-node task scheduling method, device and storage medium
Through RMI, the remote method call design for task scheduling is solved, and the existing system is difficult to dynamically allocate tasks in simple scenarios is realized, and lightweight and flexible distributed task scheduling is achieved, with low resource utilization and cross-platform support.
Patent Information
- Application Number
- CN202410965326.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-07-18
AI Technical Summary
The existing distributed task scheduling system is difficult to dynamically allocate tasks in simple scenarios, cannot flexibly expand according to node load and task requirements, and has a large resource consumption.
The RMI is used to build a remote method call design for task scheduling, including building RMI server, client, task queue and intermediate state storage, storing task data through a relational database, and real-time, dynamic, multi-node acquisition and processing of task data through the RMI remote interface.
It realizes lightweight distributed task scheduling, has low resource usage, supports cross-platform, and is easy to expand nodes. It can dynamically adjust task processing capabilities according to application scenarios.
Smart Images

Figure CN118860688B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of task scheduling technology, and in particular to a distributed cluster, multi-node task scheduling method, device and storage medium. Background Art
[0002] With the development of network technology and the increase of data processing demand, distributed computing systems are becoming more and more important. In distributed systems, task scheduling is a key issue, which requires allocating tasks to various nodes in order to use computing resources efficiently.
[0003] At present, there are some Java-based task scheduling systems, such as Apache Hadoop and Apache Spark. These systems can process large amounts of data in a distributed environment, but they usually rely on specific frameworks and libraries. For some simple tasks or specific application scenarios, they may be too complex, with high development costs and large resource consumption. At the same time, in simple scenarios, it is often impossible to dynamically allocate tasks according to the load of the nodes and the requirements of the tasks. In terms of the number of nodes, they usually only support a fixed number and cannot be easily and quickly expanded to more nodes. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a distributed cluster, multi-node task scheduling method, device and storage medium.
[0005] In a first aspect, the present invention provides a distributed cluster, multi-node task scheduling method, comprising:
[0006] The design of task scheduling remote method call based on RMI includes: constructing an RMI server and an RMI client, wherein the RMI client includes a task data extraction node; constructing task queues tasks and intermediate state waiting of the RMI server, which are respectively used to store the intermediate state information of task queue data and task processing data; constructing an RMI remote interface for processing task queue tasks and intermediate state waiting, including: a data push interface for pushing data to task queue tasks, a data extraction interface for extracting data from task queue tasks, a tasks counting interface for calculating the number of tasks in the entire task queue, a waiting counting interface for calculating the number of waiting in the entire intermediate state, a first waiting deletion interface for deleting the corresponding intermediate state information according to the key of the intermediate state information in the intermediate state waiting, a waiting clearing interface for clearing all intermediate state information in the intermediate state waiting, and a second waiting deletion interface for deleting the corresponding intermediate state information in the intermediate state waiting according to the storage time;
[0007] According to the application scenario type, the task data generated by different application scenarios are stored in a relational database, where the fields of the task data include: id: unique identifier of the task; status: task processing status; createDate: task generation time; custom task data attributes adapted to the application scenario;
[0008] The task data is extracted into the task queue tasks and the intermediate state waiting through the task data extraction node; each RMI client as a task processing node obtains the task data in real time, dynamically and multi-node by calling the RMI remote interface, and after the task data is obtained, the task is processed according to the actual application scenario and the task state is written back; the task data extraction node periodically clears the intermediate state waiting intermediate state information according to the preset parameter t; and the task data is pushed and processed in a loop.
[0009] Furthermore, a linked queue is used to create task queues tasks, which are used to store task queue data. A linked queue is a blocking queue based on a linked list. The internal part of the linked queue is implemented by a single linked list. Elements can only be taken from the head of the queue and added to the tail of the queue. The linked queue uses a lock separation technology containing two locks to achieve non-blocking entry and exit of the queue. Adding elements and getting elements have independent locks, realizing read-write separation of the linked queue and supporting parallel execution of read and write operations. A synchronous hash map is used to create an intermediate state waiting, which is used to store the intermediate state information of task processing data.
[0010] Furthermore, the RMI client as a task processing node configures a multi-threaded processing mechanism.
[0011] Furthermore, the task data is extracted into the task queue tasks and the intermediate state waiting through the task data extraction node, including: sorting the task data in the relational database in ascending order according to the task generation time;
[0012] Select the task data to be extracted based on the task data status that is not pushed.
[0013] Establish a link instance, and obtain the task data to be extracted quantitatively and regularly in a loop. First, format the task data to be extracted into a JSON string, and then use the link instance to call the data push interface to push the task data to the task queue tasks. At the same time, store the unique identification id and storage time of the current task data to be processed in the intermediate state waiting.
[0014] Furthermore, the process of extracting task data to task queues tasks and intermediate state waiting is controlled according to task queues tasks and intermediate state waiting, including:
[0015] Set the maximum data volume threshold Tmax that the task queue tasks can carry;
[0016] The data push interface obtains the length of the current task queue tasks through the tasks counting interface. The data push interface determines whether to store the pending task in the task queue tasks based on whether the length of the current task queue tasks is greater than the maximum data volume threshold Tmax that can be carried. If the length of the current task queue tasks is less than the maximum data volume threshold Tmax that can be carried, and it is not in the task queue task, the pending task is stored at the end of the task queue tasks. At the same time, the unique identification id and storage time of the current pending task data are stored in the intermediate state data waiting. If it is greater than or equal to the maximum data volume threshold Tmax that can be carried, or it is in the task queue task, the current task data will not enter the pending task queue tasks, and wait for the next batch of data to be verified again.
[0017] Furthermore, according to the ID of the pending task data in the intermediate state waiting, it is determined whether the currently pushed pending task data already exists in the task queue tasks. If it already exists, it will not be pushed to the task queue tasks again. If it does not exist, the pending task data will be pushed to the task queue tasks.
[0018] Furthermore, each RMI client as a task processing node obtains task data in real time, dynamically, and through multiple nodes by calling the RMI remote interface. After the task data is obtained, the task is processed according to the actual application scenario and the task status is written back, including:
[0019] Start the thread pool and independently start an infinite loop monitoring thread. The monitoring thread executes the actual task push logic through try, catch and finally methods. In the finally method, the thread sleep time is set to limit the interval time of each loop, so as to control the frequency of the current task processing node obtaining task data. In the try method, the current number of active threads in the thread pool is obtained and compared with the number of threads in the thread pool. If the current number of active threads in the thread pool is equal to the number of threads initialized by threadCount, then the pending tasks will no longer be obtained. Otherwise, a remote connection is established to obtain the pending tasks from the RMI server task queue tasks, and a thread is activated from the thread pool to execute the current task obtained. After the task is processed, the status of the task data in the relational database is written back, and the status attribute of the task data is updated to change it to the processed state.
[0020] Furthermore, the task data extraction node periodically clears the intermediate state waiting intermediate state information according to the preset parameter t, including:
[0021] The configuration parameter t is used as the threshold of the maximum task processing time in the current application scenario. A scheduled task is started, and the first waiting deletion interface or the second waiting deletion interface is used to periodically and batch clean up the intermediate state information of the intermediate state waiting that exceeds the maximum task processing time t. This ensures that the task data will not be pushed repeatedly, and also specifically cleans up the intermediate state information of the exception processing task in the intermediate state waiting, allowing the exception processing task to be extracted and processed again.
[0022] In a second aspect, the present invention provides a device for implementing a distributed cluster and multi-node task scheduling method, comprising: at least one processing unit, the processing unit is connected to a storage unit via a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the distributed cluster and multi-node task scheduling method is implemented.
[0023] In a third aspect, the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the distributed cluster and multi-node task scheduling method is implemented.
[0024] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art:
[0025] The present invention is based on RMI to build a remote method call design for task scheduling, including: building an RMI server and an RMI client; building a task queue tasks and intermediate state waiting of the RMI server; building an RMI remote interface; storing task data generated by different application scenarios in a relational database according to the application scenario type; extracting task data into task queue tasks and intermediate state waiting through a task data extraction node; each RMI client as a task processing node acquires task data in real time, dynamically, and through multiple nodes by calling the RMI remote interface, and after the task data is acquired, processes the task according to the actual application scenario and writes back the task state; the task data extraction node periodically clears the intermediate state waiting intermediate state information according to a preset parameter t; and pushes and processes the cyclic task data. The present invention is based on RMI to build a remote method call design for task scheduling, and provides a set of distributed cluster and multi-node task scheduling methods for lightweight distributed task scheduling demand scenarios, which has low resource occupancy, is lightweight, supports cross-platform, and is easy to expand nodes. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0027] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0028] Figure 1 A flowchart of a distributed cluster and multi-node task scheduling method provided by an embodiment of the present invention;
[0029] Figure 2 A schematic diagram of an apparatus for implementing a distributed cluster and multi-node task scheduling method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0032] Example 1
[0033] like Figure 1 As shown, the technology of the present invention implements a distributed cluster and multi-node task scheduling method to make up for the shortcomings of the existing technology. It includes:
[0034] In order to realize the distributed cluster and multi-node task scheduling method of this application, it is necessary to build a task scheduling remote method call design based on RMI. RMI (Remote Method Invocation) is a remote method call framework that adopts the client / server communication mode and deploys remote Java objects that provide various services on the server. RMI provides the client with a method to request access to the remote Java object on the server. RMI allows communication between Java virtual machines, so that objects running in different Java virtual machines can call methods of remote objects like calling local methods. Building a task scheduling remote method call design based on RMI includes: building an RMI server and an RMI client, wherein the RMI client contains a task data extraction node; building the task queue tasks and intermediate state waiting of the RMI server, which are used to store the intermediate state information of the task queue data and the task processing data respectively; and building a remote interface for processing task queue tasks and intermediate state waiting.
[0035] The task scheduling remote method call design of this application includes:
[0036] RMI server design. Create the server class TaskServer, first get the task object and assign it to ITaskitask, then create and export the Registry instance on the local host that accepts requests for the specified port (port) through LocateRegistry.createRegistry(port), such as port 8888, and finally bind the remote call task object through the following method: Naming.bind(" / / host:port / name", h), such as Naming.bind("rmi: / / localhost:8888 / ITask", itask).
[0037] The Naming class provides methods for storing and obtaining remote object references in the object registry. Each method of the Naming class can take a name as one of its parameters. The name is in the URL format of java.lang.String:host:port / name.
[0038] host: the host where the registry is located (remote or local). If omitted, the default is the local host. port: the port number that the registry accepts calls. If omitted, the default is 1099. name: a simple string that is not interpreted by the registry.
[0039] Build the public variables of the RMI server: task queue tasks and intermediate state waiting, which store the intermediate state information of task queue data and task processing data respectively.
[0040] In this application, task queues and intermediate states are used as public variables for data storage, transfer, and consumption in the entire project. The code example for creating public variables on the RMI server is as follows:
[0041] "public static LinkedBlockingQueue <Map<String,Object> >tasks = newLinkedBlockingQueue <Map<String,Object> >();
[0042] public static Map<String,Object> waiting = new ConcurrentHashMap<String,Object> (); ".
[0043] Use LinkedBlockingQueue to create task queues tasks, which are used to store task queue data. Use ConcurrentHashMap to create intermediate state waiting, which is used to store intermediate state information of task processing data. ConcurrentHashMap supports high concurrent updates and queries.
[0044] Among them, the link queue is a blocking queue based on a linked list. By default, the size of the link queue is Integer.MAX_VALUE. The link queue has almost no boundaries and can grow dynamically as elements are added. At the same time, the link queue is implemented by a single linked list, and elements can only be taken from the head of the queue and added to the tail of the queue. The link queue uses a lock separation technology containing two locks to achieve mutual non-blocking of entry and exit. Adding elements and obtaining elements have independent locks, realizing the read-write separation of the link queue and supporting parallel execution of read and write operations.
[0045] This application utilizes the advantages of linked queues and synchronous hash maps for high concurrent execution to meet the processing requirements of intermediate state information of task queue data and task processing data.
[0046] RMI client design.
[0047] The client includes task processing nodes and task data extraction nodes.
[0048] A single task processing node processes and configures a multi-threaded processing mechanism.
[0049] When the task processing node is started, it creates threads by creating a thread pool and assigns them to the variable ExecutorService pushthreads = Executors.newFixedThreadPool(threadCount), where threadCount is the number of threads created. The threadCount parameter should be set according to the specific application scenario and server configuration capabilities.
[0050] Then, a monitoring thread is started independently. After the monitoring thread is started, it starts to process tasks. An infinite loop is started through for (;;) to monitor the data of the task queue tasks of the RMI server. In the loop body, the actual task push logic is executed through try, catch and finally methods. In the finally method, Thread.sleep(time) is used to achieve the interval time of each loop, so as to control the frequency of the current task processing node to obtain task data. Time is a threshold parameter that can be adjusted according to the actual situation.
[0051] In the try method, when calling to obtain task data through the RMI remote interface, first obtain the current number of active threads in the thread pool, that is, ((ThreadPoolExecutor) pushthreads).getActiveCount(), and compare the current number of active threads in the thread pool with the number of threads in the thread pool, threadCount. If the current number of active threads in the thread pool is equal to threadCount, then no more pending tasks will be obtained to prevent the task processing threads from running out but still continuously obtaining task data from the task queue, causing a single node to obtain tasks without limit, resulting in task processing backlog or memory exception problems. If the current number of active threads in the thread pool is less than the number of initialized threads, a remote connection is established to obtain pending task data from the RMI server-side task queue tasks.
[0052] That is, ITask itask =(ITask) Naming.lookup("rmi: / / host:port / ITask "), where host is the IP address of the server and port is the port.
[0053] Then through Map<String,Object> The task data is obtained by taskmap = itask.poll(). If taskmap is empty, it means that there is no data in the current remote call task queue. If taskmap is not empty, a thread is activated from the thread pool to execute the current task obtained. The specific task execution process is processed according to the actual application scenario. After the task is processed, the status of the task data in the relational database is written back, and the status attribute of the task data is updated to change to the processed status.
[0054] Therefore, the number of task processing nodes can be increased or decreased arbitrarily according to the application scenario. At the same time, the task processing capacity of each node can be dynamically set through dynamic configuration, thread pool thread number threadCount, and task data extraction frequency according to the performance configuration of the node server, thereby realizing multi-node, high-concurrency, and dynamically adjusted task processing in the application scenario.
[0055] Task data extraction node settings.
[0056] The functions implemented by the task data extraction node include two aspects:
[0057] The first aspect: sort the task data in the relational database in ascending order according to the task generation time, select the task data with a status of not pushed as the task data to be extracted, and the single extracted data volume pushCount and the time interval pushTime should be set according to the capacity of the task processing node, and should also consider the actual application scenario, which is a configurable dynamic threshold parameter.
[0058] Establish a link instance, i.e. ITask itask =(ITask) Naming.lookup("rmi: / / host:port / ITask "), obtain the task data to be extracted quantitatively and regularly according to pushCount and pushTime in a loop, first format the task data to be extracted into a JSON string, and then use the link instance to call the data push interface to push the task data to the task queue tasks, and at the same time store the unique identification id and storage time of the current task data to be processed in the intermediate state waiting.
[0059] The second aspect: In the task processing node, the status of the task data in the relational database is updated after the task is processed, but the intermediate state information of the intermediate state waiting is not deleted immediately.
[0060] The configuration parameter t is used as the threshold of the maximum task processing time in the current application scenario. A scheduled task is started, and the first waiting deletion interface or the second waiting deletion interface is used to periodically and batch clean up the intermediate state information of the intermediate state waiting that exceeds the maximum task processing time t. This ensures that the task data will not be pushed repeatedly, and also specifically cleans up the intermediate state information of the exception processing task in the intermediate state waiting, allowing the exception processing task to be extracted and processed again.
[0061] Define the remote interface based on RMI remote method call. Define the remote interface class ITask based on RMI remote method call. The interfaces involved in the remote interface class ITask include:
[0062] (1) Data push interface that pushes data to task queue tasks:
[0063] public void push(String mapstr) throws RemoteException;
[0064] (2) Data extraction interface for extracting data from task queue tasks:
[0065] Public Map<String,Object> poll() throws RemoteException, InterruptedException;
[0066] (3) Task counting interface to calculate the number of tasks in the entire task queue:
[0067] public long tcount() throws RemoteException;
[0068] (4) Waiting counting interface to calculate the number of waiting in the entire intermediate state:
[0069] public long wcount() throws RemoteException;
[0070] (5) Delete the first waiting deletion interface of the corresponding intermediate state information according to the key of the intermediate state information in the intermediate state waiting:
[0071] public void delete(String key) throws RemoteException;
[0072] (6) Clear the waiting clearing interface for all intermediate state information in the intermediate state waiting:
[0073] public long clearwait() throws RemoteException;
[0074] (7) The second waiting deletion interface for deleting the corresponding intermediate state information in the intermediate state waiting according to the storage time:
[0075] public void removeMapByTime(Integer t) throws RemoteException;
[0076] Implement the specific functions of the RMI remote interface.
[0077] Set the RMI remote interface implementation class to TaskImpl. Define the function implementation of each RMI remote interface in the RMI remote interface implementation class.
[0078] Implementation of data push interface push(String mapstr):
[0079] (1) The mapstr parameter is a map string representing a task data to be processed. In order to be more compatible with more scenarios, the mapstr must contain the id attribute as the unique identifier of the current data, the status attribute as the task processing status, and the createDate as the task data creation time. The content of the custom task data attribute can be set according to any actual scenario. The map string is converted into a JSONObject through the JSONObject.parseObject(mapstr) method and assigned to the variable jsonObject, and then (Map<String,Object> )jsonObject is finally converted to Map<String,Object> Type and assign it to the variable mapData.
[0080] (2) When extracting task data with a status of pending and entering the task queue tasks, the id attribute data and storage time of the current task will also be stored in the intermediate state waiting, so as to monitor whether there is duplication in the task processing data and prevent duplicate processing.
[0081] Specifically, waiting.containsKey(mapData.get("id").toString()) determines whether the pending task data mapData currently pushed over already exists in the task queue tasks according to the id of the pending task data in the intermediate state waiting. If it already exists, it will not be pushed to the task queue tasks again. If it does not exist, the pending task data mapData will be pushed to the task queue tasks.
[0082] The linked queue to which the task queue tasks belongs is an unbounded queue, which means it has almost no boundaries. The queue can grow dynamically as elements are added. If there is no remaining memory, the queue will throw an exception. Therefore, in order to avoid the situation where the task queue tasks are too large and cause server load or memory fullness, the maximum data volume threshold Tmax that the task queue can carry is set, which can be dynamically adjusted according to the actual application scenario and server configuration.
[0083] Get the length of the current task queue tasks through the tasks counting interface or tasks.size() operation. Whether the length of the current task queue tasks is greater than the maximum data volume threshold Tmax that can be carried determines whether to store the pending task in the task queue tasks. If it is less than the maximum data volume threshold Tmax that can be carried, the task is stored at the end of the task queue tasks through the task.add(mapData) method, and the unique identification id and storage time of the current pending task data are stored in the intermediate state data waiting through the waiting.put(map.get("id").toString(), new Date()) method. If it is greater than or equal to the maximum data volume threshold Tmax that can be carried, the current task data will not enter the pending task queue tasks, and wait for the next batch of data to be verified again.
[0084] Data extraction interface Map<String,Object> Implementation of poll():
[0085] The poll operation of the linked queue is used to take the first element of the task queue tasks. If the taking fails, false is returned. If successful, the corresponding element can be returned. Execute "tasks.poll();" to obtain the task data to be processed from the task queue tasks.
[0086] Implementation of tasks counting interface tcount():
[0087] Get the length of the task queue tasks by linking the queue's size operation "tasks.size();".
[0088] Implementation of the waitting count interface wcount():
[0089] The Map length of the intermediate state watting is returned by the size operation "waitting.size();" of the synchronized hash map;
[0090] The first waiting delete interface delete(String key) implementation:
[0091] The remove operation "waitting.remove(key);" of the synchronous hash map is used to delete elements in the map by key. A scheduled task is started, and the first waiting deletion interface is used to perform scheduled and batch cleaning of the intermediate state information of the intermediate state waiting that exceeds the maximum task processing time t, so as to ensure that the task data will not be pushed repeatedly. At the same time, the intermediate state information of the exception handling task in the intermediate state waiting is also cleared in a targeted manner, allowing the exception handling task to be extracted and processed again. The scheduled task regularly cleans up the intermediate state information in the intermediate state waiting that is greater than the maximum task processing time t, specifically through for(Map.Entry<String,Object> entry:waiting.entrySet()) traverses the intermediate state information of the intermediate state waiting, and compares the storage time (Date) entry.getValue() of the traversed intermediate state information with the current system time. If the time difference between the two is greater than or equal to the maximum task processing time t, then obtain the key according to entry.getKey(), and delete the intermediate state information of the task intermediate state waiting according to the key through the first waiting deletion interface.
[0092] Implementation of waitting clearing interface clearwait():
[0093] Deleting all elements in the map is achieved through the clear operation "waitting. clear();" of the synchronized hash map.
[0094] The second wait removes the implementation of the interface removeMapByTime(Integer t):
[0095] In the design, different pending task data may be generated in various application scenarios, and the pending task data is stored in a relational database. Then, according to the status attribute of the task data, the pending task data is pushed to the task queue tasks at a regular interval. After the qualified task data is added to the end of the task queue tasks, the attributes of the synchronized task data will also be added to the intermediate state waiting through waiting.put(map.get("id").toString(), new Date()).
[0096] When the attributes of the task data are added to the intermediate state waiting by means of waiting.put(map.get("id").toString(), new Date()), the creation time corresponding to the task data key is recorded. This application needs to configure the parameter t as the threshold of the maximum task processing time in the current application scenario, start a scheduled task, and regularly delete the second waiting interface to implement scheduled and batch cleaning of the intermediate state information of the intermediate state waiting that exceeds the maximum task processing time t, so as to ensure that the task data will not be pushed repeatedly, and at the same time, the intermediate state information of the exception handling task in the intermediate state waiting is cleared in a targeted manner, allowing the exception handling task to be extracted and processed again. The scheduled task regularly cleans up the intermediate state information in the intermediate state waiting that is greater than the maximum task processing time t, specifically through for(Map.Entry<String,Object> entry: waitting.entrySet()) traverses the intermediate state information of the intermediate state waiting, compares the storage time (Date) entry.getValue() of the traversed intermediate state information with the current time, and if the time difference between the two is greater than or equal to the maximum task processing time t, the intermediate state information of the task intermediate state waiting is deleted according to the storage time through the second waiting deletion interface.
[0097] It is set to store the task data generated by different application scenarios into a relational database according to the application scenario type. Each task data contains the following fields: id: unique identifier of the task; status: task processing status; createDate: task generation time; custom task data attributes, among which the custom task data attributes can be arbitrarily configured according to actual conditions.
[0098] Then the task data is extracted into the task queue tasks and the intermediate state waiting through the task data extraction node.
[0099] Each task processing node calls the RMI remote interface to obtain task data in real time, dynamically, and through multiple nodes based on the task queue tasks and intermediate state waiting conditions. After the task data is obtained, the task is processed and the task status is written back based on the actual application scenario. The intermediate state information in the intermediate state waiting is cleared periodically based on the preset threshold.
[0100] Push and process cyclic task data.
[0101] Here is a scenario where a company's "sales battle report" message is automatically pushed:
[0102] Set up a relational database table im_sent_log for message push. The main fields include id (task unique identifier), touid (push recipient), create_by (creator), create_date (creation time), status (message status), msg (push result description), taskid (message source task), end_date (push end time), im_content (content), im_text (task text, when the task is of image type, im_text is a dynamic json string), and im_type (task type).
[0103] Package the RMI server into a jar file, then start the RMI server and establish a remote call service. In the command prompt box, enter: java -jar rmi.jar. The service is started successfully, and the remote link binding is successful. The registered service port number is 8888, and the server IP address is 10.20.132.66.
[0104] Task data generation:
[0105] Task 1: Create a sales report push task through the automatic push system. The system user logs in to the system and uploads the designed PPT template. The template content includes two indicator data: the current day's protection premium and the current day's long-term insurance performance manpower. Then configure the push frequency to push once every hour, and then fill in the data acquisition logic. Use the SQL extraction script to extract the employee data whose total premium ammnt exceeds 1,000 yuan: select the current day's protection premium, the current day's long-term insurance performance manpower, salesno from im_sales where amnt>1000, where salesno is the employee number, and the task is started. Task id: d999f562-a727-4664-98de-f7ce23c03c4c.
[0106] Task 2: Create a text message reminder task for employee birthday wishes through the automatic push system. The system user logs in to the system and sets the push content as:
[0107] Dear Partner:
[0108] This year is your {1}th birthday, I wish you a happy birthday!
[0109] Then configure the push frequency to push once a day at 10 o'clock, and then fill in the data acquisition logic: extract the script through SQL; select age, salesno from employee where birthday=convert(varchar(10), getdate(), 120), start the task, task id: f41f653d-370f-4291-a243-4addee79b3b0.
[0110] On any day, when the task starts and reaches the push time, the data to be pushed is automatically obtained according to the data extraction script, and then the messages are dynamically organized. The battle report data to be pushed can be generated in batches according to the actual amount of data obtained, and the battle report data is stored in the im_sent_log table of the SQL Server relational database table. The id is the automatically generated uuid, the default push status is to be pushed, and touid is the receiving object.
[0111] The following is an example of a relational database table im_sent_log for message data:
[0112]
[0113] Start the task data extraction node. The task data extraction node is deployed on the server 10.20.132.67 and is interconnected with the RMI server 10.20.132.67. Start the scheduled task, extract data every 10 seconds, and extract 2000 data each time. Extract according to the creation time and push status. The earlier the task time, the priority will be extracted. Only the data in the state of being pushed will be extracted. The currently extracted ids are c13ed33e-48e7-4187-a6f5-f39de2de7554, 6796f01f-b8f0-4caf-9433-e4f5d61769d7, eb78c568-8c4a-441a-a526-0fae5154fd2d, 861114f4-e2a The four tasks to be pushed of 6-4e72-8925-00721c8a5b55 come from the tasks with task id d999f562-a727-4664-98de-f7ce23c03c4c and f41f653d-370f-4291-a243-4addee79b3b0. That is, two tasks generated two push tasks at a time, which will be pushed to employees with work numbers 13708812, 13708813, 13708814 and 13708815 through the company's APP message interface respectively.
[0114] Then, ITask itask = (ITask) Naming.lookup ("rmi: / / host:port / ITask") is used to establish a link instance, where host is the server's IP address 10.20.132.66, and port is the server's port 8888, to establish a remote call link with the server's RMI message queue. Then, through a for loop, the imtasks list is traversed, and the extracted battle report task data is added to the task queue tasks and the intermediate state waiting to be pushed by calling the data push interface.
[0115] After the task data extraction node is started, the scheduled task is started every 10 minutes. By calling the second waiting deletion interface removeMapByTime(t), the intermediate state information that exceeds the maximum task processing time t is cleaned up in batches. Currently, t=10 is set. While preventing repeated data push, the battle report data that has not been successfully pushed can also be re-entered into the task queue for special circumstances.
[0116] Deploy task push application service nodes as task processing nodes on the five servers with IP addresses 10.20.132.68, 10.20.132.69, 10.20.132.70, 10.20.132.71, and 10.20.132.72, respectively, and ensure that the server of the task push application service node is interconnected with the network 10.20.132.66 of the server where the RMI server is located. Configure the thread pool number threadCount of each task push application service node to 200, the push interval to 1000 milliseconds, start the application service of the task push application service node, and establish a remote call link with the RMI server. Immediately after the system starts, a thread is automatically started, and then an infinite loop is started through a for (;;) loop. Every 1000 milliseconds, monitor whether there are pending battle report tasks in the RMI server task queue that need to be pushed, and obtain the battle report tasks through the RMI remote interface data extraction interface.
[0117] Currently, the five server nodes 68-72 are all in idle state, and each node thread pool has idle threads. In the next second, the two nodes 10.20.132.68 and 10.20.132.69 extracted the IDs c13ed33e-48e7-4187-a6f5-f39de2de7554 and 6796f01f-b8f0-4caf-9433-e4f5d respectively. For task 61769d7, the two nodes 10.20.132.70 and 10.20.132.71 pulled the data of two tasks with IDs eb78c568-8c4a-441a-a526-0fae5154fd2d and 861114f4-e2a6-4e72-8925-00721c8a5b55 respectively. The node 10.20.132.72 is temporarily idle.
[0118] 10.20.132.68, 10.20.132.69 two node servers, after getting the task data, analyze the Map<String,Object> The im_type field in the task record is found to be in picture format, so the company's APP picture message interface is called to push the pictures to 13708812 and 13708813. After the push is successful, the status field of the task record is updated to success, msg is push success, and end_date is the storage time according to the id.
[0119] 10.20.132.70, 10.20.132.71 two node servers, after getting the task data, analyze the Map<String,Object> The im_type field in the task record is found to be a text message, so the company's APP text message interface is called to push the text message to 13708814 and 13708815. After the push is successful, the status field of the task record is updated to success, msg is pushed successfully, and end_date is the storage time according to the id.
[0120] The scheduled task is executed every 10 minutes. The scheduled task deletes the intermediate state information of the intermediate state waiting that exceeds the maximum task processing time t by calling the second waiting deletion interface. Specifically, for(Map.Entry<String,Object> entry:waiting.entrySet()) traverses the intermediate state information in the intermediate state waiting, and compares the current storage time (Date) entry.getValue() with the current time. If it is greater than 10, the task data is deleted through the second waiting deletion interface. The ids currently in waiting are c13ed33e-48e7-4187-a6f5-f39de2de7554, 6796f01f-b8f0-4caf-9433-e4f5d61769d7, eb78c568-8c4a-441a-a526-0fae5154fd2d, and 861114f4-e2a6-4e72-8925-00721c8a5b55, all of which meet the condition that the existence time exceeds 10 minutes. So the four elements are deleted from the intermediate state waiting.
[0121] The loop steps realize real-time dynamic processing of generated message tasks, and push them to enterprise app user messages efficiently and quickly in batches, multi-nodes, and distributed manner.
[0122] Example 2
[0123] See also Figure 2 As shown, an embodiment of the present invention provides a device for implementing a distributed cluster and multi-node task scheduling method, including: at least one processing unit, the processing unit is connected to a storage unit via a bus unit, the storage unit is a computer-readable storage medium, and can be used to store software programs, computer executable programs, and modules, such as software programs, computer executable programs, and modules corresponding to a distributed cluster and multi-node task scheduling method in an embodiment of the present invention. The processing unit implements the above-mentioned distributed cluster and multi-node task scheduling method by running the software programs, computer executable programs, and modules stored in the storage unit, including:
[0124] The design of task scheduling remote method call based on RMI includes: constructing an RMI server and an RMI client, wherein the RMI client includes a task data extraction node; constructing task queues tasks and intermediate state waiting of the RMI server, which are respectively used to store the intermediate state information of task queue data and task processing data; constructing an RMI remote interface for processing task queue tasks and intermediate state waiting, including: a data push interface for pushing data to task queue tasks, a data extraction interface for extracting data from task queue tasks, a tasks counting interface for calculating the number of tasks in the entire task queue, a waiting counting interface for calculating the number of waiting in the entire intermediate state, a first waiting deletion interface for deleting the corresponding intermediate state information according to the key of the intermediate state information in the intermediate state waiting, a waiting clearing interface for clearing all intermediate state information in the intermediate state waiting, and a second waiting deletion interface for deleting the corresponding intermediate state information in the intermediate state waiting according to the storage time;
[0125] According to the application scenario type, the task data generated by different application scenarios are stored in a relational database, where the task data contains the following fields: id: unique identifier of the task; status: task processing status; createDate: task generation time; custom task data attributes adapted to the application scenario;
[0126] The task data is extracted into the task queue tasks and the intermediate state waiting through the task data extraction node; each RMI client as a task processing node obtains the task data in real time, dynamically and multi-node by calling the RMI remote interface, and after the task data is obtained, the task is processed according to the actual application scenario and the task state is written back; the task data extraction node periodically clears the intermediate state waiting intermediate state information according to the preset parameter t; and the task data is pushed and processed in a loop.
[0127] Of course, the computer program stored in the storage unit of the device for implementing the distributed cluster and multi-node task scheduling method provided by an embodiment of the present invention is not limited to the method operations described above, but can also execute related operations in the distributed cluster and multi-node task scheduling method provided by any embodiment of the present invention.
[0128] Example 3
[0129] An embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program. When the computer program is executed, the distributed cluster and multi-node task scheduling method is implemented, including:
[0130] The design of task scheduling remote method call based on RMI includes: constructing an RMI server and an RMI client, wherein the RMI client includes a task data extraction node; constructing task queues tasks and intermediate state waiting of the RMI server, which are respectively used to store the intermediate state information of task queue data and task processing data; constructing an RMI remote interface for processing task queue tasks and intermediate state waiting, including: a data push interface for pushing data to task queue tasks, a data extraction interface for extracting data from task queue tasks, a tasks counting interface for calculating the number of tasks in the entire task queue, a waiting counting interface for calculating the number of waiting in the entire intermediate state, a first waiting deletion interface for deleting the corresponding intermediate state information according to the key of the intermediate state information in the intermediate state waiting, a waiting clearing interface for clearing all intermediate state information in the intermediate state waiting, and a second waiting deletion interface for deleting the corresponding intermediate state information in the intermediate state waiting according to the storage time;
[0131] According to the application scenario type, the task data generated by different application scenarios are stored in a relational database, where the task data contains the following fields: id: unique identifier of the task; status: task processing status; createDate: task generation time; custom task data attributes adapted to the application scenario;
[0132] The task data is extracted into the task queue tasks and the intermediate state waiting through the task data extraction node; each RMI client as a task processing node obtains the task data in real time, dynamically and multi-node by calling the RMI remote interface, and after the task data is obtained, the task is processed according to the actual application scenario and the task state is written back; the task data extraction node periodically clears the intermediate state information of the intermediate state waiting according to the preset parameter t; and the push and processing of task data are cyclic.
[0133] A computer-readable storage medium provided in an embodiment of the present invention stores a computer program which is not limited to the method operations described above, but can also execute related operations in a distributed cluster and multi-node task scheduling method provided in any embodiment of the present invention.
[0134] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, structures or units, which can be electrical, mechanical or other forms.
[0135] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0136] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0137] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A distributed cluster, multi-node task scheduling method, characterized in that: include: The design of task scheduling remote method call based on RMI includes: constructing RMI server and RMI client, wherein the RMI client includes task processing node and task data extraction node; constructing task queue tasks and intermediate state waiting of RMI server, which are respectively used to store intermediate state information of task queue data and task processing data, and taking task queue tasks and intermediate state waiting as public variables for data storage, transfer and consumption in the whole project; constructing RMI remote interface for processing task queue tasks and intermediate state waiting, including: data push interface for pushing data to task queue tasks, data extraction interface for extracting data from task queue tasks, tasks counting interface for calculating the number of tasks in the whole task queue, waiting counting interface for calculating the number of waiting in the whole intermediate state, first waiting deletion interface for deleting corresponding intermediate state information according to the key of intermediate state information in intermediate state waiting, waiting clearing interface for clearing all intermediate state information in intermediate state waiting, and second waiting deletion interface for deleting corresponding intermediate state information in intermediate state waiting according to storage time; According to the application scenario type, the task data generated by different application scenarios are stored in a relational database, where the fields of the task data include: id: unique identifier of the task; status: task processing status; createDate: task generation time; custom task data attributes adapted to the application scenario; The task data is extracted from the relational database into the task queue tasks and the intermediate state waiting through the task data extraction node, including: sorting the task data in the relational database in ascending order according to the task generation time; selecting the task data to be extracted according to the task data status being in the unpushed state, establishing a link instance, and obtaining the task data to be extracted quantitatively and regularly in a cyclic manner, first formatting the task data to be extracted into a JSON string, and then using the link instance to call the data push interface to push the task data into the task queue tasks, and at the same time storing the unique identification id and storage time of the current task data to be processed in the intermediate state waiting, so as to judge whether the task data to be processed currently pushed over already exists in the task queue tasks according to the id of the task data to be processed in the intermediate state waiting, and if it already exists, it will not be repeatedly pushed to the task queue tasks, and if it does not exist, the task data to be processed will be pushed to the task queue tasks; Each RMI client as a task processing node obtains task data in real time, dynamically, and through multiple nodes by calling the RMI remote interface. After obtaining the task data, it processes the task according to the actual application scenario and writes back the task status, including: When the task processing node is configured to start, a thread is created by creating a thread pool; an infinite loop monitoring thread is independently started, and the monitoring thread executes the actual task push logic through try, catch and finally methods. In the finally method, the thread sleep time is set to limit the interval time of each loop, so as to control the frequency of the current task processing node obtaining task data; in the try method, the current number of active threads in the thread pool is obtained and compared with the number of threads in the thread pool. If the current number of active threads in the thread pool is equal to the number of threads initialized by threadCount, then the pending tasks will no longer be obtained. Otherwise, a remote connection is established to obtain the pending tasks from the RMI server task queue tasks, and a thread is activated from the thread pool to execute the current task obtained. After the task is processed, the task data in the relational database is written back to the status, and the status attribute of the task data is updated to change to the processed status; The task data extraction node periodically clears the intermediate state waiting intermediate state information according to the preset parameter t, wherein, in the task processing node, the state of the task data in the relational database is updated after the task processing is completed, but the intermediate state waiting intermediate state information is not immediately deleted; the configuration parameter t is used as the threshold of the maximum task processing time in the current application scenario, and a scheduled task is started, and the first waiting deletion interface or the second waiting deletion interface is periodically used to implement the scheduled and batch cleaning of the intermediate state waiting intermediate state information that exceeds the maximum task processing time t, so as to ensure that the task data will not be pushed repeatedly, and at the same time, the intermediate state information of the exception processing task in the intermediate state waiting is also specifically cleared, so as to allow the exception processing task to be extracted and processed again; Push and process cyclic task data.
2. The distributed cluster and multi-node task scheduling method according to claim 1, characterized in that: A linked queue is used to create task queues tasks, which are used to store task queue data. A linked queue is a blocking queue implemented based on a linked list. The linked queue is implemented by a single linked list. Elements can only be taken from the head of the queue and added to the tail of the queue. The linked queue uses a lock separation technology containing two locks to achieve non-blocking entry and exit of the queue. Adding elements and getting elements have independent locks, realizing read-write separation of the linked queue and supporting parallel execution of read and write operations. A synchronous hash map is used to create an intermediate state waiting, which is used to store the intermediate state information of task processing data.
3. The distributed cluster and multi-node task scheduling method according to claim 1, characterized in that: The RMI client that serves as the task processing node configures a multi-threaded processing mechanism.
4. The distributed cluster and multi-node task scheduling method according to claim 1, characterized in that: The process of extracting task data to task queues tasks and intermediate state waiting is controlled according to task queues tasks and intermediate state waiting, including: Set the maximum data volume threshold Tmax that the task queue tasks can carry. The maximum data volume threshold Tmax can be dynamically adjusted according to the actual application scenario and server configuration; The data push interface obtains the length of the current task queue tasks through the tasks counting interface. The data push interface determines whether to store the pending task in the task queue tasks based on whether the length of the current task queue tasks is greater than the maximum data volume threshold Tmax that can be carried. If the length of the current task queue tasks is less than the maximum data volume threshold Tmax that can be carried, and it is not in the task queue task, the pending task is stored at the end of the task queue tasks. At the same time, the unique identification id and storage time of the current pending task data are stored in the intermediate state data waiting. If it is greater than or equal to the maximum data volume threshold Tmax that can be carried, or it is in the task queue task, the current task data will not enter the pending task queue tasks, and wait for the next batch of data to be verified again.
5. A device for implementing a distributed cluster and multi-node task scheduling method, characterized in that: include: At least one processing unit, the processing unit is connected to a storage unit via a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the distributed cluster and multi-node task scheduling method as described in any one of claims 1-4 is implemented.
6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the distributed cluster and multi-node task scheduling method as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Task scheduling method and device, storage medium and server node
CN110377406A
Data processing method and device and storage medium
CN113806065A