Distributed batch data processing method and device and storage medium
Through the distributed batch data processing method, the problem of low efficiency and stability in large-scale business data processing is solved, efficient and stable data processing is achieved, and resource waste is avoided.
Patent Information
- Application Number
- CN202311555879.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
The prior art has low efficiency and stability when processing large-scale business data, which can easily lead to overflow of stand-alone hardware resources and waste resources.
The distributed batch data processing method is adopted, and the data is processed in batches by configuring a scheduled task list, generating a scheduled task instance list, deploying a task execution cluster, generating a scheduled task configuration information, allocating a task instance to the executor, using the modulus operation rules to process the data in batches, and summarizing the processing results and saving them into the database.
It effectively improves the efficiency and stability of data processing, avoids the failure of overflow of single-machine hardware resources, maximizes the efficiency of host usage, and reduces resource waste.
Smart Images

Figure CN120020722A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer data processing technologies, and in particular, to a distributed batch data processing method, apparatus, and storage medium. Background Art
[0002] With the steady growth of the digital industry scale, the business data that enterprises need to process will become larger and larger. In this case, the efficiency and stability of data processing will form a bottleneck for business development, unable to keep up with the speed of business development, and restricting business development. For example, when the efficiency and stability of data processing are low, it is easy to cause faults such as overflow of single-machine hardware resources, wasting resources.
[0003] Based on this, how to improve the efficiency of data processing and ensure the stability of data processing is of great significance to business development. Summary of the Invention
[0004] Embodiments of this application provide a distributed batch data processing method, apparatus, and storage medium to solve the problems existing in the related technologies. The technical solutions are as follows:
[0005] In a first aspect, embodiments of this application provide a distributed batch data processing method, including:
[0006] According to the business data processing requirements of the target business scenario, configure a timing task list for the large amount of data to be processed in the target business scenario, where the timing task list includes several timing tasks and their corresponding timing task tags;
[0007] According to the business processing logic of the target business scenario and the timing task list, generate a timing task instance list for any one of the timing tasks, where the timing task instance list includes several timing task instances and their corresponding timing task tags and timing task instance parameters;
[0008] Deploy a task execution cluster according to the timing task instance list, where the task execution cluster includes several executors;
[0009] Generate timing task configuration information according to the timing task list and the task execution cluster, where the timing task configuration information includes the association relationship between the executor and the timing task tag;
[0010] According to the timing task configuration information, assign any one of the timing task instances to the corresponding executor in the task execution cluster according to the timing task tag;
[0011] According to the timing task instance parameters, use the task execution cluster to batch process the large amount of data by adopting a modulo operation rule, and summarize each processing result and then store it in the database.
[0012] In one implementation, according to the business data processing requirements of the target business scenario, configuring the list of scheduled tasks for the large volume of data to be processed in the target business scenario includes:
[0013] Determine several processing dimensions of the large volume of data according to the business data processing requirements;
[0014] Configure the list of scheduled tasks according to the several processing dimensions.
[0015] In one implementation, according to the business processing logic of the target business scenario and the list of scheduled tasks, generating the list of scheduled task instances for any one of the scheduled tasks includes:
[0016] Determine the data volume segmentation unit of the large volume of data according to the business processing logic;
[0017] Determine the number of data sets that can be obtained after segmenting the large volume of data according to the data volume segmentation unit;
[0018] Generate the list of scheduled task instances for any one of the scheduled tasks according to the number of data sets and the list of scheduled tasks.
[0019] In one implementation, deploying the task execution cluster according to the list of scheduled task instances includes:
[0020] According to the list of scheduled task instances, use the containerized deployment tool to instantiate several application instances to obtain the task execution cluster, where one application instance is one executor.
[0021] In one implementation, generating the scheduled task configuration information according to the list of scheduled tasks and the task execution cluster includes:
[0022] According to the list of scheduled tasks, configure the execution period, number, and task name for any one of the scheduled tasks, and then select an executor from the task execution cluster to execute any one of the scheduled tasks to generate the scheduled task configuration information.
[0023] In one implementation, using the task execution cluster to batch process the large volume of data according to the modulo operation rule according to the scheduled task instance parameters includes:
[0024] According to the scheduled task instance parameters, obtain the business data corresponding to any one of the scheduled task instances, where the business data corresponding to all the scheduled task instances constitutes the large volume of data;
[0025] Initialize a number of sub - threads according to the modulo operation rule by using the executor corresponding to any one of the said timing task instances;
[0026] Process the said service data in batches through a number of the said sub - threads.
[0027] In one implementation, summarizing each processing result and then storing it in the database includes:
[0028] When it is monitored that a number of the said sub - threads have all completed data processing, collect the processing results of a number of the said threads;
[0029] Summarize the processing results of a number of the said sub - threads, and then store the summarized result in the said database.
[0030] In a second aspect, the embodiments of the present application also provide a distributed batch data processing device, including:
[0031] A configuration unit, used to configure a timing task list of a large amount of data to be processed in the said target service scenario according to the service data processing requirements of the target service scenario, where the timing task list includes a number of timing tasks and their corresponding timing task tags;
[0032] A deployment unit, used to generate a timing task instance list of any one of the said timing tasks according to the service processing logic of the said target service scenario and the timing task list, where the timing task instance list includes a number of timing task instances and their corresponding timing task tags, timing task instance parameters; deploy a task execution cluster according to the timing task instance list, where the task execution cluster includes a number of executors;
[0033] A processing unit, used to generate timing task configuration information according to the timing task list and the task execution cluster, where the timing task configuration information includes the association relationship between the executor and the timing task tag; according to the timing task configuration information, assign any one of the said timing task instances to the corresponding executor in the task execution cluster according to the timing task tag; according to the timing task instance parameters, use the task execution cluster to batch - process the large amount of data by adopting the modulo operation rule, and summarize each processing result and then store it in the database.
[0034] In one implementation, the said configuration unit is specifically used for:
[0035] Determine a number of processing dimensions of the large amount of data according to the service data processing requirements;
[0036] Configure the timing task list according to a number of the said processing dimensions.
[0037] In one implementation, the said deployment unit is specifically used for:
[0038] Determine the data volume segmentation unit of the large volume of data according to the service processing logic;
[0039] Determine the number of data sets that can be obtained after segmenting the large volume of data according to the data volume segmentation unit;
[0040] Generate a list of scheduled task instances for any one of the scheduled tasks according to the number of data sets and the list of scheduled tasks.
[0041] In one implementation, the deployment unit is specifically configured to:
[0042] Instantiate a number of application instances by using a containerized deployment tool according to the list of scheduled task instances to obtain the task execution cluster, where one application instance is one executor.
[0043] In one implementation, the processing unit is specifically configured to:
[0044] Configure an execution period, a number, and a task name for any one of the scheduled tasks according to the list of scheduled tasks, and then select an executor for executing any one of the scheduled tasks from the task execution cluster to generate the scheduled task configuration information.
[0045] In one implementation, the processing unit is specifically configured to:
[0046] Obtain the service data corresponding to any one of the scheduled task instances according to the scheduled task instance parameters, where the service data corresponding to all the scheduled task instances constitutes the large volume of data;
[0047] Initialize a number of sub-threads according to the modulo operation rule by using the executor corresponding to any one of the scheduled task instances;
[0048] Process the service data in batches through the number of sub-threads.
[0049] In one implementation, the processing unit is specifically configured to:
[0050] When it is monitored that all the number of sub-threads have completed data processing, collect the processing results of the number of threads;
[0051] Summarize the processing results of the number of sub-threads, and then store the summary result in the database.
[0052] In a third aspect, an embodiment of the present application further provides a computer device, which includes: a memory and a processor. Instructions are stored in the memory and loaded and executed by the processor to implement the method in any one of the above aspects. Wherein, the memory and the processor communicate with each other through an internal connection path.
[0053] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the method in any one of the above aspects is implemented.
[0054] The advantages or beneficial effects in the above technical solutions at least include:
[0055] The present application can be driven by business data processing requirements, customize the timing tasks of data processing, and support specifying timing task tags for timing tasks, which is convenient to meet the corresponding business requirements of the target business scenario; it can also support horizontal expansion and contraction of the executors in the task execution cluster according to whether the business scenario is in the data peak period through timing task tags, maximizing the utilization efficiency of the host, avoiding resource waste, and can also parallel process business data through the task execution cluster, splitting a large amount of data into several small data sets, then distributing the small data sets to multiple timing tasks for parallel processing, and finally summarizing the data processing results, so as to effectively improve the efficiency of data processing and ensure the stability of data processing, and avoid the failure of single-machine hardware resource overflow.
[0056] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the above-described illustrative aspects, embodiments, and features, further aspects, embodiments, and features of the present application will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in the present application and should not be regarded as limiting the scope of the present application.
[0058] Figure 1 It is a schematic flowchart of a distributed batch data processing method provided by an embodiment of the present application;
[0059] Figure 2 It is a schematic diagram of a distributed task scheduling framework provided by an embodiment of the present application;
[0060] Figure 3Schematic diagram of another distributed task scheduling framework provided by an embodiment of the present application;
[0061] Figure 4 Flow schematic diagram of another distributed batch data processing method provided by an embodiment of the present application;
[0062] Figure 5 Structural block diagram of a distributed batch data processing device provided by an embodiment of the present application;
[0063] Figure 6 Structural block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0064] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature rather than restrictive.
[0065] Figure 1 The flowchart showing a distributed batch data processing method according to an embodiment of the present application is as follows Figure 1 As shown, the method may include the following steps:
[0066] S110. Configure a timing task list for a large amount of data to be processed in the target business scenario according to the business data processing requirements of the target business scenario.
[0067] In one implementation manner, the target business scenario may be any business scenario that requires data processing.
[0068] In one implementation manner, the timing task list may include but is not limited to: several timing tasks and their corresponding timing task tags. It can be understood that one timing task corresponds to one timing task tag.
[0069] In one implementation manner, several processing dimensions of the large amount of data may be determined according to the business data processing requirements of the target business scenario; and then the timing task list may be configured according to the several processing dimensions. It can be understood that one processing dimension corresponds to one timing task.
[0070] Exemplarily, taking the target business scenario as the order business scenario, and taking the data volume of the large amount of data to be processed in this business scenario as 50 million as an example, the large amount of data to be processed in this business scenario can be as shown in Table 1 below. In specific implementation, according to the business data processing requirements of this business scenario, N processing dimensions of the large amount of data to be processed can be determined, such as processing dimensions like cities, customer levels, prices, item classifications, etc., and then corresponding scheduled tasks can be configured according to these processing dimensions to facilitate the analysis of the sales trend of this business scenario.
[0071] Table 1
[0072] Order Identification Order Name Order Code Order Classification … Customer 1 xxxx xxxx xxxx xxxx xxxx 2 xxxx xxxx xxxx xxxx xxxx 3 xxxx xxxx xxxx xxxx xxxx 4 xxxx xxxx xxxx xxxx xxxx … … … … … … 5kw xxxx xxxx xxxx xxxx xxxx
[0073] In specific implementation, according to the business data processing requirements of the target business scenario, several scheduled tasks (TASK) for the large amount of data to be processed in the target business scenario can be configured in the scheduled task configuration center, and then a scheduled task label can be added to any scheduled task to obtain a scheduled task list.
[0074] In this application, according to the business data processing requirements of the target business scenario, configuring the scheduled task list for the large amount of data to be processed in the target business scenario can be driven by the business data processing requirements, customize the scheduled tasks for data processing, facilitate meeting the corresponding business requirements of the target business scenario, and also support specifying a scheduled task label for the scheduled task.
[0075] S120. Generate a scheduled task instance list for any scheduled task according to the business processing logic of the target business scenario and the scheduled task list.
[0076] In one implementation manner, the scheduled task instance list may include, but is not limited to: several scheduled task instances and their corresponding scheduled task labels, scheduled task instance parameters.
[0077] In one implementation manner, according to this business processing logic, the data volume segmentation unit of the large amount of data to be processed in the target business scenario can be determined; then, according to this data volume segmentation unit, the number of data sets that can be obtained after the large amount of data is segmented can be determined; finally, according to this number of data sets and the scheduled task list, a scheduled task instance list for any scheduled task is generated.
[0078] Exemplarily, if the data volume of this large amount of data is 100 million and the data is split by 10 million, then any scheduled task can correspond to 10 scheduled task instances. To improve the data analysis efficiency, it can also be split by a smaller data volume, split by 1 million, and any scheduled task can then correspond to 100 task instances.
[0079] As an example, combined with the foregoing order business scenario and as shown in Table 1, the list of scheduled task instances can be as shown in Table 2 below.
[0080] Table 2
[0081]
[0082] In the present application, by generating a list of scheduled task instances for any scheduled task according to the business processing logic of the target business scenario and the list of scheduled tasks, it is possible to facilitate the subsequent splitting of a large amount of data into N data sets, which helps to improve the efficiency of data processing. It can also support the subsequent horizontal expansion of the executor according to the scheduled task tags, that is, the more scheduled task instances, the more executors can be deployed.
[0083] S130. Deploy a task execution cluster according to the list of scheduled task instances.
[0084] In one implementation, the task execution cluster includes a number of executors (execution hosts).
[0085] In one implementation, according to the list of scheduled task instances, a number of application instances can be instantiated using a containerized deployment tool, and thus the task execution cluster can be obtained. Among them, one application instance is one executor.
[0086] Exemplarily, if the amount of the large amount of data is 100 million and the data is split by 10 million, then any scheduled task can correspond to 10 scheduled task instances. At this time, 3 application instances (executors) can be instantiated. If the data is split by 1 million, any scheduled task can correspond to 100 task instances. At this time, 30 application instances (executors) can be instantiated. When multiple scheduled task instances are processed concurrently by the executor, they can be processed faster.
[0087] As an example, the project name of the corresponding required service of the target business scenario can be configured (for example, Order_Analysis_Excuter), and then according to the project name, the corresponding required service can be deployed and managed using a containerized deployment tool (such as Kubernetes), and the project name can also be used as the tag of the executor.
[0088] It can be understood that the task execution cluster is an executor application with the same service name composed of N container instances, and one executor application is one service instance.
[0089] As an example, the corresponding required service may be associated with the above processing dimension. For example, when the processing dimension is the customer level dimension, the corresponding required service may be the customer data analysis service.
[0090] In this application, by using a containerized deployment tool to deploy a task execution cluster, it is possible to support quickly and dynamically adjusting (scaling up or down) the number of executors horizontally without the need to manually log in to the host for operation, which is convenient and fast. For example, during the peak period of a business scenario, 10 executors are deployed. After the peak period, the 10 executors can be dynamically adjusted to 3 executors through the containerized deployment tool.
[0091] During specific implementation, the executors can be automatically registered into the task scheduling cluster of the task scheduling center. As an example, as Figure 2 shown, the task scheduling cluster can include several scheduling threads. Subsequently, the scheduled tasks can be scheduled to different executors through these several scheduling threads.
[0092] Exemplarily, taking the task execution cluster including 3 executors as an example, after these 3 executors are registered into the task scheduling cluster, the 3 executors received by the task scheduling cluster can be as shown in Table 3 below.
[0093] Table 3
[0094] Actuator Identification Actuator Label Actuator Access Address 1 Order_Analysis_Excuter 192.168.0.1 2 Order_Analysis_Excuter 192.168.0.2 3 Order_Analysis_Excuter 192.168.0.3
[0095] In this application, by deploying the task execution cluster according to the scheduled task instance list, it is possible to support the horizontal expansion of the executors according to the scheduled task tags. For example, it is possible to support horizontally expanding the task execution cluster of a certain peak-period business scenario according to the scheduled task tags. The scheduled task tags can be reasonably allocated according to the current business processing logic, such as quickly responding to the needs of the business scenario and efficiently and quickly processing business data. After this business scenario passes the peak period, some executors can be dynamically recycled, other scheduled task tags can be added, and other executors can be deployed, thereby ensuring the maximum utilization of the executors and the most efficient processing of business data, and avoiding resource waste.
[0096] S140. Generate scheduled task configuration information according to the scheduled task list and the task execution cluster.
[0097] In one implementation manner, the scheduled task configuration information may include, but is not limited to, the association relationship between the executors and the scheduled task tags.
[0098] In one implementation manner, the execution period, number, and task name can be configured for any scheduled task according to the scheduled task list, and then the executors used to execute any scheduled task can be selected from the task execution cluster, and thus the scheduled task configuration information can be generated.
[0099] As an example, in combination with Tables 1-3 above and Figure 2 - Figure 3 shown, the scheduled task configuration information can be as shown in Table 4 below.
[0100] Table 4
[0101]
[0102] In this application, by generating timing task configuration information according to the timing task list and the task execution cluster, it is possible to support scheduling an executor that matches the timing task label to execute the timing task according to the execution period of the timing task subsequently.
[0103] S150. According to the timing task configuration information and the timing task label, assign any timing task instance to the corresponding executor in the task execution cluster.
[0104] As an example, as shown in Table 2 - Table 4 above, through the task scheduling center, it is possible to determine the executor label for the customer data analysis timing task configuration according to the timing task label such as 1000, for example, Order_Analysis_Excuter, and then assign the timing task instance to 3 executors for execution.
[0105] In this application, by assigning any timing task instance to the corresponding executor in the task execution cluster according to the timing task configuration information and the timing task label, it is possible to schedule an executor that matches the timing task label to execute the timing task.
[0106] S160. According to the timing task instance parameters, use the task execution cluster to perform modulo operation rules to batch - process a large amount of data, summarize the respective processing results, and then store them in the database.
[0107] In one implementation, it is possible to obtain the service data corresponding to any timing task instance according to the timing task instance parameters, where the service data corresponding to all timing task instances constitutes a large amount of data; then use the executor corresponding to any timing task instance to initialize a number of sub - threads according to the modulo operation rules; finally, batch - process the service data through a number of sub - threads.
[0108] As an example, the executor corresponding to any timing task instance can execute the assigned timing task according to the execution period of the timing task. The executor corresponding to any timing task instance can obtain the service data corresponding to any timing task instance according to the timing task instance parameters; then according to the modulo operation rules (such as M % N = K, where M is the data volume of the service data, N takes the value of 10, and K is an integer), divide the service data into 10 batches of data. Then, the executor corresponding to any timing task instance can initialize 10 sub - threads in the thread pool, where each sub - thread can be default - assigned an ID, allocated from 1 to 10. The executor corresponding to any timing task instance then batch - processes a number of batches of data through a number of sub - threads, where each sub - thread processes the service data whose primary key modulo is equal to the current sub - thread ID. It can be understood that one sub - thread processes one batch of data.
[0109] Exemplarily, in combination with the foregoing order business scenario, Table 1-4, and Figure 2 - 4 as shown, the scheduled task instances are randomly assigned to 3 executors (Order_Analysis_Excuter) for execution. Taking the scheduled task instance identifier 1 as an example, the executor (Order_Analysis_Excuter) obtains the scheduled task instance list, parses the scheduled task instance parameters (in json format), and retrieves the data with the order identifier between 1 and 1kw (including 1 and 1kw) from the order data table (order); the executor (Order_Analysis_Excuter) distributes the business data to the sub-threads for processing according to the modulo operation rule (M%N=K, here modulo is taken by 10 and processed by 10 sub-threads). At this time, the data distribution algorithm and the data processing process are as follows:
[0110] Initialize 10 sub-threads through the thread pool. Each sub-thread identifier ranges from 1 to 10, and the instances of each current sub-thread (Thread) are stored through a List object threadList; create a public Hash object resultMap to store the processing results of each sub-thread in a K-V manner;
[0111] The main thread starts taking the modulo by the order identifier % 10. If the result is 1, put the data into a LIST collection (list1); if the result is 2, put the data into a LIST collection (list2); and so on. The main thread divides the 1kw data into 10 LIST collections (list1 to 10);
[0112] The main thread gives list1 to the sub-thread with the sub-thread identifier 1 for processing, gives list2 to the sub-thread with the sub-thread identifier 2 for processing, and so on, and distributes the data set to be processed to the corresponding sub-threads for processing;
[0113] The sub-thread starts processing the distributed data set, and stores the result in the public Hash object resultMap. K stores the sub-thread identifier, and V stores the result (T) of the sub-thread data processing.
[0114] In this application, by taking the modulo operation rule according to the timed task instance parameters and using the task execution cluster to batch process a large amount of data, data can be isolated and processed at the physical level through the configuration of task instance parameters, avoiding the scenario of deadlocks caused by multiple threads fetching data simultaneously. At the same time, it also supports an algorithm mechanism for flexible segmentation and processing of business data through the modulo operation rule (M % N = K). The thread pool size is initialized according to N, and the threads match through the modulo result K to process the data that matches the modulo result, further avoiding the possible deadlock scenario of multiple threads fetching data, thereby ensuring the security and stability of data processing to the greatest extent and avoiding the failure of single-machine hardware resource overflow.
[0115] In one implementation, when it is monitored that a number of child threads have all completed data processing, the processing results of the number of threads are collected. For example, the main thread of the executor monitors the status of the child threads. When the status of all child threads is the dead state (i.e., the data processing completion state), the processing result of each child thread is collected; the processing results of a number of child threads are summarized (e.g., merged), and then the summarized result is stored in the database.
[0116] As an example, combined with the foregoing data distribution algorithm and data processing process and Figure 3 - 4 As shown, the main thread parses the List object threadList to store the instances of each current child thread (Thread), judges the status of the child thread (Thread.State). When the status (Thread.State) of each child thread is TERMINATED, it means that the data sets allocated to all current child threads have been processed; then traverse the public Hash object resultMap, obtain V (the result (T) of the child thread data processing) by traversing K (child thread identifier), and finally summarize the processing results and store them in the database.
[0117] Combined with Figure 3 - 4 As shown, the timed task configuration center supports users to formulate tasks for timed processing of business data according to business scenario requirements; the task scheduling cluster supports distributing tasks to idle machines for execution according to certain strategies based on the timed tasks configured in the task configuration center; the task execution cluster supports executing the tasks assigned by the task scheduling cluster and records the final execution results. That is, the task scheduling center reads the timed task set from the database, distributes the timed tasks to the task execution cluster that matches the label according to the timed task label. The execution class sets the modulo rule (M % N = K) according to the data volume size, initializes the thread pool (set according to N), processes business data with multiple threads, the main thread records the status of multiple threads, and after the status of all threads is the dead state (the thread processing is completed), collects the processing results of each child thread, summarizes the results, and records the results.
[0118] That is, this application supports a distributed task scheduling framework implemented based on Spring Boot. Through the task scheduling center, it supports task configuration and task monitoring. The framework is mainstream and is quick and convenient to learn and use.
[0119] As can be seen from the above description, this application can be driven by business data processing requirements to customize timed tasks for data processing, and support specifying timed task tags for timed tasks, facilitating the meeting of corresponding business requirements in the target business scenario; it can also support horizontal expansion and contraction of executors in the task execution cluster according to whether the business scenario is in the data peak period through timed task tags, maximizing the host usage efficiency, avoiding resource waste, and can also parallel process business data through the task execution cluster, splitting a large batch of data into several small data sets, then distributing the small data sets to multiple timed tasks for parallel processing, and finally summarizing the data processing results, thereby effectively improving the data processing efficiency and ensuring the stability of data processing, avoiding the failure of single-machine hardware resource overflow.
[0120] Figure 5 The structural block diagram of a distributed batch data processing device according to an embodiment of the present application is shown. As Figure 5 shown, the device may include:
[0121] A configuration unit 210, configured to configure a timed task list for a large batch of data to be processed in a target business scenario according to the business data processing requirements of the target business scenario, where the timed task list includes several timed tasks and their corresponding timed task tags;
[0122] A deployment unit 220, configured to generate a timed task instance list for any timed task according to the business processing logic and timed task list of the target business scenario, where the timed task instance list includes several timed task instances and their corresponding timed task tags and timed task instance parameters; deploy a task execution cluster according to the timed task instance list, where the task execution cluster includes several executors;
[0123] A processing unit 230, configured to generate timed task configuration information according to the timed task list and the task execution cluster, where the timed task configuration information includes the association relationship between the executor and the timed task tag; according to the timed task configuration information, assign any timed task instance to the corresponding executor in the task execution cluster according to the timed task tag; according to the timed task instance parameters, use the task execution cluster to batch process a large batch of data by adopting a modulo operation rule, and summarize each processing result and then store it in the database.
[0124] In one implementation manner, the configuration unit 210 is specifically configured to:
[0125] Determine several processing dimensions for a large volume of data according to the business data processing requirements;
[0126] Configure a list of scheduled tasks according to several processing dimensions.
[0127] In one implementation, the deployment unit 220 is specifically configured to:
[0128] Determine the data volume segmentation unit for a large volume of data according to the business processing logic;
[0129] Determine the number of data sets that can be obtained after the large volume of data is segmented according to the data volume segmentation unit;
[0130] Generate a list of scheduled task instances for any scheduled task according to the number of data sets and the list of scheduled tasks.
[0131] In one implementation, the deployment unit 220 is specifically configured to:
[0132] Instantiate several application instances using a containerized deployment tool according to the list of scheduled task instances to obtain a task execution cluster, where one application instance is an executor.
[0133] In one implementation, the processing unit 230 is specifically configured to:
[0134] Configure the execution period, number, and task name for any scheduled task according to the list of scheduled tasks, and then select an executor from the task execution cluster for executing any scheduled task to generate scheduled task configuration information.
[0135] In one implementation, the processing unit 230 is specifically configured to:
[0136] Obtain the business data corresponding to any scheduled task instance according to the scheduled task instance parameters, where the business data corresponding to all scheduled task instances constitutes a large volume of data;
[0137] Initialize several sub-threads according to the modulo operation rule using the executor corresponding to any scheduled task instance;
[0138] Process the business data in batches through several sub-threads.
[0139] In one implementation, the processing unit 230 is specifically configured to:
[0140] When it is monitored that all several sub-threads have completed data processing, collect the processing results of the several sub-threads;
[0141] Summarize the processing results of the several sub-threads, and then store the summary result in the database.
[0142] In the embodiments of the present application, the functions of the units of the distributed batch data processing device can be referred to the corresponding descriptions in the above methods, which will not be elaborated here.
[0143] Figure 6 The structural block diagram of a computer device according to an embodiment of the present application is shown. As Figure 6 shown, the computer device includes: a memory 310 and a processor 320. Instructions are stored in the memory 310, and the instructions are loaded and executed by the processor 320 to implement the distributed batch data processing method in the above embodiments. The number of the memory 310 and the processor 320 can be one or more.
[0144] The computer device further includes:
[0145] a communication interface 330, configured to communicate with external devices and perform data interaction and transmission.
[0146] If the memory 310, the processor 320, and the communication interface 330 are implemented independently, the memory 310, the processor 320, and the communication interface 330 can be interconnected through a bus and complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0147] Optionally, in a specific implementation, if the memory 310, the processor 320, and the communication interface 330 are integrated on a chip, the memory 310, the processor 320, and the communication interface 330 can complete communication with each other through an internal interface.
[0148] The embodiments of the present application provide a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the method provided in the embodiments of the present application is implemented.
[0149] The embodiments of the present application further provide a chip, which includes a processor, configured to call and run the instructions stored in the memory, so that a communication device installed with the chip executes the method provided in the embodiments of the present application.
[0150] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is configured to execute the code in the memory. When the code is executed, the processor is configured to execute the method provided by the embodiment of the application.
[0151] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced reduced instruction set machine (ARM) architecture.
[0152] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory, and may further include a non-volatile random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0153] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0154] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0155] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0156] Any process or method description represented in the flowchart or described in other ways herein can be understood as representing a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed.
[0157] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a defined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices.
[0158] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above method embodiments can be completed by a program instructing relevant hardware, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.
[0159] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0160] As mentioned above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various changes or substitutions, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A distributed batch data processing method, characterized in that: include: According to the business data processing requirements of the target business scenario, a scheduled task list of large batches of data to be processed in the target business scenario is configured, wherein the scheduled task list includes a number of scheduled tasks and their corresponding scheduled task tags; Generate a scheduled task instance list of any scheduled task according to the business processing logic of the target business scenario and the scheduled task list, wherein the scheduled task instance list includes a number of scheduled task instances and their corresponding scheduled task tags and scheduled task instance parameters; Deploy a task execution cluster according to the scheduled task instance list, wherein the task execution cluster includes a plurality of executors; Generate scheduled task configuration information according to the scheduled task list and the task execution cluster, wherein the scheduled task configuration information includes an association relationship between an executor and a scheduled task tag; According to the scheduled task configuration information and the scheduled task label, allocating any of the scheduled task instances to the corresponding executor in the task execution cluster; According to the scheduled task instance parameters, the task execution cluster is used to process the large batches of data in batches using a modulo operation rule, and the various processing results are summarized and stored in a database.
2. The method according to claim 1, characterized in that According to the business data processing requirements of the target business scenario, the scheduled task list for the large batch of data to be processed in the target business scenario includes: Determining several processing dimensions of the large batch of data according to the business data processing requirements; The scheduled task list is configured according to a plurality of the processing dimensions.
3. The method according to claim 1, characterized in that Generating a scheduled task instance list of any scheduled task according to the business processing logic of the target business scenario and the scheduled task list includes: Determining a data volume segmentation unit for the large batch of data according to the business processing logic; Determining the number of data sets that can be obtained after the large batch of data is segmented according to the data volume segmentation unit; A scheduled task instance list of any scheduled task is generated according to the number of data sets and the scheduled task list.
4. The method according to claim 1, characterized in that: Deploying a task execution cluster according to the scheduled task instance list includes: According to the scheduled task instance list, a plurality of application instances are instantiated using a containerized deployment tool to obtain the task execution cluster, wherein one application instance is one executor.
5. The method according to claim 1, characterized in that Generating scheduled task configuration information according to the scheduled task list and the task execution cluster includes: According to the scheduled task list, an execution period, a number and a task name are configured for any of the scheduled tasks, and then an executor for executing any of the scheduled tasks is selected from the task execution cluster to generate the scheduled task configuration information.
6. The method according to any one of claims 1 to 5, characterized in that: According to the scheduled task instance parameters, using the task execution cluster to adopt a modulo operation rule to process the large batch of data in batches includes: According to the scheduled task instance parameters, obtain the business data corresponding to any of the scheduled task instances, wherein the business data corresponding to all the scheduled task instances constitute the bulk data; Using the executor corresponding to any of the scheduled task instances, initialize a number of sub-threads according to the modulo operation rule; The business data is processed in batches through a plurality of sub-threads.
7. The method according to claim 6, characterized in that Summarizing various processing results and storing them in the database includes: When monitoring that the plurality of sub-threads have completed data processing, collecting processing results of the plurality of threads; The processing results of the plurality of sub-threads are summarized, and the summarized results are stored in the database.
8. A distributed batch data processing device, characterized in that: include: A configuration unit, configured to configure a scheduled task list of large batches of data to be processed in a target business scenario according to a business data processing requirement of the target business scenario, wherein the scheduled task list includes a plurality of scheduled tasks and their corresponding scheduled task tags; A deployment unit, configured to generate a scheduled task instance list of any scheduled task according to the business processing logic of the target business scenario and the scheduled task list, wherein the scheduled task instance list includes a plurality of scheduled task instances and their corresponding scheduled task tags and scheduled task instance parameters; and deploy a task execution cluster according to the scheduled task instance list, wherein the task execution cluster includes a plurality of executors; A processing unit is used to generate scheduled task configuration information according to the scheduled task list and the task execution cluster, wherein the scheduled task configuration information includes the association relationship between the executor and the scheduled task label; according to the scheduled task configuration information, any of the scheduled task instances is assigned to the corresponding executor in the task execution cluster according to the scheduled task label; according to the scheduled task instance parameters, the task execution cluster is used to adopt the modulus operation rule to batch process the large batch data, and the various processing results are summarized and stored in the database.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory stores instructions, and the instructions are loaded and executed by the processor to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed on a computer, the method according to any one of claims 1 to 7 is implemented.