Data batch processing method, device, equipment and computer storage medium
By determining ready batches in the big data platform and using collection system resources and tokens for processing, the problem of inefficient batch execution of resource data in the existing technology is solved, and efficient utilization of system resources and flexible adjustment of collection risk rules are achieved.
Patent Information
- Application Number
- CN201911246706.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-05
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2039-12-05
AI Technical Summary
In the prior art, data processing systems in the financial industry cannot fully utilize system resources when processing multiple resource data, resulting in inefficient batch execution processing of resource data and the inflexible control of the order and concurrency of resource data batches.
By obtaining ready batches in the big data platform, determining the optimal batches, and using the system resources of the collection system and available tokens for execution until all batches are completed, optimizing collection risk rules in combination with the identifiable policy package of the policy engine.
It improves the efficiency of batch execution and processing of resource data in the collection system, ensures the rational use of system resources, and supports flexible adjustment of collection risk rules.
Smart Images

Figure CN111008078B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial technology (Fintech), and in particular to a method, apparatus, device and computer storage medium for batch processing of data. Background Art
[0002] With the development of computer technology, more and more technologies are being applied in the financial sector. The traditional financial industry is gradually transforming into financial technology (Fintech), and data processing technology in big data is no exception. However, the security and real-time requirements of the financial industry also place higher demands on technology. For example, traditional data processing technology currently processes customer lists and related target data according to fixed rules for non-expired reminders and overdue processing access rules. Based on the target data, the system calls a policy engine, which implements case management based on the risk level of the case. In addition, the current system processes target data at the resource data level. When multiple resource data are processed simultaneously in a system, they are mostly processed serially, which fails to fully utilize system resources and cannot conveniently and efficiently control the order and concurrency of batch execution of each resource data. Therefore, how to improve the processing efficiency of batch execution of resource data in the system has become a technical problem that needs to be solved urgently. Summary of the Invention
[0003] The main purpose of the present invention is to provide a data batch processing method, device, equipment and computer storage medium, aiming to improve the processing efficiency of batch execution processing of resource data in the system.
[0004] To achieve the above object, the present invention provides a method for batch processing of data, which comprises the following steps:
[0005] Obtaining batches corresponding to different types of resource data in the big data platform and determining whether each batch is in a ready state;
[0006] If there is a ready batch in a ready state in each of the batches, and there are multiple ready batches, then the number of cases in each of the ready batches and the system resources of the collection system are obtained, and the optimal batch is determined based on the system resources and the number of cases in each of the ready batches;
[0007] Obtain available tokens in the collection system, execute processing on the optimal batch based on the available tokens, and after the optimal batch processing is completed, continue to execute the step of obtaining batches of multiple resource data in big data until each batch processing is completed.
[0008] Optionally, the step of obtaining system resources of the debt collection system includes:
[0009] A first batch that is completed first and in a ready state is obtained based on each batch, and execution processing is performed on the first batch. System resources of the collection system are determined based on the execution processing result of the first batch.
[0010] Optionally, the step of determining the optimal batch size based on the system resources and the number of cases in each of the ready batches includes:
[0011] Calculate the usage growth rate of the system resources and determine whether the usage growth rate is less than a preset alarm threshold,
[0012] If it is less than, the optimal batch is determined based on the number of cases in each of the ready batches.
[0013] Optionally, the step of calculating the usage growth rate of the system resources includes:
[0014] The positive relationship between the increase in usage of the system resources and the increase in the number of cases in the first batch is calculated based on the execution processing result of the first batch, and the usage growth rate of the system resources is calculated according to the positive relationship.
[0015] Optionally, the step of determining the optimal batch size based on the number of cases in each of the ready batch sizes includes:
[0016] Calculating an available resource value of the system resources according to the alarm threshold, and calculating a maximum target case load that can be admitted to the system resources according to the available resource value;
[0017] Determining a case level corresponding to each of the ready batches according to the number of cases in each of the ready batches, and determining a target level in each of the case levels based on the target case volume;
[0018] A target ready batch at a target level is obtained from each of the ready batches. If there are multiple target ready batches, a target ready batch with the largest number of cases is obtained from each of the target ready batches, and the target ready batch with the largest number of cases is used as the optimal batch.
[0019] Optionally, the step of obtaining available tokens in the collection system and executing the optimal batch based on the available tokens includes:
[0020] Determining a target number of tokens when the optimal batch is to be executed, obtaining an available token number of available tokens in the collection system, and determining whether the target number of tokens is greater than the available number of tokens;
[0021] If it is less than or equal to, executing the optimal batch according to the available tokens.
[0022] Optionally, after the step of determining whether the target number of tokens is greater than the number of available tokens, the method further comprises:
[0023] If it is greater, new tokens generated by the collection system based on a preset rate are obtained until the sum of the number of new tokens and the number of available tokens is equal to the target number of tokens, and the optimal batch is executed based on the available tokens and the new tokens.
[0024] In addition, to achieve the above-mentioned purpose, the present invention further provides a data batch processing device, the data batch processing device comprising:
[0025] an acquisition module for acquiring batches corresponding to different types of resource data in the big data platform and determining whether each batch is in a ready state, wherein the big data platform acquires the batches based on a recognizable policy package in the policy engine, wherein the recognizable policy package includes collection risk rules extracted from the big data processing logic based on a preset algorithm;
[0026] a determination module configured to obtain the number of cases in each batch and the system resources of the collection system if all the batches are in a ready state, and determine the optimal batch according to the system resources and the number of cases in each batch;
[0027] A processing module is used to obtain available tokens in the collection system, execute processing on the optimal batch based on the available tokens, and after the optimal batch processing is completed, continue to execute the steps of obtaining batches corresponding to different types of resource data in the big data platform until each batch processing is completed.
[0028] In addition, to achieve the above-mentioned purpose, the present invention also provides a data batch processing device, which includes: a memory, a processor, and a data batch processing program stored in the memory and capable of running on the processor, and when the data batch processing program is executed by the processor, the steps of the data batch processing method described above are implemented.
[0029] In addition, to achieve the above objectives, the present invention also provides a computer storage medium, on which a data batch processing program is stored. When the data batch processing program is executed by a processor, the steps of the data batch processing method described above are implemented.
[0030] The present invention obtains batches corresponding to different types of resource data in a big data platform and determines whether each of the batches is in a ready state; if there is a ready batch in a ready state in each of the batches, and there are multiple ready batches, then the number of cases of each ready batch and the system resources of the collection system are obtained, and the optimal batch is determined based on the system resources and the number of cases of each ready batch; the available tokens in the collection system are obtained, and the optimal batch is executed based on the available tokens. After the optimal batch processing is completed, the steps of obtaining batches corresponding to different types of resource data in the big data platform are continued until the processing of each batch is completed. By obtaining batches corresponding to resource data in a big data platform, and when each batch is in a ready state, the optimal batch is determined based on the system resources of the collection system and the number of cases of each ready batch, and the optimal batch is processed based on the available tokens until all batches are processed, it is ensured that the collection system obtains the usage of system resources in real time, and processes the optimal batch in combination with the tokens, thereby improving the efficiency of batch execution of resource data in the collection system. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention;
[0032] Figure 2 This is a flow chart of a first embodiment of a method for batch processing of data according to the present invention;
[0033] Figure 3 A schematic diagram of a device module of a device for batch processing of data according to the present invention;
[0034] Figure 4 Schematic diagram of batch processing in the batch processing method of data of the present invention;
[0035] Figure 5 Schematic diagram of the processing of the policy engine in the batch processing method of data of the present invention;
[0036] Figure 6 The data structure of resource data in the batch processing method of data of the present invention;
[0037] Figure 7 Schematic diagram of the batch processing flow based on the token algorithm in the batch processing method of data of the present invention;
[0038] Figure 8 A schematic diagram of batch processing of multiple resource data in the batch processing method of data of the present invention;
[0039] Figure 9 Schematic diagram of the process of parallel batch processing in the batch processing method of data of the present invention;
[0040] Figure 10 This is a batch state flow diagram in the batch processing method of data of the present invention.
[0041] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0042] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0043] like Figure 1 As shown, Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention.
[0044] The batch processing device for data in the embodiment of the present invention may be a PC or a server device, on which a Java virtual machine runs.
[0045] like Figure 1 As shown, the batch processing device of the data may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0046] Those skilled in the art will understand that Figure 1 The device structure shown in the figure does not constitute a limitation of the device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0047] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a data batch processing program.
[0048] exist Figure 1In the device shown, the network interface 1004 is mainly used to connect to the background server and communicate data with the background server; the user interface 1003 is mainly used to connect to the client (user end) and communicate data with the client; and the processor 1001 can be used to call the batch processing program of the data stored in the memory 1005 and perform the operations in the batch processing method of the following data.
[0049] Based on the above hardware structure, an embodiment of a method for batch processing of data of the present invention is proposed.
[0050] Reference Figure 2 , Figure 2 This is a flow chart of a first embodiment of a method for batch processing of data according to the present invention, the method comprising:
[0051] Step S10: obtaining batches corresponding to different types of resource data in the big data platform, and determining whether each batch is in a ready state;
[0052] In this embodiment, the strategy engine, centered around technologies such as decision trees, champion challengers, and neural networks, is a tool capable of performing real-time prediction and analysis of big data, performing scoring modeling, risk identification, and intelligent decision-making based on specific business rules. Big data processing is the process of extracting, transforming, calculating, and indexing big data through a big data platform. The platform retrieves the batches based on identifiable strategy packages within the strategy engine. These identifiable strategy packages include collection risk rules extracted from the big data processing logic based on a pre-set algorithm.
[0053] At present, the access rules for judging whether users need to be reminded in advance or to initiate collection are judged in the big data processing logic. For example, based on the customer's repayment date, an advance reminder is made on T-3 day. However, there are two problems with this: (a) it targets all customers, causing harassment to customers who do not need reminders; (b) because it targets all customers, it is necessary to send text messages and voice messages to all customers, which is costly and wastes a lot of resources. The purpose of collection is to notify customers as early as possible within a certain human resource and collection cost, predict the risk of customers turning into bad debts, and reduce the occurrence of bad debts. Since risks themselves are difficult to predict, complex and changeable, hard-coding this risk logic and coupling it to the big data processing data retrieval script cannot meet business needs. For example, Figure 4As shown, different resource data are aggregated within the big data platform, including auto loan accounts, IOUs, mortgage accounts, IOUs, and credit loan data. Specifically, this data is aggregated into the auto loan source layer, the mortgage source layer, and the credit loan source layer. This data is then processed, including big data-derived variable processing, a big data execution strategy engine, and collection data processing. This collection data processing involves acquiring collection information from collection documents and enabling collection robots to perform collection based on the new collection documents generated through big data processing. Collection information also includes industry-wide collection activities, including batch resource data, application services, and databases. Therefore, in this embodiment, by using the preset strategy development software, creating strategy input and output variables, formulating scoring cards, decision trees, decision tables, champion challengers and other strategies, the complex and changeable debt collection risk rules are extracted from the Hive (data warehouse tool) processing logic and maintained in the strategy engine. The strategy is converted into an identifiable strategy package (such as a SER package) that can be recognized by the application, and placed in the HDFS path of the big data platform for HQL to call. The interface for calling the SER package is converted into a jar package through a custom UDF function and placed in the big data platform to realize the function of Hive statements directly calling the strategy engine. For example, Figure 5 As shown, first, the data in the Hive table is converted into the policy engine input transmission object, and then the input transmission object is converted into the built-in input parameter of the policy engine. The policy engine is executed to obtain the corresponding policy output and store it in the database.
[0054] To fully understand, analyze, and predict customer risks, we process relevant lending data indicators for customers who have not yet registered a default, in addition to those with normal overdue payments. This allows us to predict their overdue risk in advance. For example, we process data from a customer's 24 historical billing cycles to obtain indicators such as first overdue payment and first payment, as well as risk indicators to build a customer behavior scoring model. This model uses characteristic variables from different dimensions of the user to determine the customer's future overdue risk.
[0055] In this embodiment, different resource data have different data structures. For example, car loans have special information such as mileage and driver's license, but the basic information is basically the same. Therefore, in this embodiment, the data structure is divided into common basic information and data structures derived according to the characteristics of different resource data. Figure 6As shown, case data for resource data includes customer base data, account data, IOU data, repayment plan data, transaction history data, and contact information data. The unique and differentiated data for each resource data type is maintained in separate tables. Furthermore, to logically isolate data, collection batches are organized by institution and resource data type. After the big data platform organizes different types of resource data into tables and sends them to the collection system, the collection system determines the corresponding batch based on the individual resource data types, with one batch corresponding to each type of resource data. The collection system then determines whether all received batches are in a ready state. Different actions are then performed based on the different determination results.
[0056] Step S20: If there is a ready batch in the ready state among the batches, and there are multiple ready batches, the number of cases in each ready batch and the system resources of the collection system are obtained, and the optimal batch is determined based on the system resources and the number of cases in each ready batch;
[0057] When it is determined that there are ready batches in the ready state in each batch, it is necessary to determine whether there are multiple ready batches. If there are multiple, the first ready batch (i.e. the first batch) needs to be executed, and the current system resources of the collection system will be monitored at this time to provide a reference for subsequent parallel batch scheduling. And because parallel processing can provide system resource utilization and improve the timing of batch running (where batch running is batch processing with multiple reads, multiple calculation processes, and one write), but too high concurrency may cause data loss, and system resources are limited, so in this example, the system resources of the collection system and their utilization rate will be comprehensively considered. The system resources include memory, server CPU (i.e. virtual machine CPU), database CPU and database IO, etc. For example, as shown in the table below, the 4G memory and 6-core CPU allocated to the server for the collection system are used as an example.
[0058]
[0059] Due to system resource depletion and changes in system resources caused by other applications, you can choose to obtain the average system resource usage over a preset time period. Furthermore, to prevent system crashes due to excessive load, a unified alarm threshold (maxAlarmRate), such as 80%, is set for each system resource. Only when the utilization of a particular system resource exceeds the alarm threshold will the system stop acquiring new batches for execution. Furthermore, when the utilization of a system resource falls below the alarm threshold, the system selects the optimal batch from among the ready batches. The optimal batch selection method first calculates the available system resources. For example, the current available memory (remainMemoryRate) = alarm threshold - current used memory rate; the current available CPU (remainVMCpuRate) = alarm threshold - current used CPU rate; the current available database CPU (remainDBCpuRate) = alarm threshold - current used database CPU rate; and the current available database (remainIORate) = alarm threshold - current used database I / O rate. Based on the available system resources, the maximum number of cases that can be admitted to the system resources (maxCaseCount) is calculated, i.e., maxCaseCount = the minimum value of Min(remainMemoryRate / perMemoryRate, remainingVMCpuRate / perVMCpuRate, remainingDBCpuRate / perDBCpuRate, remainingIORate / perIORate). The ready batches are then sorted by case volume for each resource. If the case volumes are equal, the batch corresponding to the resource with the largest case volume is added to the running queue (this is the optimal batch).
[0060] Step S30, obtain the available tokens in the collection system, execute the optimal batch based on the available tokens, and after the optimal batch processing is completed, continue to execute the steps of obtaining batches corresponding to different types of resource data in the big data platform until each batch processing is completed.
[0061] After obtaining the optimal batch, it is also necessary to obtain the available tokens in the collection system, and execute the optimal batch according to the available tokens. After the optimal batch processing is completed, continue to execute the above steps of obtaining batches of multiple resource data in the big data until each batch processing is completed. Among them, the way to obtain available tokens is that the collection system first produces tokens at a preset rate, where the time for each token to be generated is 1 / rate, such as 1 token per minute. In addition, in order to limit the number of parallel batches, if the generated tokens have not been used, when the number of tokens in the token bucket reaches a certain value, the first generated unused token will be discarded. For example, only the number of tokens generated within 5 minutes is retained, then the maximum number of tokens available in the token bucket is 5, and the corresponding business logic indicates that the number of batches running at the same time is 5, and the current upper limit of data for parallel batch processing is 500,000. For example, Figure 7 As shown, the mortgage approval signal, car loan approval signal and mall installment approval signal are obtained in the big data platform, and the data is sent upstream to complete the batch readiness. That is, after the mortgage batch, car loan batch and mall batch in the collection system are in the ready state, token 1, token 2...token N are obtained, and the batch execution point is determined (only N tokens are allowed at the same time), that is, batch execution is performed according to the number of tokens.
[0062] In addition, to assist in understanding the principle of batch processing of resource data in this embodiment, an example is given below.
[0063] For example, Figure 8 As shown, when performing batch processing of multi-resource data, it includes a batch scheduling system, big data processing and a collection system. The batch scheduling system is used to start data processing and perform big data processing, such as car loans, mortgages and mall installment loans. It also determines the account to be collected, collects data, processes the processing completion status, etc., and sends it to the collection system to determine whether the resource data (i.e., product data) processing is completed? Whether the batch is ready, etc., wherein the batch readiness is determined by the batch scheduling system to determine whether the collection system starts to batch car loans, mortgages, and mall installment batches. In the collection system, a token is generated every ten minutes and placed in a token bucket. Car loan batches, mortgage batches, etc. are polled to obtain tokens to determine whether the tokens in the bucket are greater than the requested token number N. If not, continue to obtain tokens. If so, the number of tokens in the bucket is reduced by N, and a case is created or updated. The policy engine is executed, buttons are assigned, the case is sent to the partner, the batch completion signal is reported, and the batch completion is determined in the batch scheduling system. For another example, if Figure 9As shown, when multiple batches arrive at the collection system, they are centrally managed through the scheduling system of the scheduling platform. The scheduling system completes the daily cut (i.e., it is updated once a day). After the collection system completes the daily cut, all resource data for that day is in a waiting state. When the big data pushes the customer, account, IOU, transaction, contact information and other data of the resource data to the collection system, it sends a signal to the collection system to inform it that the push is complete. Whether the batch date of the transaction big data pushed by the collection system is equal to the system daily cut date, and after verifying that the amount and number of cases of the resource data are correct, the resource data batch is converted to a ready state. When the first ready batch runs, the optimal batch is selected from the ready queue based on the system resources, and an attempt is made to obtain a token for batch execution. When the token is obtained, the batch will be suspended until the suspension time ends, and then the batch will be run according to the token. After the batch run is completed, information is reported until it ends.
[0064] In this embodiment, batches corresponding to different types of resource data in the big data platform are obtained, and it is determined whether each of the batches is in a ready state; if there is a ready batch in a ready state in each of the batches, and there are multiple ready batches, the number of cases of each ready batch and the system resources of the collection system are obtained, and the optimal batch is determined based on the system resources and the number of cases of each ready batch; available tokens in the collection system are obtained, and the optimal batch is executed based on the available tokens. After the optimal batch processing is completed, the step of obtaining batches corresponding to different types of resource data in the big data platform is continued until the processing of each batch is completed. By obtaining batches corresponding to resource data in the big data platform, and when each batch is in a ready state, the optimal batch is determined based on the system resources of the collection system and the number of cases of each batch, and the optimal batch is processed based on the available tokens until all batches are processed, it is ensured that the collection system obtains the usage of system resources in real time, and processes the optimal batch in combination with the tokens, thereby improving the efficiency of batch execution of resource data in the collection system. In this embodiment, the batches in the big data platform are obtained based on the identifiable policy packages of the policy engine, and the identifiable policy packages are collection risk rules extracted from the big data processing logic based on a preset algorithm. That is, if the collection risk rules need to be modified, it is only necessary to directly adjust the identifiable policy package, avoiding the phenomenon in which the collection risk rules cannot be flexibly changed in the existing technology.
[0065] Furthermore, based on the first embodiment of the method for batch processing data of the present invention, a second embodiment of the method for batch processing data of the present invention is proposed. This embodiment is a refinement of step S10 of the first embodiment of the present invention, the step of obtaining system resources of the collection system, and includes:
[0066] Step a: obtaining the first batch that is completed and ready first based on each batch, and executing the first batch, and determining the system resources of the collection system based on the execution result of the first batch.
[0067] When all batches in the collection system are in the ready state, it is necessary to obtain the first batch that has completed the ready state first among all batches, and directly execute the first batch. Based on the execution processing results of this first batch, the system resources of the collection system are determined, such as the usage growth rate of system resources, remaining system resources, etc.
[0068] In this embodiment, the first batch that completes the ready state first is executed, and the system resources are determined according to the execution processing results, thereby ensuring the accuracy of obtaining the system resources at the current moment.
[0069] Furthermore, the step of determining the optimal batch size based on the system resources and the number of cases in each of the ready batches includes:
[0070] Step b: calculating the usage growth rate of the system resources and determining whether the usage growth rate is less than a preset alarm threshold.
[0071] After obtaining system resource information based on the first batch calculation, we can also calculate the positive relationship between system resource growth and caseload growth, that is, the rate of increase in system resource usage based on batch processing. To prevent excessive system resource usage from causing a system crash, we need to set a preset alarm threshold for system resources and determine whether the rate of increase in system resource usage is less than the preset alarm threshold. Different actions are then taken based on the different determination results.
[0072] Step c: if it is less than, then determine the optimal batch based on the number of cases in each of the ready batches.
[0073] When it is determined that the growth rate of system resource usage is less than the preset alarm threshold, it can be considered that the current system resources are sufficient for the batch system to execute the ready batches. Therefore, the optimal batch can be determined based on the number of cases in each ready batch.
[0074] In this embodiment, the growth rate of system resource usage is calculated, and when the growth rate is less than the alarm threshold, the optimal batch is determined based on the number of cases in each ready batch, thereby ensuring the normal operation of the batch system and avoiding the occurrence of system resource overload.
[0075] Specifically, the step of calculating the usage growth rate of the system resources includes:
[0076] Step d: calculating the positive relationship between the increase in the use of the system resources and the increase in the number of cases in the first batch based on the execution processing result of the first batch, and calculating the growth rate of the use of the system resources according to the positive relationship.
[0077] In this embodiment, detailed information of system resources is determined based on the execution processing results of the first batch, and based on this detailed information, the positive relationship between the increase in the use of system resources and the increase in the number of cases in the first batch is calculated, and the growth rate of the use of system resources is calculated based on this positive relationship.
[0078] In this embodiment, by determining the positive relationship between the growth in system resource usage and the increase in case volume, and then determining the growth rate of system resource usage based on the positive relationship, the accuracy of the obtained growth rate of system resource usage is ensured.
[0079] Specifically, the step of determining the optimal batch size based on the number of cases in each batch includes:
[0080] Step e: calculating the available resource value of the system resources according to the alarm threshold, and calculating the maximum admissible target case volume of the system resources according to the available resource value;
[0081] When the growth rate of system resource usage is less than the preset alarm threshold, the available resource value of the system resource needs to be calculated based on the alarm threshold, such as the current available memory (remainMemoryRate) = alarm threshold - current used memory rate; the current available CPU (remainVMCpuRate) = alarm threshold - current used CPU usage rate; the current available database CPU (remainDBCpuRate) = alarm threshold - current used database CPU usage rate; the current available database (remainIORate) = alarm threshold - current used database IO rate; and then calculate the maximum accessible target case volume of the system resource based on the available resource value, and compare the number of cases in each batch with the target case volume to determine the optimal batch.
[0082] Step m, determining a case level corresponding to each of the ready batches according to the number of cases in each of the ready batches, and determining a target level in each of the case levels based on the target case volume;
[0083] After calculating the maximum target caseload for system resources, we need to obtain the caseload counts for each ready batch and categorize each batch based on the caseload count to determine the corresponding caseload level. For example, 100,000-200,000 is a caseload level, and 200,000-500,000 is a caseload level. We then determine the caseload level within which the target caseload falls and use this level as the target level.
[0084] Step n: obtaining a target ready batch at a target level from each of the ready batches; if there are multiple target ready batches, obtaining a target ready batch with the largest number of cases from each of the target ready batches, and taking the target ready batch with the largest number of cases as the optimal batch.
[0085] Once the target level is determined, you need to determine which ready batches within each ready batch are at the target level and use them as the target ready batch. If there is only one target ready batch, this target ready batch can be used as the optimal batch. However, if there are multiple target ready batches, the number of cases in each target ready batch must be obtained and compared in sequence. The target ready batch with the largest number of cases is selected and used as the optimal batch.
[0086] In this embodiment, by calculating the available resource value of system resources based on the alarm threshold, determining the target case volume based on the available resource value, and determining the optimal batch based on the target case volume, it is ensured that the system resources are comprehensively considered when the batch system runs the batch, thereby improving the efficiency of the batch operation.
[0087] Furthermore, based on any one of the first to second embodiments of the data batch processing method of the present invention, a third embodiment of the data batch processing method of the present invention is proposed. This embodiment is a refinement of step S30 of the first embodiment of the present invention, which obtains available tokens in the collection system and executes processing on the optimal batch based on the available tokens, including:
[0088] Step f, determining a target number of tokens when the optimal batch is to be executed, obtaining the available token number of available tokens in the collection system, and determining whether the target number of tokens is greater than the available token number;
[0089] In the collection system, it is necessary to determine the target number of tokens required for the optimal batch to be executed, such as 1, 2, etc., and determine whether to execute the optimal batch immediately based on the number of available tokens in the collection system. That is, when the collection system obtains the optimal batch in the ready queue, it will try to obtain the execution token of the optimal batch (i.e., the target token) and query the number of available tokens in the current token bucket. If the target number of tokens is less than the number of available tokens, the optimal batch will be executed, the optimal batch will be added to the running queue, and the number of tokens in the token bucket will be reduced by the target number of tokens. If the target number of tokens is greater than the number of available tokens, it will continue to wait for the collection system to generate new tokens until the target number of tokens are available in the collection system and then execute the optimal batch.
[0090] Step g: If it is less than or equal to, executing the optimal batch according to the available tokens.
[0091] When it is determined that the target number of tokens is less than or equal to the number of available tokens, the optimal batch is executed based on the available tokens.
[0092] In this embodiment, by determining that the target number of tokens when the optimal batch is to be executed is less than or equal to the number of available tokens, the optimal batch is processed, thereby ensuring the normal operation of the optimal batch processing.
[0093] Furthermore, after the step of determining whether the target number of tokens is greater than the number of available tokens, the method further includes:
[0094] Step h: If it is greater than, obtain new tokens generated by the collection system based on a preset rate until the sum of the number of new tokens and the number of available tokens is equal to the target number of tokens, and execute the optimal batch according to the available tokens and the new tokens.
[0095] When it is determined that the target number of tokens is greater than the number of available tokens, it is necessary to obtain new tokens generated by the collection system based on the preset rate until the sum of the number of new tokens and the number of available tokens is equal to the target number of tokens, and the optimal batch will be processed based on the available tokens and the new tokens. For example, if the target number of tokens for the optimal batch is 3, and the number of available tokens in the current token bucket is 2, and the collection system generates tokens every 1 minute, it only needs to wait for one minute to obtain new tokens before processing the optimal batch. Figure 10 As shown in the flow diagram of the batch status, the batch goes from the waiting state to the ready state, then to the running state, and finally to the end. While in the waiting state, the batch needs to determine whether the collection system verification has passed. If the collection system verification fails, the batch enters the failed state. Furthermore, when the batch enters the running state from the ready state, it needs to obtain a token. If the token acquisition fails, the batch enters the suspended state and enters the ready state at the end of the suspended time. If a system abnormality is detected while the batch is running, the batch enters the failed state.
[0096] In this embodiment, when the target token quantity is greater than the available token quantity, new tokens are acquired and then the optimal batch is executed, thereby ensuring the normal operation of the optimal batch execution process.
[0097] The embodiment of the present invention also provides a data batch processing device, referring to Figure 3 , the data batch processing device includes:
[0098] An acquisition module is used to obtain batches corresponding to different types of resource data in the big data platform and determine whether each batch is in a ready state;
[0099] a determination module configured to, if there is a ready batch in a ready state in each of the batches, and if there are multiple ready batches, obtain the number of cases in each of the ready batches and the system resources of the collection system, and determine the optimal batch according to the system resources and the number of cases in each of the ready batches;
[0100] A processing module is used to obtain available tokens in the collection system, execute processing on the optimal batch based on the available tokens, and after the optimal batch processing is completed, continue to execute the steps of obtaining batches corresponding to different types of resource data in the big data platform until each batch processing is completed.
[0101] Optionally, the determining module is further configured to:
[0102] A first batch that is completed first and in a ready state is obtained based on each batch, and execution processing is performed on the first batch. System resources of the collection system are determined based on the execution processing result of the first batch.
[0103] Optionally, the determining module is further configured to:
[0104] Calculate the usage growth rate of the system resources and determine whether the usage growth rate is less than a preset alarm threshold,
[0105] If it is less than, the optimal batch is determined based on the number of cases in each of the ready batches.
[0106] Optionally, the determining module is further configured to:
[0107] The positive relationship between the increase in usage of the system resources and the increase in the number of cases in the first batch is calculated based on the execution processing result of the first batch, and the usage growth rate of the system resources is calculated according to the positive relationship.
[0108] Optionally, the determining module is further configured to:
[0109] Calculating an available resource value of the system resources according to the alarm threshold, and calculating a maximum target case load that can be admitted to the system resources according to the available resource value;
[0110] Determining a case level corresponding to each of the ready batches according to the number of cases in each of the ready batches, and determining a target level in each of the case levels based on the target case volume;
[0111] A target ready batch at a target level is obtained from each of the ready batches. If there are multiple target ready batches, a target ready batch with the largest number of cases is obtained from each of the target ready batches, and the target ready batch with the largest number of cases is used as the optimal batch.
[0112] Optionally, the processing module is further configured to:
[0113] Determining a target number of tokens when the optimal batch is to be executed, obtaining an available token number of available tokens in the collection system, and determining whether the target number of tokens is greater than the available number of tokens;
[0114] If it is less than or equal to, executing the optimal batch according to the available tokens.
[0115] The data batch processing device further includes:
[0116] If it is greater, new tokens generated by the collection system based on a preset rate are obtained until the sum of the number of new tokens and the number of available tokens is equal to the target number of tokens, and the optimal batch is executed based on the available tokens and the new tokens.
[0117] The methods executed by the above-mentioned program modules can refer to the various embodiments of the batch data processing method of the present invention, and will not be described in detail here.
[0118] The present invention also provides a computer storage medium.
[0119] The computer storage medium of the present invention stores a data batch processing program, which implements the steps of the data batch processing method described above when executed by a processor.
[0120] The method implemented when the batch processing program for data running on the processor is executed can refer to the various embodiments of the batch processing method for data of the present invention, and will not be described in detail here.
[0121] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0122] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0124] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for batch processing of data, characterized in that: The data batch processing method comprises the following steps: Obtaining batches corresponding to different types of resource data in the big data platform and determining whether each batch is in a ready state; If there is a ready batch in a ready state in each of the batches, and there are multiple ready batches, then the number of cases in each of the ready batches and the system resources of the collection system are obtained, and the optimal batch is determined based on the system resources and the number of cases in each of the ready batches; Obtaining available tokens in the collection system, executing processing on the optimal batch based on the available tokens, and after the optimal batch processing is completed, continuing to execute the step of obtaining batches corresponding to different types of resource data in the big data platform until each batch processing is completed; The step of determining the optimal batch size based on the system resources and the number of cases in each ready batch size includes: Calculate the usage growth rate of the system resources and determine whether the usage growth rate is less than a preset alarm threshold, If it is less than, calculating the available resource value of the system resources according to the alarm threshold, and calculating the maximum admissible target case volume of the system resources according to the available resource value; Determining a case level corresponding to each of the ready batches according to the number of cases in each of the ready batches, and determining a target level in each of the case levels based on the target case volume; Obtaining a target ready batch at the target level from each of the ready batches; if there are multiple target ready batches, obtaining a target ready batch with the largest number of cases from each of the target ready batches, and taking the target ready batch with the largest number of cases as the optimal batch; The step of obtaining system resources of the debt collection system includes: Obtaining a first batch that is completed first and in a ready state based on each batch, performing execution processing on the first batch, and determining system resources of the collection system based on the execution processing result of the first batch; The step of calculating the usage growth rate of the system resources includes: The positive relationship between the increase in usage of the system resources and the increase in the number of cases in the first batch is calculated based on the execution processing result of the first batch, and the usage growth rate of the system resources is calculated according to the positive relationship.
2. The method for batch processing of data according to claim 1, wherein: The step of obtaining available tokens from the collection system and executing the optimal batch based on the available tokens includes: Determining a target number of tokens when the optimal batch is to be executed, obtaining an available token number of available tokens in the collection system, and determining whether the target number of tokens is greater than the available number of tokens; If it is less than or equal to, the optimal batch is executed.
3. The method for batch processing of data according to claim 2, wherein: After the step of determining whether the target number of tokens is greater than the number of available tokens, the method further comprises: If it is greater, new tokens generated by the collection system based on a preset rate are obtained until the sum of the number of new tokens and the number of available tokens is equal to the target number of tokens, and the optimal batch is executed based on the available tokens and the new tokens.
4. A data batch processing device, characterized in that: The data batch processing device includes: An acquisition module is used to obtain batches corresponding to different types of resource data in the big data platform and determine whether each batch is in a ready state; a determination module configured to, if there is a ready batch in a ready state in each of the batches, and if there are multiple ready batches, obtain the number of cases in each of the ready batches and the system resources of the collection system, and determine the optimal batch according to the system resources and the number of cases in each of the ready batches; a processing module configured to obtain available tokens from the collection system, execute processing on the optimal batch based on the available tokens, and, after the optimal batch processing is completed, continue to execute the step of obtaining batches corresponding to different types of resource data from the big data platform until the processing of each batch is completed; The determination module is specifically configured to calculate the usage growth rate of the system resources and determine whether the usage growth rate is less than a preset alarm threshold; if so, calculate the available resource value of the system resources according to the alarm threshold, and calculate the maximum accessible target case volume of the system resources according to the available resource value; determine the case level corresponding to each ready batch according to the number of cases in each ready batch, and determine a target level in each case level based on the target case volume; obtain a target ready batch at the target level in each ready batch; if there are multiple target ready batches, obtain the target ready batch with the largest number of cases in each target ready batch, and use the target ready batch with the largest number of cases as the optimal batch; The determining module is specifically configured to obtain a first batch that is completed first and in a ready state based on each batch, execute the first batch, and determine the system resources of the collection system based on the execution result of the first batch; The determination module is specifically used to calculate the positive relationship between the increase in the use of the system resources and the increase in the number of cases in the first batch based on the execution processing results of the first batch, and calculate the growth rate of the use of the system resources according to the positive relationship.
5. A data batch processing device, characterized in that: The data batch processing device includes: a memory, a processor, and a data batch processing program stored in the memory and executable on the processor. When the data batch processing program is executed by the processor, the steps of the data batch processing method according to any one of claims 1 to 3 are implemented.
6. A computer storage medium, characterized in that The computer storage medium stores a data batch processing program, which, when executed by a processor, implements the steps of the data batch processing method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Method and device for categorical data batch processing
CN101017546A