Dynamic data sharding processing method, system, medium and device

By adding data shard processing nodes in business batch processing and dynamically adjusting shards, the problems of difficult changes in shard settings and waste resources in the existing technology are solved, flexible data sharding and merging are realized, data aggregation needs are supported, and batch running efficiency is improved.

CN115033562BActive Publication Date: 2025-05-06IND BANK CO +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210503900.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-10
Publication Date
2025-05-06
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

The existing data sharding algorithm is not easy to change after sharding is set, and can only increase exponentially when adding shards, resulting in waste of resources and complex deployment models, making it difficult to support data aggregation needs.

Method used

Add data shard processing nodes in business batch processing. By setting the number of subtasks and registering the corresponding number of subtasks, selecting partition keys and obtaining the paging keywords of the partition keys according to preset rules, dynamically increasing or decreasing shards to achieve flexible sharding and merging of data.

Benefits of technology

The ability to dynamically adjust shards is realized without stopping services or performing data migration, reducing resource waste and deployment complexity, supporting flexible data aggregation needs, and improving batch running efficiency of each shard.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115033562B_ABST
    Figure CN115033562B_ABST
Patent Text Reader

Abstract

The present invention provides a dynamic data sharding processing method, system, medium and equipment, including: step 1: adding a data sharding processing node in a business batch; step 2: setting the number of subtasks and registering a corresponding number of subtasks; step 3: selecting a partition key, obtaining the paging keyword of the partition key according to a preset rule, and recording the paging keyword in a subtask parameter; step 4: pulling up subtasks according to a preset concurrency number on any number of task processors; step 5: applying the business processing logic on the original trunk to the subtask for processing; step 6: checking the processing status of all subtasks, and continuing to execute in the original task order after all subtasks are processed. The present invention does not change the original architecture, does not require supporting services, has a short implementation cycle, can arbitrarily change the sharding key, evenly divides the sharding data, and ensures the efficiency of each sharding batch.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data sharding processing, and in particular to a method, system, medium and device for dynamic data sharding processing. Background Art

[0002] Conventional data sharding algorithms require planning sharding keys in advance, dividing shards according to sharding keys, establishing sharding databases, and developing sharding routing front-end services. While improving efficiency, this also causes many problems:

[0003] 1) Sharding is not easy to change after it is set up. It is necessary to stop the service to migrate the data, which is prone to errors and takes a long time when the amount of data is large;

[0004] 2) Shards cannot be reduced, only increased;

[0005] 3) Sharding can only be multiplied, which is easy to waste resources;

[0006] 4) The sharding key cannot be changed after it is selected. If the selection is inappropriate, the sharding will be invalid;

[0007] 5) The deployment model is complex and requires a lot of supporting services;

[0008] 6) If there is a need for data aggregation, it is difficult to support.

[0009] Patent document CN109085995A (application number: CN201710445789.5) discloses a storage method, device and system for dynamic data sharding. The method includes: responding to the access request initiated by the client, performing business analysis, obtaining the node parameters and storage parameters of the configuration management template; collecting the network bandwidth of each storage node in real time; determining the current number of file shards according to the node parameters, storage parameters and the network bandwidth of each storage node; sharding or merging the data according to the current number of file shards, so that the routing controller can access the data slices of each corresponding storage node. However, this patent still cannot solve the above problems. Summary of the invention

[0010] In view of the defects in the prior art, the object of the present invention is to provide a dynamic data sharding processing method, system, medium and device.

[0011] The dynamic data sharding processing method provided by the present invention includes:

[0012] Step 1: Add data sharding processing nodes in business batch processing;

[0013] Step 2: Set the number of subtasks and register the corresponding number of subtasks;

[0014] Step 3: Select a partition key, obtain the paging keyword of the partition key according to the preset rules, and record the paging keyword in the subtask parameters;

[0015] Step 4: Start subtasks on any number of task processors according to the preset concurrency number;

[0016] Step 5: Apply the original business processing logic on the main trunk to the subtask for processing;

[0017] Step 6: Check the processing status of all subtasks, and continue to execute in the original task order after all subtasks are processed.

[0018] Preferably, step 3 comprises:

[0019] Step 3.1: Filter and clean the business data to be processed;

[0020] Step 3.2: Sort by partition key and deduplicate by partition key;

[0021] Step 3.3: Get the minimum and maximum virtual nodes, and record a partition key for every preset number of partition key records.

[0022] Preferably, if the partition keyword is not obtained, the minimum and maximum virtual nodes are used as the only partition keywords. Otherwise, the minimum partition keyword is used as the starting node of the first partition, the maximum partition keyword is used as the ending node of the last partition, and the middle partition keyword is used as the ending node of the previous partition and the starting node of the next partition. All partitions are left-open and right-closed intervals.

[0023] Preferably, step 5 comprises:

[0024] Step 5.1: Obtain the subtask to be processed through the business processing machine;

[0025] Step 5.2: After obtaining ownership of a subtask, load the partition data according to the partition key;

[0026] Step 5.3: Perform business processing on the partitioned data. During the processing, divide the task into multiple steps, and divide each step into multiple segments for processing. If the processing conditions are not met temporarily, interrupt the task and restart it at any time when the conditions are met.

[0027] The dynamic data sharding processing system provided by the present invention comprises:

[0028] Module M1: Add data sharding processing nodes in business batch processing;

[0029] Module M2: Set the number of subtasks and register the corresponding number of subtasks;

[0030] Module M3: Select a partition key, obtain the paging keyword of the partition key according to the preset rules, and record the paging keyword into the subtask parameters;

[0031] Module M4: Start subtasks on any number of task processors according to the preset concurrency number;

[0032] Module M5: Apply the original business processing logic on the main trunk to the subtask for processing;

[0033] Module M6: Check the processing status of all subtasks, and continue to execute in the original task order after all subtasks are processed.

[0034] Preferably, the module M3 includes:

[0035] Module M3.1: Filter and clean the business data to be processed;

[0036] Module M3.2: Sort by partition key and deduplicate by partition key;

[0037] Module M3.3: Get the minimum and maximum virtual nodes, and record a partition key for every preset number of partition key records.

[0038] Preferably, if the partition keyword is not obtained, the minimum and maximum virtual nodes are used as the only partition keywords. Otherwise, the minimum partition keyword is used as the starting node of the first partition, the maximum partition keyword is used as the ending node of the last partition, and the middle partition keyword is used as the ending node of the previous partition and the starting node of the next partition. All partitions are left-open and right-closed intervals.

[0039] Preferably, the module M5 comprises:

[0040] Module M5.1: Obtain the subtasks to be processed through the business processor;

[0041] Module M5.2: After obtaining ownership of a subtask, load partition data according to the partition keyword;

[0042] Module M5.3: Perform business processing on partitioned data. During the processing, the task is divided into multiple segments for processing. If the processing conditions are not met temporarily, the task is interrupted and can be restarted for processing at any time after the conditions are met.

[0043] According to the computer-readable storage medium storing a computer program provided by the present invention, the steps of the method described above are implemented when the computer program is executed by a processor.

[0044] The dynamic data slicing processing device provided by the present invention includes: a controller;

[0045] The controller includes a computer-readable storage medium storing a computer program, and the computer program implements the steps of the dynamic data slicing processing method when executed by a processor; or, the controller includes the dynamic data slicing processing system.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] The present invention proposes a technology that is easy to implement and can dynamically increase or decrease shards according to the amount of data. It does not change the original architecture, does not require supporting services, has a short implementation cycle, can arbitrarily change the shard key, evenly divides the shard data, and ensures the efficiency of each shard batch. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0049] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0050] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0051] Example:

[0052] The present invention is suitable for processing large quantities of data. First, a suitable partition key is selected, and the paging keywords of the partition key are obtained according to certain rules and recorded in the segmentation subtask. The segmentation subtask obtains the processed partition data according to the paging keywords and performs related business processing.

[0053] The present invention isolates data sharding from data processing. Data sharding does not depend on the number of data processing machines. Increasing or reducing data processing machines does not involve data migration, nor does it need to use special processing algorithms such as consistent hashing algorithms. Data processing tasks do not need to pay attention to which shard they are processing.

[0054] like Figure 1 The present invention provides an easy-to-implement dynamic data sharding processing method, comprising the following steps:

[0055] Step 1: Add data sharding processing nodes in business batch processing;

[0056] Step 2: Select a partition key, obtain the paging keyword of the partition key according to certain rules, and record the paging keyword in the subtask parameters;

[0057] Step 3: Register the corresponding number of subtasks according to the number of subtasks;

[0058] Step 4: Start subtasks on any number of task processors according to a certain number of concurrency;

[0059] Step 5: Put the business processing logic on the original trunk into the subtask for processing;

[0060] Step 6: The parent task checks the processing status of all subtasks and continues with the original steps after all subtasks are processed.

[0061] Wherein step 2 comprises the following steps:

[0062] Step 2.1: Filter and clean the business data to be processed;

[0063] Step 2.2: Sort by partition key and deduplicate by partition key;

[0064] Step 2.3: Get the minimum and maximum virtual nodes, and record a partition key every certain number of partition key records;

[0065] Step 2.4: If the partition key is not obtained, the minimum and maximum virtual nodes are used as the only partition key. Otherwise, the minimum partition key is used as the starting node of the first partition, the maximum partition key is used as the ending node of the last partition, the middle partition key is used as the ending node of the previous partition and the starting node of the next partition. All partitions are left-open and right-closed intervals.

[0066] Step 5 includes the following steps:

[0067] Step 5.1: All business processors compete for the subtasks to be processed;

[0068] Step 5.2: After obtaining ownership of a subtask, load the partition data according to the partition key;

[0069] Step 5.3: Perform business processing on partitioned data. The processing can be divided into multiple steps, and each step can be divided into multiple segments. The framework of each subtask is a copy of the framework of the parent task, that is, logically they are parent and child tasks, but they are equivalent in processing, which reduces the complexity of scheduling.

[0070] Step 5.4: If the processing conditions are not met temporarily, the task can be interrupted and restarted at any time after the conditions are met.

[0071] The dynamic data sharding processing system provided by the present invention includes: module M1: adding data sharding processing nodes in business batch processing; module M2: setting the number of subtasks and registering a corresponding number of subtasks; module M3: selecting a partition key, obtaining a paging keyword of the partition key according to a preset rule, and recording the paging keyword in a subtask parameter; module M4: pulling up subtasks on any number of task processors according to a preset concurrency number; module M5: applying the business processing logic on the original trunk to the subtask for processing; module M6: checking the processing status of all subtasks, and continuing to execute in the original task order after all subtasks have been processed.

[0072] The module M3 includes: module M3.1: filtering and cleaning the business data to be processed; module M3.2: sorting by partition key and deduplication according to the partition key; module M3.3: obtaining the minimum and maximum virtual nodes, and recording a partition keyword according to the number of partition key records for every preset number. If the partition keyword is not obtained, the minimum and maximum virtual nodes are used as the only partition keywords, otherwise the minimum partition keyword is used as the starting node of the first partition, the maximum partition keyword is used as the ending node of the last partition, and the middle partition keyword is used as the ending node of the previous partition and the starting node of the next partition, and all partitions are left-open and right-closed intervals.

[0073] The module M5 includes: module M5.1: obtaining the subtask to be processed through the business processor; module M5.2: after obtaining the ownership of a subtask, loading the partition data according to the partition keyword; module M5.3: performing business processing on the partition data, dividing the task into multiple segments for processing during the processing, and interrupting the task if the processing conditions are not met temporarily, and restarting the task for processing at any time after the conditions are met.

[0074] According to the computer-readable storage medium storing a computer program provided by the present invention, the steps of the method described above are implemented when the computer program is executed by a processor.

[0075] The dynamic data sharding processing device provided according to the present invention comprises: a controller; the controller comprises a computer-readable storage medium storing a computer program, and the computer program implements the steps of the dynamic data sharding processing method when executed by a processor; or the controller comprises the dynamic data sharding processing system.

[0076] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.

[0077] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A dynamic data sharding processing method, characterized in that: include: Step 1: Add data sharding processing nodes in business batch processing; Step 2: Set the number of subtasks and register the corresponding number of subtasks; Step 3: Select a partition key, obtain the paging keyword of the partition key according to the preset rules, and record the paging keyword in the subtask parameters; Step 4: Start subtasks on any number of task processors according to the preset concurrency number; Step 5: Apply the original business processing logic on the main trunk to the subtask for processing; Step 6: Check the processing status of all subtasks, and continue to execute in the original task order after all subtasks are processed; The step 3 comprises: Step 3.1: Filter and clean the business data to be processed; Step 3.2: Sort by partition key and deduplicate by partition key; Step 3.3: Get the minimum and maximum virtual nodes, and record a partition key for every preset number of partition key records; If the partition key is not obtained, the minimum and maximum virtual nodes are used as the only partition key. Otherwise, the minimum partition key is used as the starting node of the first partition, the maximum partition key is used as the ending node of the last partition, and the middle partition key is used as the ending node of the previous partition and the starting node of the next partition. All partitions are left-open and right-closed intervals. The step 5 comprises: Step 5.1: Obtain the subtask to be processed through the business processing machine; Step 5.2: After obtaining ownership of a subtask, load the partition data according to the partition key; Step 5.3: Perform business processing on the partitioned data. During the processing, divide the task into multiple steps, and divide each step into multiple segments for processing. If the processing conditions are not met temporarily, interrupt the task and restart it at any time when the conditions are met.

2. A dynamic data sharding processing system, characterized in that: include: Module M1: Add data sharding processing nodes in business batch processing; Module M2: Set the number of subtasks and register the corresponding number of subtasks; Module M3: Select a partition key, obtain the paging keyword of the partition key according to the preset rules, and record the paging keyword into the subtask parameters; Module M4: Start subtasks on any number of task processors according to the preset concurrency number; Module M5: Apply the original business processing logic on the main trunk to the subtask for processing; Module M6: Check the processing status of all subtasks, and continue to execute in the original task order after all subtasks are processed; The module M3 comprises: Module M3.1: Filter and clean the business data to be processed; Module M3.2: Sort by partition key and deduplicate by partition key; Module M3.3: Get the minimum and maximum virtual nodes, and record a partition key for every preset number of partition key records; If the partition key is not obtained, the minimum and maximum virtual nodes are used as the only partition key. Otherwise, the minimum partition key is used as the starting node of the first partition, the maximum partition key is used as the ending node of the last partition, and the middle partition key is used as the ending node of the previous partition and the starting node of the next partition. All partitions are left-open and right-closed intervals. The module M5 comprises: Module M5.1: Obtain the subtasks to be processed through the business processor; Module M5.2: After obtaining ownership of a subtask, load partition data according to the partition keyword; Module M5.3: Perform business processing on partitioned data. During the processing, the task is divided into multiple segments for processing. If the processing conditions are not met temporarily, the task is interrupted and can be restarted for processing at any time after the conditions are met.

3. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to claim 1 are implemented.

4. A dynamic data sharding processing device, characterized in that: include: Controller; The controller includes the computer-readable storage medium storing the computer program as described in claim 3, and when the computer program is executed by the processor, the steps of the dynamic data sharding processing method as described in claim 1 are implemented; or, the controller includes the dynamic data sharding processing system as described in claim 2.

Citation Information

Patent Citations

  • Storage method, device and system for data dynamic fragmentation

    CN109085995A

  • Data dynamic fragmentation storage methods, devices and systems

    CN109085995B

  • Data storage method and device, equipment and storage medium

    CN112231398A