Multi-server data processing method, device, computer equipment, storage medium

By dividing the transaction details file into multiple data shards and determining the processing shards according to the server load, multi-server collaborative processing is realized, solving the problems of low efficiency and poor fault tolerance under the traditional single-server architecture, and improving the efficiency and fault tolerance of data processing.

CN114490537BActive Publication Date: 2025-07-01INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210137996.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-15
Publication Date
2025-07-01
Estimated Expiration
2042-02-15

AI Technical Summary

Technical Problem

Traditional single-server architectures have problems of inefficiency and poor fault tolerance when processing transaction details files, especially in multiple server environments, which are difficult to achieve efficient data processing and fault tolerance.

Method used

Multi-server collaborative processing is realized by dividing the data file into multiple data shards and determining the data shards it processes based on the load information of each server. The specific steps include: the first server obtains the corresponding relationship between the server and the data shard, the second server divides the data files and determines the shard relationship, and the first server obtains the data files from the second server and processes the corresponding data shard.

Benefits of technology

Improves the efficiency and fault tolerance of data processing, can efficiently process large-scale transaction details files in a multi-server environment, and quickly identify and reprocess wrong data shards when errors occur.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114490537B_ABST
    Figure CN114490537B_ABST
Patent Text Reader

Abstract

The present application relates to a multi-server data processing method, apparatus, computer device, storage medium, and computer program product. The method includes: a first server obtains the correspondence between the server and the data shards, where the correspondence between the server and the data shards is determined based on the data file and the number of shards; the first server obtains the data file from a second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards; the first server processes the target data shard. By using this method, multiple first servers can be called. Each first server respectively obtains the data file from the second server, and then determines the target data shard responsible for each first server according to the correspondence between the server and the data shards. Each first server traverses the data file and only processes the data within the range of the corresponding target data shard, which can improve the efficiency and fault tolerance of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly to a multi-server data processing method, device, computer device, storage medium, and computer program product. Background Art

[0002] With the development of the economy and the large-scale application of mobile Internet technology, the frequency of customer transactions is increasing, and the transaction scale is growing. Due to the limited computing power of a single server, the server pressure increases, and the time taken by the bank to generate bills is also longer, and the time to complete bill generation is later. The single-node architecture can no longer meet the customer's requirements for the timeliness of receiving bills. The traditional method in the industry for generating bank bills is to receive transaction detail files on a single server. Each line in the file is the transaction detail of a customer, and the transaction details of the same customer are arranged in order together. The program reads the file line by line and generates bills. For some types of bills, a bill is generated for each line of transaction details, such as electronic receipts. For some types of bills, multiple lines of transaction details of the same customer are summarized to generate a bill, such as a merchant statement. Multiple lines of transaction details of the same merchant are displayed in one bill, and information such as the total transaction amount of this merchant is summarized and displayed.

[0003] The current method of processing transaction detail files has the following problems: Traditional batch jobs can only run on one server. If the server crashes, the job will be unexecutable. If the job needs to be rerun due to data problems, all data will be rerun. There is not enough fault tolerance space for processing transaction detail files, and the processing efficiency is low. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide a multi-server data processing method, device, computer device, computer-readable storage medium, and computer program product that can improve data processing efficiency.

[0005] In a first aspect, the present application provides a multi-server data processing method. The method includes:

[0006] The first server obtains the correspondence between the server and the data shard, and the correspondence between the server and the data shard is determined based on the data file and the number of shards;

[0007] The first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shard;

[0008] The first server processes the target data shard.

[0009] In one embodiment, the correspondence between the server and the data shard is determined based on the data file and the number of shards, including:

[0010] If the data type of the data file is a context-independent type, the second server divides the data file into the number of data shards, determines the correspondence between each first server and each data shard according to the load information of each first server, obtains the correspondence between the server and the data shard, or determines the correspondence between each first server and each data shard according to the load information of each first server, and determines the correspondence between each second server and each data shard according to the load information of each second server, to obtain the correspondence between the server and the data shard.

[0011] In one embodiment, dividing the data file into the number of data shards includes:

[0012] The second server obtains the data file and calculates the total number of data rows of the data file;

[0013] The second server calculates the number of shards according to the number of first servers and the attributes of each first server, and calculates the start row number and end row number of each data shard according to the number of shards and the total number of data rows;

[0014] The second server obtains multiple data shards of the data file according to the start row number and end row number of each data shard.

[0015] In one embodiment, the correspondence between the server and the data shard is determined based on the data file and the number of shards, including:

[0016] If the data type of the data file is a context-dependent type, the second server obtains multiple shard numbers according to the number of shards, determines the correspondence between each first server and each shard number according to the load information of each first server, obtains the correspondence between the server and the data shard, or determines the correspondence between each first server and each shard number according to the load information of each first server, and determines the correspondence between each second server and each shard number according to the load information of each second server, to obtain the correspondence between the server and the data shard.

[0017] In one embodiment, the first server obtains the data file from the second server and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shard, including:

[0018] The first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the shard number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of shards;

[0019] The first server obtains the shard number corresponding to the first server according to the correspondence between the server and the data shards, determines each piece of data to be processed corresponding to the first server according to the shard number corresponding to the first server and the shard number corresponding to each piece of data to be processed, and obtains the target data shard corresponding to the first server according to each piece of data to be processed corresponding to the first server.

[0020] In one embodiment, the first server processes the target data shard, including:

[0021] The first server processes each piece of data to be processed in the target data shard to generate corresponding first information;

[0022] According to the primary key value corresponding to each piece of data to be processed, determine the primary key value corresponding to each first information;

[0023] Integrate the first information with the same primary key value to generate second information.

[0024] In a second aspect, the present application also provides a multi-server data processing device. The device includes:

[0025] A relationship acquisition module, configured to enable the first server to obtain the correspondence between the server and the data shards, where the correspondence between the server and the data shards is determined based on the data file and the number of shards;

[0026] A data confirmation module, configured to enable the first server to obtain the data file from the second server and determine the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards;

[0027] A data processing module, configured to enable the first server to process the target data shard.

[0028] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0029] The first server obtains the correspondence between the server and the data shards, where the correspondence between the server and the data shards is determined based on the data file and the number of shards;

[0030] The first server obtains the data file from the second server and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards;

[0031] The first server processes the target data shard.

[0032] Fourthly, the present application also provides a computer-readable storage medium. On the computer-readable storage medium, there is a computer program stored, and when the computer program is executed by a processor, the following steps are implemented:

[0033] The first server obtains the correspondence between the server and the data shards, and the correspondence between the server and the data shards is determined based on the data file and the number of shards;

[0034] The first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards;

[0035] The first server processes the target data shard.

[0036] Fifthly, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0037] The first server obtains the correspondence between the server and the data shards, and the correspondence between the server and the data shards is determined based on the data file and the number of shards;

[0038] The first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards;

[0039] The first server processes the target data shard.

[0040] In the above multi-server data processing method, device, computer device, storage medium and computer program product, the first server obtains the correspondence between the server and the data shards, and the correspondence between the server and the data shards is determined based on the data file and the number of shards; the first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards; the first server processes the target data shard. By the second server pre-obtaining the data file to be processed, dividing the data file into a certain number of data shards, and pre-establishing the correspondence between the server and the data shards for the data file, when it is necessary to process the data file, multiple first servers are called, each first server respectively obtains the data file from the second server, and then determines the target data shard responsible for each first server according to the correspondence between the server and the data shards. Each first server traverses the data file and only processes the data within the range of the corresponding target data shard, which can improve the efficiency and fault tolerance of data processing. Description of the Drawings

[0041] Figure 1Schematic flowchart of a multi-server data processing method in an embodiment;

[0042] Figure 2 Schematic flowchart of selecting a primary server in an embodiment;

[0043] Figure 3 Schematic flowchart of dividing a data file into data shards in an embodiment;

[0044] Figure 4 Schematic flowchart of determining a target data shard corresponding to a first server in an embodiment;

[0045] Figure 5 Block diagram of a multi-server data processing apparatus in an embodiment;

[0046] Figure 6 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0047] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] In one embodiment, as Figure 1 shown, a multi-server data processing method is provided. In this embodiment, the method is described by taking the application of the method to a terminal as an example. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0049] Step 102, a first server obtains the correspondence between the server and the data shards, and the correspondence between the server and the data shards is determined based on the data file and the number of shards.

[0050] Among them, a data shard refers to dividing a large data set into multiple small data sets obtained. Each small data set is a data shard, and the data within each data shard range does not overlap with each other, and all data shards added together are equal to the complete large data set; the correspondence between the server and the data shards is used to represent which data shards each first server processes.

[0051] Specifically, for the data file to be processed, the number of shards is determined in advance according to the number of available first servers and the attribute status of each first server. Then, the data file is divided into data shards with the determined number of shards, and the corresponding relationship between the servers and the data shards is established. When processing the data file is required, each first server obtains the pre-established corresponding relationship between the servers and the data shards.

[0052] In a possible implementation, the second server determines the number of shards in advance according to the number of available first servers and the attribute status of each first server. Then, the data file is divided into data shards with the determined number of shards, and the corresponding relationship between the servers and the data shards is established. And the corresponding relationship between the servers and the data shards is stored in a remote database that can be accessed by all servers. When processing the data file is required, each first server accesses the remote database to obtain the pre-established corresponding relationship between the servers and the data shards.

[0053] Step 104: The first server obtains the data file from the second server and determines the target data shard corresponding to the first server from the data file according to the corresponding relationship between the servers and the data shards.

[0054] Among them, the second server can be a primary server selected from multiple first servers or a primary server selected from outside all first servers.

[0055] Specifically, the second server pre-stores the data file to be processed, determines the number of shards according to the number of available first servers and the attribute status of each first server, and then divides the data file into data shards with the determined number of shards. When processing the data file is required, each first server obtains the data file from the second server, and then determines the target data shards corresponding to each first server from the data file according to the corresponding relationship between the servers and the data shards. One first server can correspond to multiple data shards.

[0056] Step 106: The first server processes the target data shards.

[0057] Specifically, each first server traverses the entire data file, but only processes the data within the range of its own corresponding target data shards. Each server will generate at least one data processing result.

[0058] In the above multi-server data processing method, the first server obtains the correspondence between the server and the data shard, and the correspondence between the server and the data shard is determined based on the data file and the number of shards; the first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shard; the first server processes the target data shard. By having the second server pre-obtain the data file to be processed, divide the data file into a certain number of data shards, and establish in advance the correspondence between the server and the data shard, when it is necessary to process the data file, multiple first servers are called. Each first server respectively obtains the data file from the second server, and then determines the target data shard responsible for each first server according to the correspondence between the server and the data shard. Each first server traverses the data file and only processes the data within the range of the corresponding target data shard, which can improve the efficiency and fault tolerance of data processing.

[0059] In one embodiment, as Figure 2 shown, select one server from multiple first servers as the master server, and this master server is the second server. The method for selecting the master server includes:

[0060] Step 202, obtain the thread pool capacity pre-configured for each server.

[0061] Specifically, set a fixed thread pool capacity for each server in advance. The thread pool capacity refers to the maximum number of threads that the corresponding server can carry. The thread pool capacities of each server can be the same or different.

[0062] Step 204, check the remaining capacity of the thread pool queue of each server.

[0063] Specifically, for each server, determine the remaining capacity of the thread pool queue of the server according to the thread pool capacity of the server and the number of existing tasks in the thread pool queue of the corresponding server; that is, for each server, according to the maximum number of threads of the server and the number of threads in the thread pool queue of the corresponding server, obtain the number of threads that can still be added in the thread pool queue of the corresponding server, so as to determine the remaining capacity of the thread pool queue of the server. Then determine the idle degree of each server according to the remaining capacity of the thread pool queue of the server.

[0064] Step 206, select the most idle server as the master server, that is, select a server with the largest remaining capacity of the thread pool queue as the master server.

[0065] In this embodiment, by obtaining the thread pool capacity pre-configured for each server; checking the remaining capacity of the thread pool queue of each server; selecting the most idle server as the master server. It can ensure the data processing efficiency of the master server.

[0066] In one embodiment, the correspondence between the server and the data shards is determined based on the data file and the number of shards, including:

[0067] If the data type of the data file is context-independent, the second server divides the data file into the number of data shards, determines the correspondence between each first server and each data shard according to the load information of each first server, and obtains the correspondence between the server and the data shard, or determines the correspondence between each first server and each data shard according to the load information of each first server, and determines the correspondence between each second server and each data shard according to the load information of each second server, and obtains the correspondence between the server and the data shard.

[0068] Among them, the context-independent type means that there is no association between each piece of data in the data file, and each piece of data is an independent piece of information.

[0069] Specifically, the information included in the correspondence between the server and the data shard is different for different types of data files. If the data type of the data file is context-independent, the second server first calculates the number of shards according to the number of available first servers and the attributes of each first server, and then divides the data file into the number of data shards. Finally, obtain the load information of each first server, and use the server load balancing algorithm to assign at least one data shard to each first server, so as to determine the correspondence between each first server and each data shard, and obtain the correspondence between the server and the data shard.

[0070] In a possible implementation manner, if the data type of the data file is context-independent, the second server first calculates the number of shards according to the number of available first servers, the number of second servers, the attributes of each first server, and the attributes of each second server, and then divides the data file into the number of data shards. Finally, obtain the load information of each first server and the load information of each second server, and use the server load balancing algorithm to assign at least one data shard to each first server and each second server, so as to determine the correspondence between each first server and each data shard, and the correspondence between each second server and each data shard, and obtain the correspondence between the server and the data shard.

[0071] In one embodiment, if the data type of the data file is context-independent, the server load balancing algorithm adopted includes: setting a fixed thread pool for each server in advance, and a data shard processing task is a task in the thread pool queue. When allocating the number of shards that each server needs to process, it is necessary to check the remaining capacity of the server's thread pool queue. At the same time, unavailable servers will be detected. If a server crashes and becomes unavailable, the data shards will be allocated to the available servers for processing. The fewer the remaining capacity of the thread pool queue, the fewer the number of allocated shards. The more the remaining capacity of the thread pool queue, the more the number of allocated data shards. For example, there are m data shards in total, n servers in total, and the remaining capacity of the thread pool queue of the i-th server is n i , and the sum of the remaining tasks in the thread pool queues of all servers is s, then the number of shard tasks for each server is m*(n i / s).

[0072] In one embodiment, as Figure 3 shown, dividing the data file into the number of shards of data shards includes:

[0073] Step 302, the second server obtains the data file and calculates the total number of data rows in the data file.

[0074] Specifically, usually, each row in the data file records one piece of data.

[0075] Step 304, the second server calculates the number of shards according to the number of the first servers and the attributes of each first server, and calculates the start row number and the end row number of each data shard according to the number of shards and the total number of data rows.

[0076] Among them, the attribute of each first server refers to the number of cores of the CPU of each first server. Usually, the number of cores of a server's CPU determines the number of data shards that the server can process.

[0077] Specifically, the second server first calculates the number of shards according to the number of the first servers and the number of cores of the CPU of each first server. For example, if there are k first servers and the CPU of each first server has 4 cores, the number of shards is set to 4*k. Ideally, each server processes 4 shards concurrently, and each core of the CPU can process one shard.

[0078] Further, the second server calculates the start row number and the end row number of each data shard according to the number of shards and the total number of data rows. For example, if the total number of data rows is h, the number of shards is n, the start row number of the i-th shard (i starts counting from 0, 0≤i≤n-1) is (h / n)*i + 1. When i < n-1, the end row number of the i-th shard is (h / n)*(i + 1). When i = n-1, the end row number of the i-th shard is h.

[0079] Step 306: The second server obtains multiple data shards of the data file according to the starting row number and ending row number of each data shard.

[0080] Specifically, the second server divides the data file into multiple data shards according to the starting row number and ending row number of each data shard.

[0081] In this embodiment, the second server obtains the data file and calculates the total number of data rows of the data file; the second server calculates the number of shards according to the number of the first servers and the attributes of each first server, and calculates the starting row number and ending row number of each data shard according to the number of shards and the total number of data rows; the second server obtains multiple data shards of the data file according to the starting row number and ending row number of each data shard. The entire data file can be divided into multiple data shards, and multiple servers process each data shard respectively, improving the data processing efficiency. And when an error occurs in data processing, only the data shard where the error occurs needs to be determined, so as to determine the corresponding server and reprocess the data of this data shard, which also improves the fault tolerance rate of data processing.

[0082] In one embodiment, the correspondence between the server and the data shard is determined based on the data file and the number of shards, including:

[0083] If the data type of the data file is context-related, the second server obtains multiple shard numbers according to the number of shards, determines the correspondence between each first server and each shard number according to the load information of each first server, and obtains the correspondence between the server and the data shard, or determines the correspondence between each first server and each shard number according to the load information of each first server, and determines the correspondence between each second server and each shard number according to the load information of each second server, and obtains the correspondence between the server and the data shard.

[0084] Among them, the context-related type means that there is an association between some data in the data file. For example, there is business data of multiple enterprises in the data file, and there is an association between the business data belonging to the same enterprise. In the data file, each piece of data has a corresponding primary key, and the primary key values of the business data of the same enterprise are the same.

[0085] Specifically, for different types of data files, the information included in the corresponding relationship between the server and the data shards is different. If the data type of the data file is a context-related type, the second server first calculates the number of shards based on the number of available first servers and the attributes of each first server, and then configures multiple shard numbers for the data file according to the number of shards. Each shard number represents a data shard to be determined. Finally, the load information of each first server is obtained, and the server load balancing algorithm is used to allocate at least one shard number to each first server, so as to determine the corresponding relationship between each first server and each shard number, and obtain the corresponding relationship between the server and the data shard.

[0086] In a possible implementation, if the data type of the data file is a context-related type, the second server first calculates the number of shards based on the number of available first servers, the number of second servers, the attributes of each first server, and the attributes of each second server, and then configures multiple shard numbers for the data file according to the number of shards. Each shard number represents a data shard to be determined. Finally, the load information of each first server and the load information of each second server are obtained, and the server load balancing algorithm is used to allocate at least one shard number to each first server and each second server, so as to determine the corresponding relationship between each first server and each shard number, and the corresponding relationship between each second server and each shard number, and obtain the corresponding relationship between the server and the data shard.

[0087] In an embodiment, if the data type of the data file is a context-related type, the server load balancing algorithm used includes: setting a fixed thread pool for each server in advance, and a data shard processing task is a task in the thread pool queue. When allocating the shard numbers and the number of shard numbers that each server needs to process, it is necessary to check the remaining capacity of the server's thread pool queue. At the same time, unavailable servers will be detected. If a server crashes and becomes unavailable, the shard numbers will be allocated to the available servers for processing. The fewer the remaining capacity of the thread pool queue, the fewer the allocated shard numbers. The more the remaining capacity of the thread pool queue, the more the allocated shard numbers.

[0088] In an embodiment, as Figure 4 shown, the first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the corresponding relationship between the server and the data shard, including:

[0089] Step 402, the first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the shard number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of shards.

[0090] Specifically, after the first server obtains the data file, it traverses the entire file to identify the primary key values corresponding to each line of data in the data file. Usually, each line is a piece of data to be processed. The maximum value of the primary key value is not less than the number of shards. At least one group of data to be processed with the same primary key value is allocated for each shard number, so as to allocate the corresponding shard number for each piece of data to be processed. The number of shards each server is responsible for can be different.

[0091] In a possible implementation, the primary key value corresponding to each piece of data to be processed is modulo-divided by the number of shards to obtain a modulo result, and the shard number corresponding to this piece of data to be processed is determined according to the modulo result. For example, the number of shards is n, and the value range of the shard number i is from 0 to n - 1. Starting from the first line, the data file is traversed line by line. According to the primary key value m of each piece of data to be processed, it is determined whether this piece of data to be processed belongs to the i-th shard. When m % n = i (m divided by n to take the remainder), it means this record belongs to the i-th shard; assume the number of shards n = 5, and the values of the shard number i are 0, 1, 2, 3, 4. There is at least one piece of data to be processed with the primary key value m = 8, then the shard number corresponding to this part of the data to be processed is 8 % 5 = 3, and each piece of data to be processed with the primary key value of 8 belongs to the data shard with the shard number of 3.

[0092] Step 404, the first server obtains the shard number corresponding to the first server according to the corresponding relationship between the server and the data shard, determines each piece of data to be processed corresponding to the first server according to the shard number corresponding to the first server and the shard number corresponding to each piece of data to be processed, and obtains the target data shard corresponding to the first server according to each piece of data to be processed corresponding to the first server.

[0093] Specifically, each first server determines the shard number it is responsible for according to the corresponding relationship between the server and the data shard, and at the same time identifies the shard number corresponding to each piece of data to be processed in the data file. If the shard number corresponding to a piece of data to be processed is the same as the shard number corresponding to the current first server, then this piece of data to be processed is a piece of data to be processed corresponding to the current first server, and the target data shard corresponding to the current first server is obtained according to all the data to be processed corresponding to the current first server.

[0094] In this embodiment, the first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the shard number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of shards; the first server obtains the shard number corresponding to the first server according to the corresponding relationship between the server and the data shards, determines each piece of data to be processed corresponding to the first server according to the shard number corresponding to the first server and the shard number corresponding to each piece of data to be processed, and obtains the target data shard corresponding to the first server according to each piece of data to be processed corresponding to the first server. It can enable each server to only process the data within the range of the target data shard it is responsible for, improving the data processing efficiency. And when an error occurs in data processing, only the shard number where the error occurs needs to be determined, so as to determine the corresponding server and reprocess the data of this data shard, which also improves the fault tolerance rate of data processing.

[0095] In one embodiment, the first server processes the target data shard, including: the first server processes each piece of data to be processed in the target data shard to generate a corresponding first piece of information; according to the primary key value corresponding to each piece of data to be processed, determines the primary key value corresponding to each first piece of information; integrates the first pieces of information with the same primary key value to generate a second piece of information.

[0096] Specifically, if the data type of the data file is a context-independent type, each first server processes each piece of data to be processed in the corresponding target data shard, and generates a corresponding first piece of information for each piece of data to be processed.

[0097] Further, if the data type of the data file is a context-dependent type, each first server processes each piece of data to be processed in the corresponding target data shard, and generates a corresponding first piece of information for each piece of data to be processed; each first server identifies the primary key value corresponding to each piece of data to be processed in the target data shard, so as to determine the primary key value corresponding to each first piece of information, and integrates the first pieces of information with the same primary key value to generate a second piece of information, and the number of second pieces of information is the same as the number of types of primary keys (primary keys with the same key value are of one type).

[0098] In this embodiment, the first server processes the target data shard, including: the first server processes each piece of data to be processed in the target data shard to generate a corresponding first piece of information; according to the primary key value corresponding to each piece of data to be processed, determines the primary key value corresponding to each first piece of information; integrates the first pieces of information with the same primary key value to generate a second piece of information. It can generate the first piece of information according to each piece of data to be processed, and integrate all the related first pieces of information according to the related information of the data to be processed to obtain the second piece of information, improving the efficiency of subsequent data processing.

[0099] In a feasible embodiment, a multi-server data processing method includes:

[0100] If the data type of the data file is context-independent, the second server obtains the data file and calculates the total number of data rows in the data file; the second server calculates the number of shards according to the number of first servers and the attributes of each first server, and calculates the start row number and end row number of each data shard according to the number of shards and the total number of data rows; the second server obtains multiple data shards of the data file according to the start row number and end row number of each data shard.

[0101] The second server divides the data file into the number of shards, determines the correspondence between each first server and each data shard according to the load information of each first server, to obtain the correspondence between the server and the data shard, or determines the correspondence between each first server and each data shard according to the load information of each first server, and determines the correspondence between each second server and each data shard according to the load information of each second server, to obtain the correspondence between the server and the data shard.

[0102] The first server obtains the correspondence between the server and the data shard. The first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shard. The first server processes each piece of data to be processed in the target data shard to generate the corresponding first information.

[0103] In a feasible embodiment, a multi-server data processing method includes:

[0104] If the data type of the data file is context-dependent, the second server obtains multiple shard numbers according to the number of shards, determines the correspondence between each first server and each shard number according to the load information of each first server, to obtain the correspondence between the server and the data shard, or determines the correspondence between each first server and each shard number according to the load information of each first server, and determines the correspondence between each second server and each shard number according to the load information of each second server, to obtain the correspondence between the server and the data shard.

[0105] The first server obtains the correspondence between the server and the data shards. The first server obtains a data file from the second server, and the first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the shard number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of shards; the first server obtains the shard number corresponding to the first server according to the correspondence between the server and the data shards, determines each piece of data to be processed corresponding to the first server according to the shard number corresponding to the first server and the shard number corresponding to each piece of data to be processed, and obtains the target data shard corresponding to the first server according to each piece of data to be processed corresponding to the first server. The first server processes each piece of data to be processed in the target data shard to generate corresponding first information; according to the primary key value corresponding to each piece of data to be processed, determines the primary key value corresponding to each first information; integrates the first information with the same primary key value to generate second information.

[0106] In one embodiment, a multi-server data processing method, taking the generation of a bill according to a transaction detail file applied to a server as an example, the method includes:

[0107] The transaction detail file is in a shared directory. The files in this shared directory can be accessed simultaneously by programs on multiple servers. Elect a primary server from multiple servers. The primary server selects different sharding methods according to the type of the transaction detail file, divides the transaction detail file into multiple pieces, and then determines the correspondence between each server and the shards according to the server load balancing algorithm, and records this correspondence in a remote database. This remote database is accessible to all servers. The program on each server first accesses the remote database to obtain the correspondence between each server and the shards, then reads the transaction detail file, and only processes the data set belonging to its own shard to generate a bill. After the program generates the bill, through various channels, such as email, file transfer, each server sends the bill generated on its own server to the customer. Among them, the transaction detail file is equivalent to the data file, the primary server is equivalent to the second server, each server is equivalent to the first server, and the shard is equivalent to the data shard.

[0108] Each line in the transaction detail file is the detailed data of a transaction. The types of transaction detail files can be divided into two types: context-independent and context-dependent. The primary server selects different sharding methods for different types of transaction detail files.

[0109] The primary server is responsible for allocating the shards that each server needs to process. The method for electing the primary server is: check the remaining capacity of the thread pool queue of each server, and select the most idle server as the primary server, that is, select a server with the largest remaining capacity of the thread pool queue as the primary server.

[0110] When the master server assigns shards to each server, the required server load balancing algorithms include: setting a fixed thread pool for each server, and a shard processing task is a task in the thread pool queue. When allocating the number of shards that each server needs to process, it is necessary to check the remaining capacity of the server's thread pool queue. At the same time, unavailable servers will be detected. If a server crashes and becomes unavailable, then the shards will be assigned to available servers for processing. The fewer the remaining capacity of the thread pool queue, the fewer the number of shards assigned. The more the remaining capacity of the thread pool queue, the more the number of shards assigned. Suppose there are m shards in total, n servers in total, and the remaining capacity of the thread pool queue of the i-th server is n i , and the sum of the remaining tasks in the thread pool queues of all servers is s. Then the number of shard tasks for each server is m*(n i / s).

[0111] For transaction detail files that are not contextually related, a bill is generated for each transaction detail in each line of the transaction detail file. According to the number of cores of the server CPU, the number of shards is set in advance. For example, for k servers, if each server's CPU has 4 cores, the total number of shards is set to 4*k. Ideally, each server processes 4 shards concurrently, and each core of the CPU can process one shard. Traverse the transaction detail file line by line, calculate the number of lines in the file, and then divide by the number of shards to obtain the starting line number and ending line number of each shard. For example, if the number of lines in the file is h and the number of shards is n, the starting line number of the i-th shard (i starts counting from 0, 0 ≤ i ≤ n - 1) is (h / n)*i + 1. When i < n - 1, the ending line number of the i-th shard is (h / n)*(i + 1). When i = n - 1, the ending line number of the i-th shard is h. According to the current load balancing of each server, record each server and the starting and ending line numbers of the corresponding shards to be processed in the database table. Servers with a large load can process fewer shards, and servers with a small load can process more shards. When the program on each server runs, it first reads the data in the table to obtain the starting and ending line numbers of the shards to be processed. Traverse the transaction detail file from the first line to the last line, and only process the data within the shard range. The data outside the range is directly skipped without any processing. Each server sends the bill generated on its own server to the customer. Among them, the generated bill is equivalent to the first information.

[0112] For context-related transaction detail files, a bill is generated for multiple consecutive lines of transaction details. Taking the merchant statement as an example, the statement shows the records of multiple lines of transaction details for the same merchant, and also displays the statistical data of the transaction, such as the total transaction amount and the total number of transactions. In the merchant statement, each statement contains all the transaction details of a merchant, and each line of transaction details will contain the merchant number. The primary key m of the transaction details corresponds to the merchant number. Transaction details with the same merchant number should appear in one merchant statement. According to the number of cores of the server CPU, the number of shards is set in advance. For example, if there are k servers and each server has 4 cores, the total number of shards is set to 4*k. Ideally, each server processes 4 shards concurrently, and each CPU core can process one shard. According to the load balancing of the current server, each server and the corresponding shard sequence number to be processed are recorded in the database. Servers with heavy load can process fewer shards, and servers with light load can process more shards. When the program on each server runs, it first obtains the shard number i that it needs to process from the database. Assume that the total number of shards is n, and the value range of the shard number i is 0 to n-1. Starting from the first line, it traverses the transaction details file line by line, and determines whether this line of transaction record belongs to the i-th shard based on the statement primary key m. For example, a merchant statement shows all the transaction record information of a merchant number, so the value of the primary key m can be the merchant number. When m%n=i, it means that this record belongs to the i-th shard (for example, the total number of shards n is 5, and the value of the shard number i is 0,1,2,3,4. Then the shard corresponding to the merchant number 8 is 8%5=3, and the merchant belongs to the shard with sequence number 3). When the program on each server traverses the transaction details file from the first line, when determining whether this line of record belongs to the shard that it needs to process, it also needs to record the primary key information and aggregate the transaction records with the same primary key into one bill. Since the sharding is based on the value of the primary key m, the same primary key m (such as the merchant number) must be in the same shard, so that each server can obtain all the data of the same statement and send it after generating the statement. The generated statement is equivalent to the second information.

[0113] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0114] Based on the same inventive concept, an embodiment of the present application also provides a multi-server data processing device for implementing the multi-server data processing method described above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the multi-server data processing device provided below can refer to the limitations on the multi-server data processing method in the above text, and will not be repeated here.

[0115] In one embodiment, as Figure 5 shown, a multi-server data processing device 500 is provided, including: a relationship acquisition module 501, a data confirmation module 502, and a data processing module 503, where:

[0116] The relationship acquisition module 501 is used for the first server to obtain the correspondence between the server and the data shard, and the correspondence between the server and the data shard is determined based on the data file and the number of shards;

[0117] The data confirmation module 502 is used for the first server to obtain the data file from the second server, and determine the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shard;

[0118] The data processing module 503 is used for the first server to process the target data shard.

[0119] In one embodiment, the device further includes a relationship construction module 504:

[0120] The relationship construction module 504 is configured to, if the data type of the data file is context-independent, the second server divides the data file into a number of data slices equal to the number of slices, determines the correspondence between each first server and each data slice according to the load information of each first server, and obtains the correspondence between the server and the data slice, or determines the correspondence between each first server and each data slice according to the load information of each first server, and determines the correspondence between each second server and each data slice according to the load information of each second server, and obtains the correspondence between the server and the data slice.

[0121] In one embodiment, the relationship construction module 504 is further configured to the second server obtains the data file and calculates the total number of data rows of the data file; the second server calculates the number of slices according to the number of first servers and the attributes of each first server, and calculates the start row number and the end row number of each data slice according to the number of slices and the total number of data rows; the second server obtains a plurality of data slices of the data file according to the start row number and the end row number of each data slice.

[0122] In one embodiment, the apparatus further includes a relationship construction module 504:

[0123] The relationship construction module 504 is configured to, if the data type of the data file is context-dependent, the second server obtains a number of slice numbers according to the number of slices, determines the correspondence between each first server and each slice number according to the load information of each first server, and obtains the correspondence between the server and the data slice, or determines the correspondence between each first server and each slice number according to the load information of each first server, and determines the correspondence between each second server and each slice number according to the load information of each second server, and obtains the correspondence between the server and the data slice.

[0124] In one embodiment, the data confirmation module 502 is further configured to the first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the slice number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of slices; the first server obtains the slice number corresponding to the first server according to the correspondence between the server and the data slice, determines each piece of data to be processed corresponding to the first server according to the slice number corresponding to the first server and the slice number corresponding to each piece of data to be processed, and obtains the target data slice corresponding to the first server according to each piece of data to be processed corresponding to the first server.

[0125] In one embodiment, the data processing module 503 is further configured to process each piece of data to be processed in the target data shard by the first server to generate corresponding first information; determine the primary key value corresponding to each first information according to the primary key value corresponding to each piece of data to be processed; and integrate the first information with the same primary key value to generate second information.

[0126] Each module in the multi-server data processing device described above can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0127] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for the processor to exchange information with external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a multi-server data processing method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0128] Those skilled in the art can understand that Figure 6 the structure shown in

[0129] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0130] The first server obtains the correspondence between the server and the data shards. The correspondence between the server and the data shards is determined based on the data file and the number of shards.

[0131] The first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards.

[0132] The first server processes the target data shard.

[0133] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0134] If the data type of the data file is a context-independent type, the second server divides the data file into the number of data shards, determines the correspondence between each first server and each data shard according to the load information of each first server, and obtains the correspondence between the server and the data shards, or determines the correspondence between each first server and each data shard according to the load information of each first server, and determines the correspondence between each second server and each data shard according to the load information of each second server, and obtains the correspondence between the server and the data shards.

[0135] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0136] The second server obtains the data file and calculates the total number of data rows of the data file.

[0137] The second server calculates the number of shards according to the number of first servers and the attributes of each first server, and calculates the start row number and end row number of each data shard according to the number of shards and the total number of data rows.

[0138] The second server obtains multiple data shards of the data file according to the start row number and end row number of each data shard.

[0139] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0140] If the data type of the data file is a context - related type, the second server obtains multiple shard numbers according to the number of shards, determines the corresponding relationship between each first server and each shard number according to the load information of each first server, obtains the corresponding relationship between the server and the data shard, or determines the corresponding relationship between each first server and each shard number according to the load information of each first server, and determines the corresponding relationship between each second server and each shard number according to the load information of each second server, to obtain the corresponding relationship between the server and the data shard.

[0141] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0142] The first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the shard number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of shards;

[0143] The first server obtains the shard numbers corresponding to the first server according to the corresponding relationship between the server and the data shard, determines each piece of data to be processed corresponding to the first server according to the shard numbers corresponding to the first server and the shard numbers corresponding to each piece of data to be processed, and obtains the target data shard corresponding to the first server according to each piece of data to be processed corresponding to the first server.

[0144] In one embodiment, when the processor executes the computer program, the following steps are further implemented:

[0145] The first server processes each piece of data to be processed in the target data shard to generate corresponding first information;

[0146] According to the primary key value corresponding to each piece of data to be processed, determine the primary key value corresponding to each first information;

[0147] Integrate the first information with the same primary key value to generate second information.

[0148] In one embodiment, a computer - readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0149] The first server obtains the corresponding relationship between the server and the data shard, and the corresponding relationship between the server and the data shard is determined based on the data file and the number of shards;

[0150] The first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the corresponding relationship between the server and the data shard;

[0151] The first server processes the target data shard.

[0152] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0153] If the data type of the data file is a context-independent type, the second server divides the data file into a number of data shards equal to the number of shards, determines the correspondence between each first server and each data shard according to the load information of each first server, to obtain the correspondence between the servers and the data shards, or determines the correspondence between each first server and each data shard according to the load information of each first server, and determines the correspondence between each second server and each data shard according to the load information of each second server, to obtain the correspondence between the servers and the data shards.

[0154] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0155] The second server obtains the data file and calculates the total number of data rows of the data file;

[0156] The second server calculates the number of shards according to the number of first servers and the attributes of each first server, and calculates the starting row number and ending row number of each data shard according to the number of shards and the total number of data rows;

[0157] The second server obtains multiple data shards of the data file according to the starting row number and ending row number of each data shard.

[0158] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0159] If the data type of the data file is a context-dependent type, the second server obtains a number of shard numbers according to the number of shards, determines the correspondence between each first server and each shard number according to the load information of each first server, to obtain the correspondence between the servers and the data shards, or determines the correspondence between each first server and each shard number according to the load information of each first server, and determines the correspondence between each second server and each shard number according to the load information of each second server, to obtain the correspondence between the servers and the data shards.

[0160] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0161] The first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the shard number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of shards;

[0162] The first server obtains the shard number corresponding to the first server according to the correspondence between the server and the data shards, determines each piece of data to be processed corresponding to the first server according to the shard number corresponding to the first server and the shard number corresponding to each piece of data to be processed, and obtains the target data shard corresponding to the first server according to each piece of data to be processed corresponding to the first server.

[0163] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0164] The first server processes each piece of data to be processed in the target data shard to generate corresponding first information;

[0165] According to the primary key value corresponding to each piece of data to be processed, determine the primary key value corresponding to each first information;

[0166] Integrate the first information with the same primary key value to generate second information.

[0167] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0168] The first server obtains the correspondence between the server and the data shards, and the correspondence between the server and the data shards is determined based on the data file and the number of shards;

[0169] The first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards;

[0170] The first server processes the target data shard.

[0171] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0172] If the data type of the data file is a context-independent type, the second server divides the data file into the number of shards of data shards, determines the correspondence between each first server and each data shard according to the load information of each first server to obtain the correspondence between the server and the data shards, or determines the correspondence between each first server and each data shard according to the load information of each first server, and determines the correspondence between each second server and each data shard according to the load information of each second server to obtain the correspondence between the server and the data shards.

[0173] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0174] The second server obtains the data file and calculates the total number of data rows of the data file;

[0175] The second server calculates the number of shards based on the number of first servers and the attributes of each first server, and calculates the starting row number and ending row number of each data shard based on the number of shards and the total number of data rows;

[0176] The second server obtains multiple data shards of the data file according to the starting row number and ending row number of each data shard.

[0177] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0178] If the data type of the data file is a context-related type, the second server obtains multiple shard numbers according to the number of shards, determines the correspondence between each first server and each shard number according to the load information of each first server, obtains the correspondence between the server and the data shard, or determines the correspondence between each first server and each shard number according to the load information of each first server, and determines the correspondence between each second server and each shard number according to the load information of each second server, to obtain the correspondence between the server and the data shard.

[0179] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0180] The first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the shard number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of shards;

[0181] The first server obtains the shard numbers corresponding to the first server according to the correspondence between the server and the data shard, determines each piece of data to be processed corresponding to the first server according to the shard numbers corresponding to the first server and the shard numbers corresponding to each piece of data to be processed, and obtains the target data shard corresponding to the first server according to each piece of data to be processed corresponding to the first server.

[0182] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0183] The first server processes each piece of data to be processed in the target data shard to generate corresponding first information;

[0184] According to the primary key value corresponding to each piece of data to be processed, determine the primary key value corresponding to each first information;

[0185] Integrate the first information with the same primary key value to generate second information.

[0186] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0187] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0188] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0189] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A multi-server data processing method, characterized in that, The method includes: The first server obtains the correspondence between the servers and the data shards, where the correspondence between the servers and the data shards is determined by the second server based on the data file and the number of shards; the data types of the data files include: context-independent types and context-dependent types, where, for different types of the data files, the information included in the correspondence between the servers and the data shards is different; The first server obtains the data file from the second server, and determines the target data shard corresponding to the first server from the data file according to the correspondence between the servers and the data shards; the first server processes the target data shard; If the data type of the data file is a context-dependent type, the first server processing the target data shard includes: The first server processes each piece of data to be processed in the target data shard to generate corresponding first information; according to the primary key key value corresponding to each piece of data to be processed, determines the primary key key value corresponding to each piece of the first information; integrates the first information with the same primary key key value to generate second information.

2. The method according to claim 1, wherein The correspondence between the servers and the data shards is determined by the second server based on the data file and the number of shards, including: If the data type of the data file is a context-independent type, the second server divides the data file into the number of shards, determines the correspondence between each first server and each data shard according to the load information of each first server, to obtain the correspondence between the servers and the data shards, or determines the correspondence between each first server and each data shard according to the load information of each first server, and determines the correspondence between each second server and each data shard according to the load information of each second server, to obtain the correspondence between the servers and the data shards.

3. The method according to claim 2, wherein The dividing the data file into the number of shards includes: The second server obtains the data file and calculates the total number of data rows of the data file; The second server calculates the number of shards according to the number of first servers and the attributes of each first server, and calculates the starting row number and ending row number of each data shard according to the number of shards and the total number of data rows; The second server obtains multiple data shards of the data file according to the starting row number and ending row number of each data shard.

4. The method according to claim 1, wherein The correspondence between the servers and the data shards is determined by the second server based on the data file and the number of shards, including: If the data type of the data file is a context-dependent type, the second server obtains multiple shard numbers according to the number of shards, determines the correspondence between each first server and each shard number according to the load information of each first server, to obtain the correspondence between the servers and the data shards, or determines the correspondence between each first server and each shard number according to the load information of each first server, and determines the correspondence between each second server and each shard number according to the load information of each second server, to obtain the correspondence between the servers and the data shards.

5. The method according to claim 4, characterized in that, The first server obtains the data file from the second server, and determines a target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards, including: The first server obtains the primary key value corresponding to each piece of data to be processed in the data file, and determines the shard number corresponding to each piece of data to be processed according to the primary key value corresponding to each piece of data to be processed and the number of shards; The first server obtains the shard number corresponding to the first server according to the correspondence between the server and the data shards, determines each piece of data to be processed corresponding to the first server according to the shard number corresponding to the first server and the shard number corresponding to each piece of data to be processed, and obtains the target data shard corresponding to the first server according to each piece of data to be processed corresponding to the first server.

6. A multi-server data processing device, characterized in that, The device includes: A relationship obtaining module, configured to enable the first server to obtain the correspondence between the server and the data shards, where the correspondence between the server and the data shards is determined by the second server based on the data file and the number of shards; the data types of the data file include: context-independent types and context-dependent types, where, for different types of the data file, the information included in the correspondence between the server and the data shards is different; A data confirmation module, configured to enable the first server to obtain the data file from the second server, and determine a target data shard corresponding to the first server from the data file according to the correspondence between the server and the data shards; A data processing module, configured to enable the first server to process the target data shard; If the data type of the data file is a context-dependent type, the data processing module is further configured to enable the first server to process each piece of data to be processed in the target data shard to generate a corresponding first piece of information; determine the primary key value corresponding to each first piece of information according to the primary key value corresponding to each piece of data to be processed; and integrate the first pieces of information with the same primary key value to generate a second piece of information.

7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Distributed data processing method, device and system, and server

    CN107508901A