A data processing method and apparatus
By determining partition rules based on the calculation parallelism and partition fields in fund data processing, and distribute the reading and calculating incremental data, the problem that centralized data calculation cannot meet the high timeliness requirements is solved, and efficient data processing is achieved.
Patent Information
- Application Number
- CN202211370917.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-11-03
AI Technical Summary
Existing node centralized data calculations cannot meet the high-time requirements for fund-related business data processing.
By determining partition rules based on the set calculation parallelism and partition fields in the incremental data table, the incremental data is read distributedly and the latest historical data associated with the stock data is calculated.
It realizes distributed reading, writing and computing of data without changing the database storage mode, improves data processing efficiency and meets high timeliness requirements.
Smart Images

Figure CN115544075B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a data processing method and apparatus. Background Art
[0002] With the continuous expansion of the scale of the fund industry, the amount of data related to fund business has begun to show explosive growth. Currently, the calculation method for data related to fund business is single-node centralized data calculation (centralized calculation means obtaining data and performing data calculation at one location in the network, usually only one thread executes the calculation task). However, using single-node centralized data calculation can no longer meet the high timeliness requirements of data processing. Summary of the Invention
[0003] In view of this, the present invention provides a data processing method and apparatus to solve the problem that the existing node centralized data calculation cannot meet the high timeliness requirements of data processing. The technical solutions are as follows:
[0004] A data processing method, applied to a processing device, the method includes:
[0005] Determine a partition rule corresponding to the incremental data table according to a set calculation parallelism and a partition field selected from the incremental data tables participating in the calculation, where the partition rule is used to indicate the number of logical partitions set for the incremental data table, and the incremental data divided into each logical partition;
[0006] Distributively read the incremental data in the incremental data table into the corresponding logical partitions set according to the partition rule corresponding to the incremental data table;
[0007] Distributively read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database;
[0008] Calculate the latest historical data associated with the incremental data in each logical partition and the incremental data in each logical partition.
[0009] Optionally, the determining a partition rule corresponding to the incremental data table according to a set calculation parallelism and a partition field selected from the incremental data tables participating in the calculation includes:
[0010] Determine the number of logical partitions according to the calculation parallelism, and determine the end truncation length of the partition field;
[0011] Determine the end truncation range of the partition field according to the end truncation length of the partition field;
[0012] Determine the incremental data divided into each logical partition according to the end truncation range of the partition field, so as to obtain the partition rule corresponding to the incremental data table.
[0013] Optionally, the distributed reading of the incremental data in the incremental data table into the corresponding logical partitions set according to the partition rule corresponding to the incremental data table includes:
[0014] Generate a data query statement according to the partition rule corresponding to the incremental data table;
[0015] Use the data query statement to distributedly read the incremental quantity from the incremental data table into the corresponding logical partitions set.
[0016] Optionally, the generating of the data query statement according to the partition rule corresponding to the incremental data table includes:
[0017] Determine the incremental data to be read for each logical partition according to the partition rule corresponding to the incremental data table;
[0018] Generate a data query statement corresponding to each logical partition according to the incremental data to be read for each logical partition.
[0019] Optionally, the distributed reading of the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database includes:
[0020] Summarize multiple pieces of incremental data with the same partition field information in each logical partition into one piece, and only retain one piece of partition field information, where the partition field information is the specific field value corresponding to the partition field;
[0021] Read all the stock data of the users to which the incremental data in each logical partition belongs from the stock data of each user in the database;
[0022] Based on the set filtering dimension, filter out the latest historical data associated with the incremental data in each logical partition from all the stock data of the users to which the incremental data in each logical partition belongs.
[0023] Optionally, the reading of all the stock data of the users to which the incremental data in each logical partition belongs from the stock data of each user in the database includes:
[0024] Determine the number of interactions with the database based on the computing parallelism, the total number of partition field information in each logical partition, and the maximum query number of the data query statement;
[0025] Based on the number of interactions, batch-read all the stock data of the users to whom the incremental data in each logical partition belongs from the stock data of each user in the database.
[0026] A data processing device includes: a partition rule determination module, an incremental data reading module, a stock data reading module, and a data calculation module;
[0027] The partition rule determination module is configured to determine the partition rule corresponding to the incremental data table according to the set calculation parallelism and the partition field selected from the incremental data tables participating in the calculation, where the partition rule is used to indicate the number of logical partitions set for the incremental data table and the incremental data divided into each logical partition;
[0028] The incremental data reading module is configured to distributively read the incremental data in the incremental data table into the corresponding logical partitions set according to the partition rule corresponding to the incremental data table;
[0029] The stock data reading module is configured to distributively read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database;
[0030] The data calculation module is configured to calculate the latest historical data associated with the incremental data in each logical partition and the incremental data in each logical partition.
[0031] Optionally, the partition rule determination module includes: a partition number determination sub-module, an end truncation length determination sub-module, an end truncation range determination sub-module, and a partition rule determination sub-module;
[0032] The partition number determination sub-module is configured to determine the number of logical partitions according to the calculation parallelism;
[0033] The end truncation length determination sub-module is configured to determine the end truncation length of the partition field according to the calculation parallelism;
[0034] The end truncation range determination sub-module is configured to determine the end truncation range of the partition field according to the end truncation length of the partition field;
[0035] The partition rule determination sub-module is configured to determine the incremental data divided into each logical partition according to the end truncation range of the partition field, so as to obtain the partition rule corresponding to the incremental data table.
[0036] Optionally, the incremental data reading module includes: a query statement generation sub-module and a data reading sub-module;
[0037] The query statement generation sub-module is used to generate a data query statement according to the partitioning rule corresponding to the incremental data table;
[0038] The data reading sub-module is used to distributively read the incremental quantity from the incremental data table into the corresponding logical partitions set by using the data query statement.
[0039] Optionally, the stock data reading module includes: a data preprocessing sub-module, a stock data reading sub-module, and a data screening sub-module;
[0040] The data preprocessing sub-module is used to summarize multiple incremental data with the same partition field information in each logical partition into one, and only retain one partition field information, where the partition field information is the specific field information corresponding to the partition field;
[0041] The stock data reading sub-module is used to read all the stock data of the users to which the incremental data in each logical partition belongs from the stock data of each user in the database;
[0042] The data screening sub-module is used to screen out the latest historical data associated with the incremental data in each logical partition from all the stock data of the users to which the incremental data in each logical partition belongs based on the set screening dimension.
[0043] The data processing method and device provided by the present invention first determine the partitioning rule corresponding to the incremental data table according to the set calculation parallelism and the partitioning field selected from the incremental data tables participating in the calculation, then distributively read the incremental data in the incremental data table into the corresponding logical partitions set according to the partitioning rule corresponding to the incremental data table, then distributively read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database, and finally calculate the latest historical data associated with the incremental data in each logical partition and the incremental data in each logical partition. The data processing method and device provided by the present invention can realize the distributed reading, writing and calculation of data without changing the database storage mode, have high data processing efficiency for the distributed reading, writing and calculation of data, and can meet the high timeliness requirements of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to the provided drawings without creative efforts.
[0045] Figure 1Schematic flowchart of the data processing method provided by an embodiment of the present invention;
[0046] Figure 2 Schematic flowchart of determining a partition rule corresponding to an incremental data table according to a set computing parallelism and a partition field selected from the incremental data tables participating in the calculation provided by an embodiment of the present invention;
[0047] Figure 3 Schematic flowchart of distributively reading incremental data in an incremental data table into a corresponding logical partition set according to the partition rule corresponding to the incremental data table provided by an embodiment of the present invention;
[0048] Figure 4 Schematic diagram of distributively reading an incremental quantity from the incremental data tables participating in the calculation by using a data query statement provided by an embodiment of the present invention;
[0049] Figure 5 Schematic flowchart of distributively reading the latest historical data associated with the incremental data in each logical partition from the stock data of each user in a database provided by an embodiment of the present invention;
[0050] Figure 6 Schematic diagram of the structure of the data processing device provided by an embodiment of the present invention. Detailed implementation manners
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] In view that the single-node centralized data calculation cannot meet the high timeliness requirements of data processing, the inventors of this case conducted research. The initial idea was to optimize the table design in the original relational database or perform code refactoring on the original calculation program. For example, optimize the original data table design structure, analyze the execution plan, increase the index hit rate, etc., or perform code refactoring on the original calculation program to reduce the cost of redundant data calculation, etc. The inventors of this case studied the above ideas and found that the above ideas are limited by the technical architecture and essentially cannot solve the problem.
[0053] In view of this, the inventor of the present invention conducted further research and came up with the idea of migrating the data in the centralized storage database to the distributed database, and then selecting an adapted distributed computing engine for data interaction and calculation. Among them, for the centralized database, it contains a database file located in one location in the network, and multiple users can access this single database. Among them, for the distributed database, it contains multiple database files located in different locations in the network. In other words, the database is split into multiple files, and users can access the most recent database file, which will increase the speed of retrieving data. In addition, different users can access and operate data related to them, which can avoid user interference.
[0054] However, the above solutions bring problems: large amount of data to be migrated, high technical cost, and uncontrollable risks. The industry generally does not migrate large-scale historical data at one time, but often gradually and steadily advances technology or conducts distributed practices on new projects. It can be seen that the above solutions are time-consuming and labor-intensive.
[0055] In view of the many problems with the above-mentioned solutions, the inventor of this case continued to conduct research and finally proposed a data processing method with better effect through continuous research. This method can meet the high timeliness requirements of data processing and does not require the migration of data in the original database.
[0056] Next, the data processing method provided by the present invention is introduced through the following embodiments.
[0057] See also Figure 1 , shows a flow chart of a data processing method provided by an embodiment of the present invention, the data processing method is applied to a processing device (the processing device may be a server, or a server cluster composed of multiple servers, or a cloud computing server center), the data processing method may include:
[0058] Step S101: Determine the partitioning rule corresponding to the incremental data table according to the set calculation parallelism and the partitioning field selected from the incremental data table participating in the calculation.
[0059] Among them, the incremental data table involved in the calculation includes several incremental data, and the incremental data is the data generated by a certain user's transaction behavior on a certain day, such as transaction data, dividend data, etc.
[0060] The partitioning rule is used to indicate the number of logical partitions set for the incremental data table involved in the calculation, and the incremental data divided into each logical partition.
[0061] In this embodiment, the number of logical partitions is the same as the number of computational parallelism. For example, if the computational parallelism is 5, then 5 logical partitions need to be set for the incremental data table involved in the computation, and then the incremental data divided into each logical partition is set. It should be noted that the division of incremental data into set logical partitions mentioned in this embodiment refers to logical division.
[0062] It should be noted that the selected partition field must be random and unique. If the selected partition field is not random, some logical partitions will have a large amount of data, while some logical partitions will have a small amount of data, which will cause data skew. That is, the randomness of the partition field can greatly avoid data skew in logical partitions.
[0063] For example, if the partition field is a fund account number and the calculation parallelism is 5, then 5 logical partitions can be set for the incremental data table involved in the calculation, such as Figure 2 As shown, the last digit of the fund account number can be taken and distributed approximately evenly into 5 groups, for example, Block(1,6), Block(2,7), Block(3,8), Block(4,9), Block(5,0), where Block(1,6) means that the fund accounts with the last digits of 1 and 6 are divided into this partition, and Block(2,7) means that the fund accounts with the last digits of 2 and 7 are divided into this partition, and the same is true for the others.
[0064] Step S102: According to the partitioning rule corresponding to the incremental data table, the incremental data in the incremental data table is read in a distributed manner into the corresponding logical partitions that are set.
[0065] Since the partitioning rules are determined for the incremental data table involved in the calculation, the incremental data in the incremental data table can be read into the corresponding logical partition set by the distributed computing engine according to the partitioning rules corresponding to the incremental data table. The incremental data in the incremental data table can be read into the corresponding logical partition set by the processing device in a distributed reading manner.
[0066] Exemplarily, the processing device includes 3 servers and there are 5 logical partitions. Then, two logical partitions can be set up in the first server to store the incremental data of the fund accounts with the tail numbers 1, 6, 2, and 7. Two logical partitions can be set up in the second server to store the incremental data of the fund accounts with the tail numbers 3, 8, 4, and 9. One logical partition can be set up in the third server to store the incremental data of the fund accounts with the tail numbers 5 and 0. Assume that the incremental data tables involved in the calculation include the incremental data of the fund accounts with the tail numbers 2, 6, 8, and 5. Then, the incremental data of the fund accounts with the tail numbers 2 and 6 will be read into the logical partitions set up in the first server, the incremental data of the fund account with the tail number 8 will be read into the logical partition of the second server, and the incremental data of the fund account with the tail number 5 will be read into the logical partition of the third server.
[0067] Step S103: Distributively read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database.
[0068] Among them, the stock data is generated by continuously accumulating the new results generated in the previous period from the user incremental data. For example, the stock data is the profit and loss data of the user. After each period of incremental data is obtained, the current period incremental data is added to the previous period profit and loss data to obtain the current period profit and loss data. In this way, multiple profit and loss data will be continuously obtained.
[0069] Similar to the reading method of the incremental data in the incremental data table, in this embodiment, when reading the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database, the distributive reading method is also adopted.
[0070] Step S104: Calculate the latest historical data associated with the incremental data in each logical partition and the incremental data in each logical partition.
[0071] For each piece of incremental data in each logical partition, calculate the latest historical data associated with this piece of incremental data and this piece of incremental data.
[0072] The data processing method provided by the embodiment of the present invention first determines the partitioning rule corresponding to the incremental data table according to the set computing parallelism and the partitioning field selected from the incremental data tables participating in the calculation, and then reads the incremental data in the incremental data table into the corresponding logical partitions set in a distributed manner according to the partitioning rule corresponding to the incremental data table. Then, the latest historical data associated with the incremental data in each logical partition is read in a distributed manner from the stock data of each user in the database. Finally, the latest historical data associated with the incremental data in each logical partition is calculated with the incremental data in each logical partition. The data processing method provided by the embodiment of the present invention can achieve distributed reading, writing, and calculation of data without changing the database storage mode, has high data processing efficiency for distributed reading, writing, and calculation of data, and can meet the high timeliness requirements of data processing.
[0073] In another embodiment of the present invention, the implementation process of "step S101: Determine the partitioning rule corresponding to the incremental data table according to the set computing parallelism and the partitioning field selected from the incremental data tables participating in the calculation" in the above embodiment is introduced.
[0074] Please refer to Figure 2 , which shows a schematic flowchart of determining the partitioning rule corresponding to the incremental data table according to the set computing parallelism and the partitioning field selected from the incremental data tables participating in the calculation, and may include:
[0075] Step S201: Based on the set computing parallelism, determine the number of logical partitions and determine the trailing truncation length of the selected partitioning field.
[0076] Among them, the number of logical partitions is the same as the number of the set computing parallelism. For example, if the set computing parallelism is 5, the number of logical partitions can be determined to be 5.
[0077] Among them, the process of determining the trailing truncation length of the selected partitioning field based on the set computing parallelism includes: If the set computing parallelism is expressed as n, the trailing truncation length d of the partitioning field is the number of digits of (n - 1).
[0078] Exemplarily, the partitioning field is the fund account number, and the set computing parallelism n = 3. Then the trailing truncation length d of the partitioning field is the number of digits of (3 - 1), that is, d = 1. d = 1 indicates that the fund account numbers with the last digits from 0 to 9 are to be allocated.
[0079] Step S202: Determine the trailing truncation range of the partitioning field according to the trailing truncation length of the partitioning field.
[0080] Specifically, if the trailing truncation length of the partitioning field is expressed as d, the trailing truncation range r of the partitioning field is 10 d .
[0081] Exemplarily, the degree of computing parallelism n = 3, that is, the number of logical partitions is 3, the partition field is the fund account number, the truncation length d at the end of the partition field is 1, and the truncation range r at the end of the partition field is 10 1 = 10, r = 10 indicates that the incremental data of the fund account numbers with the last digits 0 to 9 should be approximately evenly distributed into 3 logical partitions.
[0082] Step S203: Determine the incremental data divided into each logical partition according to the truncation range at the end of the partition field.
[0083] Through the above process, the partition rule corresponding to the incremental data table can be obtained, that is, several logical partitions are set for the incremental data table, and the incremental data divided into each logical partition.
[0084] Exemplarily, the degree of computing parallelism n = 3, that is, 3 logical partitions are set, the partition field is the fund account number, the truncation length d at the end of the partition field is 1, and the truncation range r at the end of the partition field is 1, that is, the fund account numbers with the last digits 0 to 9 are approximately evenly distributed into 3 partitions. For example, the incremental data of the fund account numbers with the last digits 0, 3, 6, and 9 are assigned to partition 1, the incremental data of the fund account numbers with the last digits 1, 4, and 7 are assigned to partition 2, and the incremental data of the fund account numbers with the last digits 2, 5, and 8 are assigned to partition 3.
[0085] In another embodiment of the present invention, the specific implementation process of "Step S102: According to the partition rule corresponding to the incremental data table, distributively read the incremental data in the incremental data table into the corresponding logical partitions set" in the above embodiment is introduced.
[0086] Please refer to Figure 3 , which shows a schematic flow chart of distributively reading the incremental data in the incremental data table into the corresponding logical partitions set according to the partition rule corresponding to the incremental data table, and may include:
[0087] Step S301: Determine the data query statement according to the partition rule corresponding to the incremental data table.
[0088] Suppose there is an incremental data table A, and the data query statement for reading table A, that is, the SQL statement, is "select * from A". Since the present invention introduces the strategy of logical partitions, therefore, SQL needs to be generated for each logical partition.
[0089] Specifically, the process of determining the data query statement according to the partition rule corresponding to the incremental data table may include:
[0090] Step S3011: Determine the incremental data to be read for each partition according to the partition rule corresponding to the incremental data table.
[0091] Exemplarily, there are three rows of data in the incremental data table A (the fund account field is included in table A), and the fund accounts in these three rows of data are 001, 002, and 003 respectively. Assume that the partitioning rule corresponding to the incremental data table is to set 3 logical partitions, namely partition 1, partition 2, and partition 3. The incremental data of the fund accounts with the last digits of 0, 3, 6, and 9 are assigned to partition 1, the incremental data of the fund accounts with the last digits of 1, 4, and 7 are assigned to partition 2, and the incremental data of the fund accounts with the last digits of 2, 5, and 8 are assigned to partition 3. Then, it can be determined that the incremental data to be read for partition 1 is the incremental data of the fund account 003, the incremental data to be read for partition 2 is the incremental data of the fund account 001, and the incremental data to be read for partition 3 is the incremental data of the fund account 002.
[0092] Step S3012: Generate a data query statement corresponding to each logical partition according to the incremental data to be read for each logical partition.
[0093] Specifically, according to the incremental data to be read for each logical partition, determine the predicate statement corresponding to each logical partition, and generate a data query statement corresponding to each logical partition according to the predicate statement corresponding to each logical partition.
[0094] Exemplarily, the incremental data table is table A (the fund account field is included in table A), the partitioning rule corresponding to the incremental data table indicates that two logical partitions are set, namely partition 1 and partition 2. The partitioning rule also indicates that the incremental data assigned to partition 1 is the incremental data of the fund accounts with the last digits of 1, 3, 5, 7, and 9, and the incremental data assigned to partition 2 is the incremental data of the fund accounts with the last digits of 0, 2, 4, 6, and 8. Then, it can be determined that the predicate statement corresponding to partition 1 is "substr(id, length(id) - 1, length(id)) in ('1', '3', '5', '7', '9')", and the predicate statement corresponding to partition 2 is "substr(id, length(id) - 1, length(id)) in ('0', '2', '4', '6', '8')". That is, the final data query statement corresponding to partition 1 is "select * from A where substr(id, length(id) - 1, length(id)) in ('1', '3', '5', '7', '9')", and the data query statement corresponding to partition 2 is "select * from A where substr(id, length(id) - 1, length(id)) in ('0', '2', '4', '6', '8')", where id represents the partitioning field, that is, the fund account field, and length(id) represents the length of the field value of the fund account field.
[0095] Step S302: Use a data query statement to distributively read the incremental quantity from the incremental data tables participating in the calculation.
[0096] For the above example, as Figure 4 shown, the distributed computing engine reads the incremental data of the fund accounts with the last digits of 0, 2, 4, 6, and 8 from the incremental data table A based on the data query statement "select * from A where substr(id, length(id)-1, length(id)) in ('0', '2', '4', '6', '8')", and reads the incremental data of the fund accounts with the last digits of 1, 3, 5, 7, and 9 from the incremental data table A based on the data query statement "select * from A where substr(id, length(id)-1, length(id)) in ('1', '3', '5', '7', '9')".
[0097] In another embodiment of the present invention, the specific implementation process of "Step S103: Distributively read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database" in the above embodiment is introduced.
[0098] There are various ways to distributively read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database. In one possible implementation, a query can be directly performed at the database level based on a query statement to obtain the latest historical data associated with the incremental data in each logical partition.
[0099] Considering that directly performing a query at the database level based on a query statement requires multiple complex table joins and window function sorting, etc., and the query speed is slow, the present invention proposes another more preferred implementation. Please refer to Figure 5 , which shows a flowchart of a preferred implementation for distributively reading the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database, and may include:
[0100] Step S501: Aggregate multiple incremental data with the same partition field information in each logical partition into one, and only retain one partition field information.
[0101] Wherein, the partition field information is the specific field value corresponding to the partition field.
[0102] The purpose of step S501 is to deduplicate the partition field information in each logical partition. Assuming that the partition field is the fund account number, if the dividend increment data and transaction increment data of the same fund account number are read into a logical partition, then it is necessary to deduplicate the fund account number fields of the dividend increment data and transaction increment data.
[0103] Step S502: Read all the stock data of the users to whom the increment data in each logical partition belongs from the stock data of each user in the database.
[0104] Specifically, the process of reading all the stock data of the users to whom the increment data in each logical partition belongs from the stock data of each user in the database may include:
[0105] Step S5021: Determine the number of interactions with the database based on the computing parallelism, the total number of partition field information in each logical partition, and the maximum query number of the data query statement.
[0106] Specifically, the calculation method of the number of interactions with the database can be: (the total number of partition field information in each logical partition - 1) / (the number of partitions * N) + 1, where the number of partitions is equal to the computing parallelism, and N is the maximum query number of the data query statement.
[0107] Step S5022: Based on the determined number of interactions, batch-read all the stock data of the users to whom the increment data in each logical partition belongs from the stock data of each user in the database.
[0108] For each piece of increment data in each logical partition, in this embodiment, a coarse-grained screening method is adopted to read all the stock data of the user to whom this piece of increment data belongs from the stock data of each user into the memory.
[0109] Step S503: Based on the set screening dimension, screen out the latest historical data associated with the increment data in each logical partition from all the stock data of the users to whom the increment data in each logical partition belongs.
[0110] After obtaining all the stock data of the users to whom the increment data in each logical partition belongs through coarse-grained screening, the obtained stock data can be further screened to obtain the latest historical data associated with the increment data in each logical partition.
[0111] Specifically, for each piece of increment data in each logical partition, based on the set screening dimension, screen out the latest historical data associated with this piece of increment data from all the stock data of the user to whom this piece of increment data belongs. Exemplarily, all the stock data of the user to whom this piece of increment data belongs can be screened from dimensions such as funds and institutions, and then sorted in reverse order by date to obtain the latest historical data associated with this piece of increment data.
[0112] After obtaining the latest historical data associated with each piece of incremental data, calculations can be performed on each piece of incremental data and the latest historical data associated therewith. When performing calculations, some other data may also be utilized, such as the net asset value table data of the fund. These data have no partitioning fields and are used in the calculations of data in each partition. To avoid frequent interactions with the database, these data can be read in one go and then broadcast, so that these data can participate in the data calculations of each partition.
[0113] An embodiment of the present invention also provides a data processing device. The data processing device provided by the embodiment of the present invention will be described below. The data processing device described below can be correspondingly referred to the data processing method described above.
[0114] Please refer to Figure 6 , which shows a schematic structural diagram of the data processing device provided by the embodiment of the present invention, and may include: a partitioning rule determination module 601, an incremental data reading module 602, a stock data reading module 603, and a data calculation module 604.
[0115] The partitioning rule determination module 601 is configured to determine the partitioning rule corresponding to the incremental data table according to the set calculation parallelism and the partitioning fields selected from the incremental data tables participating in the calculation.
[0116] Wherein, the partitioning rule is used to indicate the number of logical partitions set for the incremental data table, and the incremental data divided into each logical partition.
[0117] The incremental data reading module 602 is configured to distributively read the incremental data in the incremental data table into the corresponding logical partitions set according to the partitioning rule corresponding to the incremental data table.
[0118] The stock data reading module 603 is configured to distributively read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database.
[0119] The data calculation module 604 is configured to calculate the latest historical data associated with the incremental data in each logical partition and the incremental data in each logical partition.
[0120] Optionally, the partitioning rule determination module 601 may include: a partitioning number determination sub-module, an end truncation length determination sub-module, an end truncation range determination sub-module, and a partitioning rule determination sub-module.
[0121] The partitioning number determination sub-module is configured to determine the number of logical partitions according to the calculation parallelism.
[0122] The end truncation length determination sub-module is used to determine the end truncation length of the partition field according to the calculation parallelism.
[0123] The end truncation range determination sub-module is used to determine the end truncation range of the partition field according to the end truncation length of the partition field.
[0124] The partition rule determination sub-module is used to determine the incremental data divided into each logical partition according to the end truncation range of the partition field, so as to obtain the partition rule corresponding to the incremental data table.
[0125] Optionally, the incremental data reading module 602 may include: a query statement generation sub-module and a data reading sub-module;
[0126] The query statement generation sub-module is used to generate a data query statement according to the partition rule corresponding to the incremental data table;
[0127] The data reading sub-module is used to distributively read incremental data from the incremental data table into the corresponding logical partitions set by using the data query statement.
[0128] Optionally, when the incremental data reading module 602 generates a data query statement according to the partition rule corresponding to the incremental data table, it is specifically used for:
[0129] Determine the incremental data to be read for each logical partition according to the partition rule corresponding to the incremental data table;
[0130] Generate a data query statement corresponding to each logical partition according to the incremental data to be read for each logical partition.
[0131] Optionally, the stock data reading module 603 may include: a data preprocessing sub-module, a stock data reading sub-module, and a data screening sub-module.
[0132] The data preprocessing sub-module is used to summarize multiple incremental data with the same partition field information in each logical partition into one, and only retain one partition field information, where the partition field information is the specific field information corresponding to the partition field.
[0133] The stock data reading sub-module is used to read all the stock data of the users to whom the incremental data in each logical partition belongs from the stock data of each user in the database.
[0134] The data screening sub-module is used to screen out the latest historical data associated with the incremental data in each logical partition from all the stock data of the users to whom the incremental data in each logical partition belongs based on the set screening dimension.
[0135] Optionally, when the stock data reading module 603 reads all the stock data of the users to which the incremental data in each logical partition belongs from the stock data of each user in the database, it is specifically configured to:
[0136] Determine the number of interactions with the database based on the computing parallelism, the total number of partition field information in each logical partition, and the maximum query number of the data query statement;
[0137] Based on the number of interactions, batch-read all the stock data of the users to which the incremental data in each logical partition belongs from the stock data of each user in the database.
[0138] The data processing device provided by the embodiments of the present invention first determines the partition rule corresponding to the incremental data table according to the set computing parallelism and the partition fields selected from the incremental data tables participating in the calculation, and then distributes and reads the incremental data in the incremental data table into the corresponding logical partitions set according to the partition rule corresponding to the incremental data table. Then, it distributes and reads the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database. Finally, it calculates the latest historical data associated with the incremental data in each logical partition and the incremental data in each logical partition. The data processing device provided by the embodiments of the present invention can realize distributed reading, writing and calculation of data without changing the database storage mode, has high data processing efficiency for distributed reading, writing and calculation of data, and can meet the high timeliness requirements of data processing.
[0139] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0140] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0141] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data processing method, characterized in that, applied to a processing device, the method includes: According to the set computing parallelism and the partition field selected from the incremental data tables participating in the calculation, determine the partition rule corresponding to the incremental data table, including: According to the computing parallelism, determine the number of logical partitions, and determine the end truncation length of the partition field; According to the end truncation length of the partition field, determine the end truncation range of the partition field; According to the end truncation range of the partition field, determine the incremental data divided into each logical partition to obtain the partition rule corresponding to the incremental data table; wherein, the partition rule is used to indicate the number of logical partitions set for the incremental data table, and the incremental data divided into each logical partition; According to the partition rule corresponding to the incremental data table, distribute and read the incremental data in the incremental data table into the corresponding logical partitions set; Distribute and read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database; Perform calculations on the latest historical data associated with the incremental data in each logical partition and the incremental data in each logical partition.
2. The data processing method according to claim 1, characterized in that, The step of distributing and reading the incremental data in the incremental data table into the corresponding logical partitions set according to the partition rule corresponding to the incremental data table includes: Generate a data query statement according to the partition rule corresponding to the incremental data table; Use the data query statement to distribute and read the incremental quantity from the incremental data table into the corresponding logical partitions set.
3. The data processing method according to claim 2, characterized in that, The step of generating a data query statement according to the partition rule corresponding to the incremental data table includes: According to the partition rule corresponding to the incremental data table, determine the incremental data to be read for each logical partition; Generate a data query statement corresponding to each logical partition according to the incremental data to be read for each logical partition.
4. The data processing method according to claim 1, characterized in that, The step of distributing and reading the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database includes: Summarize multiple incremental data with the same partition field information in each logical partition into one, and only retain one partition field information, where the partition field information is the specific field value corresponding to the partition field; Read all the stock data of the users to which the incremental data in each logical partition belongs from the stock data of each user in the database; Based on the set screening dimension, screen out the latest historical data associated with the incremental data in each logical partition from all the stock data of the users to which the incremental data in each logical partition belongs.
5. The data processing method according to claim 4, characterized in that, The step of reading all the stock data of the users to which the incremental data in each logical partition belongs from the stock data of each user in the database includes: Determine the number of interactions with the database based on the computing parallelism, the total number of partition field information in each logical partition, and the maximum number of queries in the data query statement; Based on the number of interactions, batch-read all the stock data of the users to whom the incremental data in each logical partition belongs from the stock data of each user in the database.
6. A data processing device, characterized in that, it includes: a partition rule determination module, an incremental data reading module, a stock data reading module, and a data calculation module; The partition rule determination module is used to determine the partition rule corresponding to the incremental data table according to the set computing parallelism and the partition fields selected from the incremental data tables participating in the calculation, wherein the partition rule is used to indicate the number of logical partitions set for the incremental data table, and the incremental data divided into each logical partition; The partition rule determination module includes: a partition number determination sub-module, an end truncation length determination sub-module, an end truncation range determination sub-module, and a partition rule determination sub-module; The partition number determination sub-module is used to determine the number of logical partitions according to the computing parallelism; The end truncation length determination sub-module is used to determine the end truncation length of the partition field according to the computing parallelism; The end truncation range determination sub-module is used to determine the end truncation range of the partition field according to the end truncation length of the partition field; The partition rule determination sub-module is used to determine the incremental data divided into each logical partition according to the end truncation range of the partition field, so as to obtain the partition rule corresponding to the incremental data table; The incremental data reading module is used to distributively read the incremental data in the incremental data table into the corresponding logical partitions set according to the partition rule corresponding to the incremental data table; The stock data reading module is used to distributively read the latest historical data associated with the incremental data in each logical partition from the stock data of each user in the database; The data calculation module is used to calculate the latest historical data associated with the incremental data in each logical partition and the incremental data in each logical partition.
7. The data processing device according to claim 6, characterized in that, The incremental data reading module includes: a query statement generation sub-module and a data reading sub-module; The query statement generation sub-module is used to generate a data query statement according to the partition rule corresponding to the incremental data table; The data reading sub-module is used to distributively read the incremental quantity from the incremental data table into the corresponding logical partitions set by using the data query statement.
8. The data processing device according to claim 6, characterized in that, The stock data reading module includes: a data preprocessing sub-module, a stock data reading sub-module, and a data screening sub-module; The data preprocessing sub-module is used to summarize multiple incremental data with the same partition field information in each logical partition into one, and only retain one partition field information, and the partition field information is the specific field information corresponding to the partition field; The stock data reading sub-module is used to read all the stock data of the users to whom the incremental data in each logical partition belongs from the stock data of each user in the database; The data screening sub-module is used to screen out the latest historical data associated with the incremental data in each logical partition from all the stock data of the users to whom the incremental data in each logical partition belongs based on the set screening dimensions.
Citation Information
Patent Citations
A partition-level connection method and apparatus for a distributed database
CN108959510A
Database early parallelism method and system
US20050131893A1