Mechanism data summarization method and system in distributed mode and storage medium
By designing an organization summary parameter table that simulates a broadcast table in a distributed database and combining the step-by-step and one-step summary methods, the problem of low efficiency of organization data summary in a distributed database is solved, and an efficient and flexible data summary process is achieved.
Patent Information
- Application Number
- CN202510717436.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-12
AI Technical Summary
In distributed databases, there are problems such as limited concurrency, low execution efficiency and difficulty in avoiding distributed transactions when aggregating institutional data, especially when aggregating large amounts of data, the efficiency is extremely low.
By designing an organization summary parameter table that simulates the broadcast table based on the organization relationship tree, a method combining step-by-step and one-step aggregation is adopted, and intermediate tables are used to increase concurrency and reduce cross-shard queries, thus avoiding distributed transactions and improving data aggregation efficiency.
It achieves efficient and flexible aggregation of institutional data in a distributed mode, reduces the workload of node configuration, improves query performance and concurrency, and fully utilizes distributed resources.
Smart Images

Figure CN120631976A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technology, and in particular to a method, system, and storage medium for aggregating institutional data in a distributed mode. Background Art
[0002] Distributed databases have key distributed capabilities such as data sharding management, distributed transactions, and read-write separation. They have a usage similar to that of centralized databases and can reduce the complexity of implementing distributed transformation of applications. They have been widely used in the financial field.
[0003] Due to the inherent limitations of distributed databases, aggregating large amounts of data across shards is extremely difficult. In the context of institutional data aggregation, because data in distributed databases is stored in different shards according to the shard key, traditional institutional data aggregation methods, such as concurrent, hierarchical aggregation by province, have very limited concurrency and cannot avoid distributed transactions, resulting in low execution efficiency. Summary of the Invention
[0004] The present invention provides a method, system and storage medium for aggregating institutional data in a distributed mode, which can improve the efficiency of aggregating institutional data in a distributed mode.
[0005] In a first aspect, an embodiment of the present invention provides a method for aggregating mechanism data in a distributed mode, which is applied to a mechanism data aggregation system, wherein the system includes at least a processing node, a first generation node, a first processing node, a second generation node, a second processing node, and an aggregation node, and the method includes:
[0006] By processing the nodes, an organization summary parameter table designed as a simulated broadcast table is obtained based on the organization relationship tree. The organization summary parameter table includes the organization code and shard key corresponding to each organization and its parent organization;
[0007] A first task is generated by a first generation node and processed by a first processing node. The first task is used to aggregate the organizational data of the grassroots organizations in the source table to be aggregated one level up to obtain an intermediate table based on the organizational aggregation parameter table. The source table to be aggregated includes the organizational code, shard key, and organizational data of each organization.
[0008] A second task is generated by a second generation node and processed by a second processing node; the second task is used to add, to the intermediate table, the institutional data of each institution in the intermediate table to the data of the corresponding superior institutions in one step, based on the institutional summary parameter table, to obtain an aggregated intermediate table;
[0009] The organization data belonging to the same organization in the aggregated intermediate table are aggregated through the aggregation node to obtain a target table.
[0010] In a second aspect, an embodiment of the present invention provides a mechanism data aggregation system, including a processing node, a first generation node, a first processing node, a second generation node, a second processing node, and an aggregation node;
[0011] A processing node is used to process the organization relationship tree to obtain an organization summary parameter table designed as a simulated broadcast table. The organization summary parameter table includes the organization code and shard key corresponding to each organization and its parent organization;
[0012] The first generation node is used to generate a first task, and the first processing node is used to process the first task. The first task is used to aggregate the organizational data of the grassroots organizations in the source table to be aggregated one level up to obtain an intermediate table according to the organizational aggregation parameter table. The source table to be aggregated includes the organizational code, shard key, and organizational data of each organization.
[0013] The second generation node is used to generate a second task, and the second processing node is used to process the second task; the second task is used to, based on the organization summary parameter table, add the organization data of each organization in the intermediate table to the data of each corresponding superior organization in one step in the intermediate table to obtain a summarized intermediate table;
[0014] The aggregation node is used to aggregate the organization data belonging to the same organization in the aggregated intermediate table to obtain a target table.
[0015] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0016] The technical solution of the embodiment of the present invention is to obtain an organization summary parameter table designed as a simulated broadcast table based on the organization relationship tree processing. According to the organization summary parameter table, the organization data of the grassroots organization with the largest data volume in the source table to be summarized is aggregated one level upward to obtain an intermediate table. The organization number data of each organization in the intermediate table is then aggregated in one step to each superior organization of each organization to obtain a summarized intermediate table. The organization data belonging to the same organization in the summarized intermediate table is then aggregated to obtain a target table. This solution combines step-by-step aggregation and one-step aggregation by simulating a broadcast table, and then uses the intermediate table to increase concurrency, reduce cross-shard queries, avoid cross-shard writes, and can fully utilize distributed resources without generating distributed transactions, thereby improving the efficiency of organization data aggregation in a distributed mode.
[0017] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 This is a flow chart of a method for aggregating institutional data in a distributed mode provided in accordance with the first embodiment of the present invention;
[0020] Figure 2 This is a flow chart of a method for aggregating institutional data in a distributed mode provided in accordance with a second embodiment of the present invention;
[0021] Figure 3 This is a schematic diagram of an organization relationship tree provided according to the second embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0023] It should be noted that the terms "first," "second," and the like in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.
[0024] First, the following content explains the relevant terms involved in the present invention:
[0025] Distributed mode: Distributed is a computing and system architecture mode. In this invention, it means deploying the program on multiple computer nodes, communicating and collaborating through the network. Each computer node processes different data, and splits the processing of all data from one server to multiple servers processing different data to improve processing efficiency.
[0026] Institutional Data Aggregation: Institutions can be branches of a bank, forming part of the bank. Examples include the head office, provincial branches (first-tier branches), second-tier branches, sub-branches, and outlets, all of which are considered bank branches and have hierarchical management relationships. Institutional data refers to the business-related data generated by these branches. Institutional data aggregation is the process of aggregating this institutional data from lower-level branches to higher-level branches.
[0027] Database sharding is the process of storing large databases across multiple machines. A single computer or database server can only store and process a limited amount of data. Database sharding overcomes this limitation by splitting data into smaller chunks (called shards) and storing them across multiple database servers. All database servers typically have the same underlying technology, working together to store and process large amounts of data.
[0028] Shard key: A shard key determines how a dataset is partitioned. A column in the dataset determines which rows are grouped together to form a shard. The database designer can select a shard key from an existing column or create a new one.
[0029] Transaction: It is a sequence of database operations that access and possibly operate on various data items. These operations are either all executed or none of them are executed. It is an indivisible unit of work.
[0030] Distributed transaction: refers to the transaction participants, transaction supporting servers, resource servers and transaction managers located on different nodes of different distributed systems.
[0031] Batch program: combines a series of programs together in sequence without user interaction. It schedules and processes the corresponding programs according to the work steps and execution sequence set by the system. It is generally used to process large amounts of data and needs to process the data step by step.
[0032] Batch node: A scheduling configuration unit for a batch task, usually including batch tasks, scheduling time, execution parameters and other key configurations.
[0033] In a distributed architecture, there are two common approaches to aggregating institutional data: one is hierarchical aggregation, and the other is one-step aggregation to the aggregation bank at each level. The following uses the example of the head office, provincial branch (i.e., first-level branch), second-level branch, sub-branch, and outlet to illustrate hierarchical aggregation and one-step aggregation.
[0034] Gradual aggregation can be understood as aggregating from the lowest level of institutions one level at a time up to the head office. Gradual aggregation can be done in a provincial concurrency manner, that is, provincial concurrency below the head office, and aggregation from the lowest level of outlets within the province to the first level branch, i.e., provincial branch. After all provincial branches are gathered, aggregation is finally done from the provincial branch to the head office. This concurrency method has low concurrency, making it difficult to utilize distributed resources, and will generate distributed transactions, which should be avoided as much as possible. Another concurrency method is to dynamically generate concurrent tasks. Compared with static provincial concurrency, it can utilize more distributed resources, and tasks do not cross shards, avoiding distributed transactions. The disadvantage of this concurrency method is that it requires the configuration of a large number of batch nodes (the number of nodes is twice the maximum number of institutional levels - 1, and the aggregation of institutional data at each level requires a dynamic task generation node and an institutional data aggregation node) and the need to associate two master data tables during aggregation, which makes query efficiency low.
[0035] Another approach to institutional aggregation is to aggregate data to all levels of summary banks in a single step. This requires first generating a flattened institutional aggregation relationship table. By linking this table, data can be aggregated to all levels of summary institutions, including the head office, in a single step. This aggregation method allows aggregation of each level of institutions to be completed by a single batch node, without the need for coordination from other batch nodes. Therefore, data can be executed in parallel without waiting. However, due to the massive amount of data involved in querying across all shards for the top-level institutions, this approach is extremely inefficient in a distributed architecture and can suffer from an extremely long tail.
[0036] In a distributed architecture, the one-step parallel aggregation solution is generally not used in big data scenarios due to the long tail phenomenon. When the amount of data to be aggregated by an organization is large, a serial level-by-level aggregation method is usually used for organization aggregation, but this may have the following drawbacks:
[0037] 1. There are many nodes and they can only be executed serially. For example, when the organization level is 5, at least 8 batch nodes are required (4 batch nodes for generating dynamic tasks and 4 batch nodes for organization summary). In actual applications, the organization level is not simply 5 levels, but more than a dozen levels, and the number of levels is not fixed (maintained online by business personnel). Assuming there are 15 levels of organizations, 28 batch nodes need to be configured. If one more level is reserved, 30 batch nodes need to be configured, and all of them are serial. If there are as many as 45 organization summary scenarios in batch processing, the total node configuration amount will be very large (45*30) and will consume a lot of batch scheduling time.
[0038] Second, poor query performance. During the institutional data aggregation process, each level of aggregation organization may have directly affiliated basic banks (outlets). Therefore, except for the first aggregation node, which only needs to associate one table to be aggregated, all other levels need to associate the table to be aggregated with the aggregated table. This results in poor query performance.
[0039] In summary, this solution has no advantages, whether from the perspective of node configuration workload or overall performance. To overcome the shortcomings of the above-mentioned technology, the present invention provides a method for aggregating institutional data in a distributed mode, which only requires the configuration of 6 nodes, is not affected by the number of institutional levels, supports dynamic concurrency, has a configurable number of concurrent nodes, is highly efficient, and can fully utilize distributed resources.
[0040] Example 1
[0041] Figure 1 This is a flowchart of a method for aggregating institutional data in a distributed mode, according to a first embodiment of the present invention. This embodiment is applicable to aggregating institutional data in a distributed mode, and the method can be applied to an institutional data aggregation system. The institutional data aggregation system includes at least a processing node, a first generation node, a first processing node, a second generation node, a second processing node, and an aggregation node. The nodes can collaborate via network communication.
[0042] like Figure 1 As shown, the method includes:
[0043] S110. By processing the nodes, an organization summary parameter table designed as a simulated broadcast table is obtained based on the organization relationship tree. The organization summary parameter table includes the organization code and shard key corresponding to each organization and its parent organization.
[0044] The processing node may be a node of the mechanism summary parameter table obtained by processing based on the mechanism relationship tree.
[0045] An institution relationship tree can be understood as a graphical tool that uses a tree-like structure to visually display the hierarchical relationships between multiple institutions. The root node can be the highest-level institution at each level, such as the head office. Below the root node can be one or more first-level branches; below each first-level branch can be one or more second-level branches; below each second-level branch can be one or more sub-branches; and below each sub-branches can be one or more outlets. In this context, outlets are grassroots institutions, the lowest level of institutions at each level. The above is an example of an institution relationship tree; the specific structure of the institution relationship tree can be determined based on actual application needs.
[0046] The organization code can be a code used to uniquely identify an organization. Different organizations have different organization codes, and there is no limitation on the encoding method of the organization code.
[0047] An organization's shard key may indicate the shard that stores the organization's data. Shard keys may be pre-assigned to each organization based on established rules, with no specific restrictions. Optionally, different organizations may be assigned the same or different shard keys.
[0048] In this step, a predetermined organization relationship tree can be obtained from a server or locally by processing nodes, and the hierarchical relationships between the organizations in the organization relationship tree, i.e., the upper and lower hierarchical relationships, can be determined. Based on the determined hierarchical relationships between the organizations, for each organization, the organization code and sharding key of the organization and the organization code and sharding key of the organization's parent organization are integrated. After integrating all organizations, an organization summary parameter table is obtained. The organization summary parameter table can be a table indicating the hierarchical relationships between multiple organizations and the relevant parameters of each organization.
[0049] Furthermore, to minimize cross-shard access, the defined organization summary parameter table is designed as a simulated broadcast table. This means that each shard in the organization data aggregation system has a copy of the organization summary parameter table. For example, if the organization data aggregation system has more than a dozen shard keys, each of these shard keys can store the same organization summary parameter table, reducing cross-shard queries in the distributed database.
[0050] S120. Generate a first task through a first generation node and process the first task through a first processing node. The first task is used to aggregate the organizational data of the grassroots organizations in the source table to be aggregated to obtain an intermediate table according to the organizational aggregation parameter table. The source table to be aggregated includes the organizational code, shard key and organizational data of each organization.
[0051] The first generating node may be a node that generates the first task, and the first processing node may be a node that processes the first task.
[0052] Optionally, both the first generation node and the first processing node support concurrent processing and can be understood as batch nodes.
[0053] The first generation node can dynamically generate concurrent tasks based on the data in the source table to be summarized, and the first processing node can concurrently process the dynamic tasks, and the number of concurrent tasks is adjustable.
[0054] The first task is described as follows:
[0055] The source table to be aggregated can be a source table that stores the organization data to be aggregated. Each row in the source table can correspond to an organization and store the organization code, shard key, and organization data. The source table to be aggregated can be pre-stored locally or on a server.
[0056] The intermediate table can be a table obtained by summarizing the grassroots organizations' organizational data in the source table. Grassroots organizations are the lowest level organizations among all levels of organizations, such as branches.
[0057] According to the institutional summary parameter table, the institutional data of the grassroots institutions in the source table to be summarized are aggregated one level to obtain an intermediate table. It can be understood that the institutional code of the grassroots institution and the upper-level institution of the grassroots institution is determined through the institutional summary parameter table. Through the determined institutional code, the institutional data of the grassroots institution is aggregated to its upper-level institution on the basis of the source table to be summarized. For example, the institutional data of each branch is aggregated to the corresponding branch, and the institutional code and sharding key of the upper-level institution are completed. The corresponding content of other institutions except the grassroots institutions in the source table to be summarized remains unchanged. Afterwards, the sharding key of the data source institution can be added to each institutional data in the table to obtain the intermediate table.
[0058] Among them, the data source agency indicates the source of the agency data. For example, if the data of a certain agency is obtained by collecting the agency data of one or more grassroots agencies one level higher, then the data source agency of the agency data is the one or more grassroots agencies; for another example, if the data of a certain agency is the agency data of a certain agency, then the data source agency is the agency.
[0059] S130. Generate a second task through a second generation node and process the second task through a second processing node; the second task is used to add, according to the organization summary parameter table, to the intermediate table, the organization data of each organization in the intermediate table to the data of the corresponding superior organizations in one step, to obtain a summarized intermediate table.
[0060] The second generating node may be a node that generates the second task, and the second processing node may be a node that processes the second task.
[0061] Optionally, both the second generation node and the second processing node support concurrent processing and can be understood as batch nodes.
[0062] The second generation node can generate concurrent tasks based on the data in the intermediate table, and the second processing node can concurrently process dynamic tasks, and the number of concurrent tasks is adjustable.
[0063] The second task is described as follows:
[0064] The aggregated intermediate table may be a table obtained by aggregating the institutional data of each institution in the intermediate table to the corresponding superior institutions in one step.
[0065] According to the institution summary parameter table, the intermediate table is newly added to aggregate the institution data of each institution in the intermediate table to the data of the corresponding superior institutions in one step to obtain the summarized intermediate table. It can be understood that any institution in the intermediate table is determined as the target institution, and the institution code and shard key of one or more superior institutions of the target institution are determined from the institution summary parameter table; one or more rows are added to the intermediate table, and each newly added row can correspond to a superior institution of the target institution. The row may include the institution code and shard key of a superior institution of the target institution, the shard key of the target institution as the data source institution, and the institution data of the target institution aggregated to the superior institution; when all the institutions in the intermediate table are processed in the above manner, the summarized intermediate table is obtained. Among them, if there are multiple superior institutions corresponding to the target institution, a new row is added for each superior institution when adding a row.
[0066] In actual application, this step can be understood as summarizing the institutional data corresponding to each branch to the corresponding second-level branch, first-level branch, and head office in one step; summarizing the institutional data corresponding to each second-level branch to the corresponding first-level branch and head office in one step; and summarizing the institutional data corresponding to each first-level branch to the head office in one step.
[0067] S140 , aggregating the organization data belonging to the same organization in the aggregated intermediate table through an aggregation node to obtain a target table.
[0068] The aggregation node aggregates the organization data in the intermediate table after aggregation. The result of the aggregation becomes the target table. Optionally, the aggregation node supports concurrent processing. Tasks in this step can be configured to support concurrency by shard key or by organization, depending on the data volume. General concurrent tasks can be used, and dynamic concurrency is not required.
[0069] In this step, through the aggregation node, the organization data corresponding to the same organization code can be determined as organization data belonging to the same organization according to the organization codes in the aggregated intermediate table, and the organization data belonging to the same organization can be aggregated to obtain the target table.
[0070] The technical solution of the embodiment of the present invention is to obtain an organization summary parameter table designed as a simulated broadcast table based on the organization relationship tree processing. According to the organization summary parameter table, the organization data of the grassroots organization with the largest data volume in the source table to be summarized is aggregated one level upward to obtain an intermediate table. The organization number data of each organization in the intermediate table is then aggregated in one step to each superior organization of each organization to obtain a summarized intermediate table. The organization data belonging to the same organization in the summarized intermediate table is then aggregated to obtain a target table. This solution combines step-by-step aggregation and one-step aggregation by simulating a broadcast table, and then uses the intermediate table to increase concurrency, reduce cross-shard queries, avoid cross-shard writes, and can fully utilize distributed resources without generating distributed transactions, thereby improving the efficiency of organization data aggregation in a distributed mode.
[0071] Example 2
[0072] Figure 2 This is a flow chart of a method for aggregating institutional data in a distributed mode according to the second embodiment of the present invention. This embodiment is a further refinement based on the above-mentioned first embodiment. Figure 2 As shown, the method includes:
[0073] S210 , by processing the nodes, based on the organization relationship tree, an organization summary parameter table designed as a simulated broadcast table is obtained.
[0074] In one embodiment, the mechanism summary parameter table includes:
[0075] The grassroots organization row is used to record the organization code and shard key of the grassroots organization, the organization code and shard key of the grassroots organization's parent organization, and the corresponding aggregation level of the parent organization;
[0076] Other institution rows are used to record the institution code and sharding key of any other institution except the grassroots institution, the institution code and sharding key of each superior institution of any other institution, and the summary level corresponding to each superior institution of any other institution.
[0077] Among them, each grassroots organization corresponds to a grassroots organization row. The organization code of the grassroots organization in the grassroots organization row is used as the organization code of this level, the sharding key of the grassroots organization is used as the sharding key of this level organization, the organization code of the upper-level organization of the grassroots organization is used as the upper-level organization code, the sharding key of the upper-level organization of the grassroots organization is used as the upper-level organization sharding key, and the level corresponding to the upper-level organization of the grassroots organization is used as the summary level.
[0078] Any other organization, except for grassroots organizations, can correspond to one or more other organization rows. For each other organization row corresponding to any other organization, the organization code of the other organization is used as the organization code at the current level, the shard key of the other organization is used as the shard key of the current level, the organization code of any parent organization of the other organization is used as the parent organization code, the shard key of any parent organization of the other organization is used as the parent organization shard key, and the level corresponding to any parent organization of the other organization is used as the summary level.
[0079] Figure 3 Schematic diagram of an organization relationship tree provided according to the second embodiment of the present invention. Figure 3As shown, institutions are divided into five levels from bottom to top: outlets, branches, second-tier branches, first-tier branches, and the head office. The head office is the highest-level institution and includes first-tier branches 1 and 2; first-tier branch 1 includes second-tier branch 1, which includes branch 1; branch 1 includes outlet 1 and outlet 2; first-tier branch 2 includes second-tier branch 2; second-tier branch 2 includes branch 2; and branch 2 includes outlet 3. This institutional relationship tree is for illustrative purposes only and can be adjusted based on actual application needs, so it is not limited here.
[0080] based on Figure 3 The institution relationship tree shown in the figure generates an institution summary parameter table as shown in Table 1. Among them, the rows with the institution codes of this level as branch 1, branch 2, and branch 3 are grassroots institution rows, and the rows displayed other than the grassroots institution rows are other institution rows.
[0081] Table 1 Summary of Mechanism Parameters
[0082]
[0083] S220. Generate a first task through a first generation node, and process the first task through a first processing node; the first task is used to determine the organization codes corresponding to the grassroots organization and the upper-level organization of the grassroots organization according to the organization summary parameter table; according to the organization codes corresponding to the grassroots organization and the upper-level organization of the grassroots organization, delete the first deleted row corresponding to the grassroots organization in the source table to be summarized, and add a first added row corresponding to the upper-level organization, the first added row includes the organization code of the upper-level organization, the sharding key of the upper-level organization and the organization data of the grassroots organization uploaded to the upper-level organization, to obtain a candidate table; in each row of the candidate table, add the sharding key of the data source organization to obtain an intermediate table, and the data source organization is the data source of the organization data in this row.
[0084] In this step, according to the institution summary parameter table, the institution code of each grassroots institution and the institution code of the grassroots institution's upper-level institution are determined, and it is determined which grassroots institution or grassroots institutions correspond to the same upper-level institution, such as branch 1 and branch 2 both correspond to branch 1, and branch 3 corresponds to branch 2; according to the aforementioned determined institution code, the first deleted row of the grassroots institutions corresponding to the same upper-level institution is deleted in the source table to be summarized. The first deleted row is the row corresponding to the grassroots institution to be deleted, that is, the current-level institution code of the row corresponds to the grassroots institution to be deleted; and the first added row of the same upper-level institution is added to the source table to be summarized. In the first added row, the institution code of the same upper-level institution is the current-level institution code, the shard key of the same upper-level institution is the current-level institution shard key, and the institution data is the institution data summarized from the grassroots institutions corresponding to the same upper-level institution, thereby obtaining a candidate table; in each row of the candidate table, the shard key of the data source institution is added to obtain an intermediate table, and the data source institution is the data source of the institution data in this row.
[0085] For example, Table 2 is the source table to be summarized, in which only part of the data is displayed, and there is also data that is not displayed.
[0086] Table 2 Source table to be summarized
[0087]
[0088] Based on Table 2 and Table 1, Table 3 is an exemplary description of the intermediate table.
[0089] Table 3 Intermediate table
[0090]
[0091] That is, based on Table 2, delete the rows in Table 2 where the institution code for this level is outlet 1 and outlet 2, and add a new row with the institution code for this level being branch 1, the sharding key for this level being branch 1 being branch 1 sharding key 2, and the institution data being the combined result of 10 for outlet 1 and 20 for outlet 2; delete the row in Table 2 where the institution code for this level being outlet 3, and add a new row with the institution code for this level being branch 2, the sharding key for this level being branch 2 being branch 2 sharding key 6, and the institution data being outlet 3; add the sharding key of the data source institution to each row to obtain the intermediate table.
[0092] It should be noted that the shard key of the organization to be aggregated is retained as the primary key in the intermediate table. This design ensures that the data processed by each task will not have repeated primary keys, making it suitable for concurrent processing.
[0093] S230. Generate a second task through a second generation node, and process the second task through a second processing node; the second task is used to determine each institution in the intermediate table as a target institution, and determine the institution code and shard key of each superior institution of the target institution according to the institution summary parameter table; add a second new row to the intermediate table, the second new row includes the institution code and shard key of a superior institution of the target institution, the shard key of the target institution and the institution data summarized by the institution data of the target institution, to obtain a summarized intermediate table.
[0094] In this step, each institution in the intermediate table is determined as a target institution, and the institution code and shard key of each superior institution of the target institution are determined according to the institution summary parameter table; for each target institution, one or more second new rows are added to the intermediate table, and each second new row corresponds to a superior institution of the target institution. The institution code of the superior institution in the row is the institution code of this level, the shard key of the superior institution is the shard key of the institution at this level, the shard key of the target institution is the shard key of the data source institution, and the institution data is the institution data summarized by the target institution; after performing the above operations on each target institution, a summarized intermediate table is obtained.
[0095] Based on Table 3, Table 4 is an exemplary description of the summarized intermediate table, which only shows part of the data.
[0096] Table 4 Summarized intermediate table
[0097]
[0098] Among them, Branch 1 is determined as the target institution, and the superior institutions corresponding to Branch 1 are Second-level Branch 1, First-level Branch 1 and the Head Office. Therefore, three second-added rows need to be added for Branch 1, which are rows 4 to 6 in Table 4 respectively; Branch 2 is determined as the target institution, and the superior institutions corresponding to Branch 2 are Second-level Branch 2, First-level Branch 2 and the Head Office. Therefore, three second-added rows need to be added for Branch 2, which are rows 7 to 9 in Table 4 respectively.
[0099] S240. Through the aggregation node, the rows in the aggregated intermediate table are merged according to the principle of consistent organization codes, and the organization data of the rows involved in the merging are aggregated to obtain the target table.
[0100] In this step, the rows with the same institution code at the same level in the aggregated intermediate table are merged through the aggregation node. During the merging, the institution data of the rows involved in the merging are aggregated to obtain the target table.
[0101] Based on Table 4, Table 5 is an exemplary description of the target table.
[0102] Table 5 Target table
[0103]
[0104] In one embodiment, the organization summary parameter table, the intermediate table, the summarized intermediate table, and the target table further include one or more of a summary date and an organization summary scenario number.
[0105] The technical solution of the embodiment of the present invention may include the following advantages:
[0106] (1) Redundant parameter support. A redundant organization summary parameter table is generated through the organization relationship tree, and the parameter implementation is simulated by the broadcast table design. This design avoids cross-shard data access as much as possible, providing support for flexible and high-performance organization aggregation that combines subsequent level-by-level aggregation and one-step aggregation.
[0107] (2) Simplified node configuration. The organization aggregation supports an unlimited number of organization levels and source organization types, and is dynamic and flexible in configuration. Compared with the existing organization aggregation scenario that requires the configuration of nearly 30 nodes, the present invention only requires the configuration of 5 batch nodes and one processing node. It also supports aggregation from any level, without the need to configure fixed-level aggregation nodes in advance, and without the need for a backup node.
[0108] (3) Simulated broadcast improves performance. The three-faceted design ensures high efficiency in executing a single node for institutional aggregation. By splitting tasks into smaller pieces along a certain dimension, the long-tail problem and the limited concurrency problem are solved. During the entire aggregation process, only one large data table (source table or intermediate table) needs to be associated, resulting in high query efficiency. The simulated broadcast design method avoids cross-shard queries as much as possible, improving performance. Furthermore, the number of serial nodes in the entire institutional aggregation process is small, which greatly improves the efficiency of institutional data aggregation in the same scenario.
[0109] Example 3
[0110] An embodiment of the present invention provides an organization data aggregation system, which can be applied to the situation of organization data aggregation in a distributed mode. The system includes a processing node, a first generation node, a first processing node, a second generation node, a second processing node and an aggregation node. The nodes communicate and collaborate through the network. There is no limitation on the connection relationship between the nodes, as long as communication can be achieved.
[0111] A processing node is used to process an organization summary parameter table designed as a simulated broadcast table based on the organization relationship tree. The organization summary parameter table includes the organization code and shard key corresponding to each organization and its parent organization;
[0112] The first generation node is used to generate a first task, and the first processing node is used to process the first task. The first task is used to aggregate the organizational data of the grassroots organizations in the source table to be aggregated one level up to obtain an intermediate table according to the organizational aggregation parameter table. The source table to be aggregated includes the organizational code, shard key, and organizational data of each organization.
[0113] The second generation node is used to generate a second task, and the second processing node is used to process the second task; the second task is used to, based on the organization summary parameter table, add the organization data of each organization in the intermediate table to the data of each corresponding superior organization in one step in the intermediate table to obtain a summarized intermediate table;
[0114] The aggregation node is used to aggregate the organization data belonging to the same organization in the aggregated intermediate table to obtain a target table.
[0115] The institutional data aggregation system provided in this embodiment processes an institutional aggregation parameter table designed as a simulated broadcast table based on an institutional relationship tree. Based on the institutional aggregation parameter table, the institutional data of the grassroots institutions with the largest amount of data in the source table to be aggregated are aggregated one level upward to obtain an intermediate table. The institutional data of each institution in the intermediate table is then aggregated in one step to each superior institution of each institution to obtain a summarized intermediate table. The institutional data belonging to the same institution in the summarized intermediate table is then aggregated to obtain a target table. This solution combines step-by-step aggregation and one-step aggregation by simulating a broadcast table, and then utilizes the intermediate table to increase concurrency, reduce cross-shard queries, and avoid cross-shard writes. This solution can fully utilize distributed resources without generating distributed transactions, thereby improving the efficiency of institutional data aggregation in a distributed mode.
[0116] Furthermore, the mechanism summary parameter table includes:
[0117] The grassroots organization row is used to record the organization code and shard key of the grassroots organization, the organization code and shard key of the grassroots organization's parent organization, and the corresponding aggregation level of the parent organization;
[0118] Other institution rows are used to record the institution code and sharding key of any other institution except the grassroots institution, the institution code and sharding key of each superior institution of any other institution, and the summary level corresponding to each superior institution of any other institution.
[0119] Furthermore, a row in the source table to be aggregated corresponds to the organization code, shard key, and organization data of an organization. According to the organization aggregation parameter table, the organization data of the grassroots organizations in the source table to be aggregated are aggregated one level up to obtain an intermediate table, including:
[0120] According to the institution summary parameter table, determine the institution codes corresponding to the grassroots institution and the institution at the next higher level of the grassroots institution;
[0121] Based on the organization codes corresponding to the grassroots organization and its parent organization, delete the first deleted row corresponding to the grassroots organization in the source table to be aggregated, and add the first added row corresponding to the parent organization. The first added row includes the parent organization code, the parent organization's shard key, and the organization data aggregated from the grassroots organization's organization data to the parent organization, to obtain a candidate table.
[0122] In each row of the candidate table, a shard key of the data source organization is added to obtain an intermediate table, where the data source organization is the data source of the organization data in this row.
[0123] Furthermore, according to the institution summary parameter table, the intermediate table is newly added with the institution data of each institution in the intermediate table being aggregated to the data of the corresponding superior institutions in one step, thereby obtaining an aggregated intermediate table including:
[0124] Identify each institution in the intermediate table as a target institution, and determine the institution code and shard key of each parent institution of the target institution based on the institution summary parameter table;
[0125] A second new row is added to the intermediate table, the second new row including the organization code and shard key of a superior organization of the target organization, the shard key of the target organization and the organization data summarized by the organization data of the target organization, to obtain a summarized intermediate table.
[0126] Furthermore, the organization data belonging to the same organization in the aggregated intermediate table are aggregated to obtain a target table, including:
[0127] The rows in the aggregated intermediate table are merged according to the principle of consistent institution codes, and the institution data of the rows involved in the merging are aggregated to obtain the target table.
[0128] Furthermore, the first generation node, the first processing node, the second generation node, the second processing node and the aggregation node all support concurrent processing.
[0129] Furthermore, the organization summary parameter table, the intermediate table, the summarized intermediate table and the target table also include one or more of the summary date and the organization summary scenario number.
[0130] The institutional data aggregation system provided by the embodiment of the present invention can execute an institutional data aggregation method in a distributed mode provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.
[0131] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed, implements a method for aggregating institutional data in a distributed mode provided by any embodiment of the present invention.
[0132] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0133] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0134] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for aggregating institutional data in a distributed mode, characterized in that: Applied to an institutional data aggregation system, the system includes at least a processing node, a first generation node, a first processing node, a second generation node, a second processing node, and an aggregation node, and the method includes: By processing the nodes, an organization summary parameter table designed as a simulated broadcast table is obtained based on the organization relationship tree. The organization summary parameter table includes the organization code and shard key corresponding to each organization and its parent organization; A first task is generated by a first generation node and processed by a first processing node. The first task is used to aggregate the organizational data of the grassroots organizations in the source table to be aggregated one level up to obtain an intermediate table based on the organizational aggregation parameter table. The source table to be aggregated includes the organizational code, shard key, and organizational data of each organization. A second task is generated by a second generation node and processed by a second processing node; the second task is used to add, to the intermediate table, the institutional data of each institution in the intermediate table to the data of the corresponding superior institutions in one step, based on the institutional summary parameter table, to obtain an aggregated intermediate table; The organization data belonging to the same organization in the aggregated intermediate table are aggregated through the aggregation node to obtain a target table.
2. The method according to claim 1, characterized in that The summary parameter table of the organization includes: The grassroots organization row is used to record the organization code and shard key of the grassroots organization, the organization code and shard key of the grassroots organization's parent organization, and the corresponding aggregation level of the parent organization; Other institution rows are used to record the institution code and sharding key of any other institution except the grassroots institution, the institution code and sharding key of each superior institution of any other institution, and the summary level corresponding to each superior institution of any other institution.
3. The method according to claim 1, characterized in that A row in the source table to be aggregated corresponds to the organization code, shard key, and organization data of an organization. According to the organization aggregation parameter table, the organization data of the grassroots organizations in the source table to be aggregated are aggregated one level up to obtain an intermediate table, including: According to the institution summary parameter table, determine the institution codes corresponding to the grassroots institution and the institution at the next higher level of the grassroots institution; Based on the organization codes corresponding to the grassroots organization and its parent organization, delete the first deleted row corresponding to the grassroots organization in the source table to be aggregated, and add the first added row corresponding to the parent organization. The first added row includes the parent organization code, the parent organization's shard key, and the organization data aggregated from the grassroots organization's organization data to the parent organization, to obtain a candidate table. In each row of the candidate table, a shard key of the data source organization is added to obtain an intermediate table, where the data source organization is the data source of the organization data in this row.
4. The method according to claim 1, wherein According to the organization summary parameter table, the organization data of each organization in the intermediate table is added to the intermediate table in one step to the data of each corresponding superior organization, and the summarized intermediate table is obtained, including: Identify each institution in the intermediate table as a target institution, and determine the institution code and shard key of each parent institution of the target institution based on the institution summary parameter table; A second new row is added to the intermediate table, the second new row including the organization code and shard key of a superior organization of the target organization, the shard key of the target organization and the organization data summarized by the organization data of the target organization, to obtain a summarized intermediate table.
5. The method according to claim 1, wherein Aggregate the institutional data belonging to the same institution in the aggregated intermediate table to obtain a target table, including: The rows in the aggregated intermediate table are merged according to the principle of consistent institution codes, and the institution data of the rows involved in the merging are aggregated to obtain the target table.
6. The method according to claim 1, characterized in that The first generation node, the first processing node, the second generation node, the second processing node, and the aggregation node all support concurrent processing.
7. The method according to claim 1, characterized in that The organization summary parameter table, the intermediate table, the summarized intermediate table and the target table further include one or more of a summary date and an organization summary scenario number.
8. An institutional data aggregation system, characterized in that: including a processing node, a first generation node, a first processing node, a second generation node, a second processing node, and a summary node; A processing node is used to process the organization relationship tree to obtain an organization summary parameter table designed as a simulated broadcast table. The organization summary parameter table includes the organization code and shard key corresponding to each organization and its parent organization; The first generating node is used to generate a first task, and the first processing node is used to process the first task; The first task is to aggregate the organizational data of the grassroots organizations in the source table to be aggregated one level up to obtain an intermediate table based on the organizational aggregation parameter table. The source table to be aggregated includes the organizational code, shard key, and organizational data of each organization. The second generating node is used to generate a second task, and the second processing node is used to process the second task; The second task is to add, to the intermediate table, the organization data of each organization in the intermediate table to the data of the corresponding superior organizations in one step according to the organization summary parameter table, to obtain a summarized intermediate table; The aggregation node is used to aggregate the organization data belonging to the same organization in the aggregated intermediate table to obtain a target table.
9. The system according to claim 8, characterized in that The summary parameter table of the organization includes: The grassroots organization row is used to record the organization code and shard key of the grassroots organization, the organization code and shard key of the grassroots organization's parent organization, and the corresponding aggregation level of the parent organization; Other institution rows are used to record the institution code and sharding key of any other institution except the grassroots institution, the institution code and sharding key of each superior institution of any other institution, and the summary level corresponding to each superior institution of any other institution.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.