Data aggregation method and device, computer equipment and readable storage medium

By dynamically adjusting thread resource allocation during data aggregation process, and determining processing priority based on the length of category elements and the length of reference elements, the problem of uneven thread resource allocation is solved and the efficiency and balance of data aggregation are improved.

CN119939500APending Publication Date: 2025-05-06HANGZHOU PINGPONG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411938339.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the data compliance governance scenario, the existing technology has the problem of unbalanced thread resource allocation, which leads to inefficiency in the data aggregation process.

Method used

By obtaining the first material template for data aggregation, traverse the category set, determine the length of reference elements and the total number of thread executions for each level of the category, determine the thread resource requirements based on the relationship between the length of the category element and the length of the reference element, and finally, the thread resource allocation is performed on the category element based on the processing priority to achieve more balanced task allocation.

Benefits of technology

By dynamically adjusting thread resource allocation, the problem of uneven thread resource allocation is solved, and the efficiency and balance of the data aggregation process are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939500A_ABST
    Figure CN119939500A_ABST
Patent Text Reader

Abstract

The invention relates to a data aggregation method and device, computer equipment and a readable storage medium. The method comprises the following steps: in response to a data aggregation processing request, obtaining a first material template for data aggregation; traversing the category set to obtain a category element set corresponding to each level of category; according to the length of each category element in the category element set and the category number of the category element set, determining the length of a reference element of each level of category, and according to the category number, determining the total thread execution number for processing each level of category; according to the relation between the length of each type of eye element and the length of the reference element, determining thread resource requirements for processing each type of eye element; according to the thread execution total number and the thread resource demand, determining the processing priority of each type of target element, performing thread resource allocation on each type of target element based on the processing priority, and executing the allocated thread resources to complete aggregation processing on the type element set. By adopting the method, the problem of uneven thread resource distribution in the data aggregation process can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, apparatus, computer device and readable storage medium for data aggregation. Background Art

[0002] In the data compliance governance scenario, in order to meet different business needs, it is necessary to extract features from massive data in multiple business scenarios and perform data aggregation on the extracted data features. However, in actual business scenarios, there will be redundant data, lack of uniformity in data format and structure, and complex multi-level business data in the business data. To address this situation, it is necessary to aggregate and process a large amount of data.

[0003] However, in scenarios where large amounts of data are traversed and aggregated, related technologies suffer from the defect of unbalanced thread resource allocation. Summary of the invention

[0004] Based on this, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium and computer program product for data aggregation that can solve the uneven distribution of thread resources during data aggregation in response to the above technical problems.

[0005] In a first aspect, the present application provides a method for data aggregation, comprising:

[0006] In response to a data aggregation processing request, obtaining a first material template for data aggregation; wherein the categories in the category set of the first material template have a hierarchical relationship;

[0007] Traversing the category set, for each level of categories, obtaining a category element set corresponding to each level of categories;

[0008] Determine the reference element length of each level of categories according to the length of each category element in the category element set and the number of categories in the category element set, and determine the total number of threads executed to process each level of categories according to the number of categories;

[0009] Determining thread resource requirements for processing each of the category elements according to a relationship between the length of each of the category elements and the length of the reference element;

[0010] According to the total number of thread executions and the thread resource requirements, the processing priority of each category element is determined, thread resources are allocated to each category element based on the processing priority, and aggregation processing of the category element set is completed by executing the allocated thread resources.

[0011] In one embodiment, the reference element length is an average element length of the category element set, and determining the thread resource requirement for processing each of the category elements according to the relationship between the length of each of the category elements and the reference element length includes:

[0012] If there is a first category element in the category element set whose length is greater than the reference element length, determining the thread resource requirement of the first category element as the first thread resource requirement;

[0013] If there is a second category element in the category element set whose length is less than the reference element length, determining the thread resource requirement of the second category element as the second thread resource requirement;

[0014] The number of thread resources occupied by the second thread resource requirement is smaller than the number of thread resources occupied by the first thread resource requirement.

[0015] In one embodiment, the category element set includes the first category element and the second category element, and determining the processing priority of each of the category elements according to the total number of thread executions and the thread resource requirements includes:

[0016] Dividing the category element set according to the reference element length to obtain subcategory element sets, and determining the proportion of the total number of thread executions according to the number of elements in the subcategory element sets;

[0017] If the number of the first category elements is greater than or equal to the product of the proportion value and the total number of thread executions, it is determined that the processing priority of the first category elements is lower than the processing priority of the second category elements;

[0018] If the number of the first category elements is less than the product of the proportion value and the total number of thread executions, it is determined that the processing priority of the first category elements is greater than the processing priority of the second category elements.

[0019] In one embodiment, the allocating thread resources to each of the category elements based on the processing priority, and completing the aggregation processing of the category element set by executing the allocated thread resources, includes:

[0020] If the processing priority of the first category element is lower than the processing priority of the second category element, allocating the available thread resources corresponding to the total number of thread executions to the second category element;

[0021] If the processing priority of the first category element is greater than the processing priority of the second category element, the available thread resources corresponding to the total number of thread executions are allocated to the first category element.

[0022] In one of the embodiments, dividing the category element set according to the reference element length to obtain a subcategory element set, and determining the proportion of the total number of thread executions according to the number of elements in the subcategory element set, includes:

[0023] Dividing the category element set according to the reference element length to obtain a first category element set and a second category element set; wherein the element length of each first category element in the first category element set is greater than the reference element length, and the element length of each second category element in the second category element set is less than the reference element length;

[0024] The proportion of the total number of thread executions is determined according to the number of elements in the first category element set and the number of elements in the second category element set.

[0025] In one embodiment, the method further comprises:

[0026] Acquire all second material templates associated with the business line, and template source data of each of the second material templates;

[0027] Traversing the template source data of each second material template, performing feature extraction on the template source data, and identifying all categories in each second material template for representing the same attribute according to the extracted features;

[0028] An aggregate category corresponding to all the categories representing the same attribute is generated, and the aggregate category is used as the category of the first material template. There is a mapping relationship between the category of the first material template and the categories of all the second material templates.

[0029] In one embodiment, the step of completing the aggregation processing of the category element set by executing the allocated thread resources includes:

[0030] By executing the allocated thread resources, according to the mapping relationship between the category of the first material template and the categories of all the second material templates, the source data corresponding to the category element set is acquired from the template source data of each of the second material templates;

[0031] Aggregation processing is performed on the source data corresponding to the category element set.

[0032] In one embodiment, the method further comprises:

[0033] For each level of categories, generating a category code for each level of categories;

[0034] The category code is determined as a thread identifier of a thread corresponding to the category element, and the thread identifier is used to query and locate thread processing progress.

[0035] In a second aspect, the present application also provides a data aggregation device, including:

[0036] A data acquisition module, configured to acquire a first material template for data aggregation in response to a data aggregation processing request; wherein the categories in the category set of the first material template have a hierarchical relationship;

[0037] A traversal module, used to traverse the category set, and for each level of categories, obtain a category element set corresponding to each level of categories;

[0038] A data processing module, configured to determine the reference element length of each level of categories according to the length of each category element in the category element set and the number of categories in the category element set, and determine the total number of threads executed to process each level of categories according to the number of categories;

[0039] Determining thread resource requirements for processing each of the category elements according to a relationship between the length of each of the category elements and the length of the reference element;

[0040] A resource allocation module is used to determine the processing priority of each of the category elements according to the total number of thread executions and the thread resource requirements, allocate thread resources to each of the category elements based on the processing priority, and complete the aggregation processing of the category element set by executing the allocated thread resources.

[0041] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0042] In response to a data aggregation processing request, obtaining a first material template for data aggregation; wherein the categories in the category set of the first material template have a hierarchical relationship;

[0043] Traversing the category set, for each level of categories, obtaining a category element set corresponding to each level of categories;

[0044] Determine the reference element length of each level of categories according to the length of each category element in the category element set and the number of categories in the category element set, and determine the total number of threads executed to process each level of categories according to the number of categories;

[0045] Determining thread resource requirements for processing each of the category elements according to a relationship between the length of each of the category elements and the length of the reference element;

[0046] According to the total number of thread executions and the thread resource requirements, the processing priority of each category element is determined, thread resources are allocated to each category element based on the processing priority, and aggregation processing of the category element set is completed by executing the allocated thread resources.

[0047] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0048] In response to a data aggregation processing request, obtaining a first material template for data aggregation; wherein the categories in the category set of the first material template have a hierarchical relationship;

[0049] Traversing the category set, for each level of categories, obtaining a category element set corresponding to each level of categories;

[0050] Determine the reference element length of each level of categories according to the length of each category element in the category element set and the number of categories in the category element set, and determine the total number of threads executed to process each level of categories according to the number of categories;

[0051] Determining thread resource requirements for processing each of the category elements according to a relationship between the length of each of the category elements and the length of the reference element;

[0052] According to the total number of thread executions and the thread resource requirements, the processing priority of each category element is determined, thread resources are allocated to each category element based on the processing priority, and aggregation processing of the category element set is completed by executing the allocated thread resources.

[0053] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0054] In response to a data aggregation processing request, obtaining a first material template for data aggregation; wherein the categories in the category set of the first material template have a hierarchical relationship;

[0055] Traversing the category set, for each level of categories, obtaining a category element set corresponding to each level of categories;

[0056] Determine the reference element length of each level of categories according to the length of each category element in the category element set and the number of categories in the category element set, and determine the total number of threads executed to process each level of categories according to the number of categories;

[0057] Determining thread resource requirements for processing each of the category elements according to a relationship between the length of each of the category elements and the length of the reference element;

[0058] According to the total number of thread executions and the thread resource requirements, the processing priority of each category element is determined, thread resources are allocated to each category element based on the processing priority, and aggregation processing of the category element set is completed by executing the allocated thread resources.

[0059] The above-mentioned data aggregation method, apparatus, computer equipment, computer-readable storage medium and computer program product, in the scenario of traversing and aggregating a large amount of data, for each level of categories in the first material template, by determining to obtain the category element set corresponding to each level of categories, according to the length of each category element in the category element set and the number of categories in the category element set, determine the reference element length of each level of categories and the total number of thread executions for processing each level of categories; for a level of categories, based on the known thread resources of each level, according to the relationship between the length of each category element and the reference element length, determine the thread resource requirements for processing each category element compared to the reference element length, and determine the thread resource requirements of each category element; on this basis, based on the total number of thread executions and the thread resource requirements of each category element, by determining the processing priority of each category element, the granularity of tasks can be adjusted to adapt to the thread resource requirements and processing performance of different category elements, thereby achieving more balanced task allocation and solving the problem of uneven thread resource allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0061] Figure 1 A diagram of an application environment of a method for data aggregation in one embodiment;

[0062] Figure 2 A schematic flow chart of a method for data aggregation in one embodiment;

[0063] Figure 3 is a schematic flow chart of step 208 in one embodiment;

[0064] Figure 4 A flowchart of a method for determining a percentage of the total number of thread executions in one embodiment;

[0065] Figure 5 A schematic diagram of the relationship between category codes and thread identifiers in one embodiment;

[0066] Figure 6A schematic diagram of thread resource allocation in one embodiment;

[0067] Figure 7 A flowchart of thread scheduling in one embodiment;

[0068] Figure 8 A timing diagram of a method for data aggregation in one embodiment;

[0069] Fig. 9 A structural block diagram of a device for data aggregation in one embodiment;

[0070] Fig.10 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0071] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0072] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0073] In the data compliance governance scenario, there are a lot of data that need to be feature extracted and aggregated. For example, when users conduct different businesses, there will be scenarios where they repeatedly register or collect user information, resulting in repeated data storage, redundant and inconsistent user data; the lack of uniformity in data format and structure increases the complexity of data parsing and processing; there is a lack of processing methods for complex multi-level aggregation management, and at the same time, due to the diversity of the data itself, the cost of data aggregation is high. In the process of aggregating data, there is no complete set of feature aggregation strategies and data aggregation generation technologies and methods, such as the aggregation of material templates, and users re-aggregating their material elements through new template definitions. Then in the scenario of traversing and aggregating a large amount of data, the amount of data processing is large, and there is a problem of uneven thread resource allocation.

[0074] To address this technical problem, a data aggregation method is proposed, which can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal responds to the data aggregation processing request and obtains the first material template for data aggregation; wherein the categories in the category set of the first material template have a hierarchical relationship; the category set is traversed, and for each level of categories, the category element set corresponding to each level of category is obtained; according to the length of each category element in the category element set and the number of categories in the category element set, the reference element length of each level of category is determined, and the total number of thread executions for processing each level of category is determined according to the number of categories; according to the relationship between the length of each category element and the reference element length, the thread resource requirements for processing each category element are determined; according to the total number of thread executions and the thread resource requirements, the processing priority of each category element is determined, and thread resources are allocated to each category element based on the processing priority, and the aggregation processing of the category element set is completed by executing the allocated thread resources.

[0075] The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, portable wearable devices, etc. The server 104 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0076] In an exemplary embodiment, Figure 2 As shown, a method for data aggregation is provided, which is applied to Figure 1 The terminal in is taken as an example to illustrate, including the following steps 202 to 206. Among them:

[0077] Step 202, in response to a data aggregation processing request, obtain a first material template for data aggregation; wherein the categories in the category set of the first material template have a hierarchical relationship.

[0078] Among them, the categories in the category set of the first material template have a hierarchical relationship. For example, the category set includes first-level to N-level categories. Correspondingly, the category elements corresponding to the N-level category include category element 1, category element 2, category element 3, ..., category element M. The number of category element sets for each level of category can be the same or different. If there is a next level under the category element of the first-level category, the category element is taken as the second-level category, and the category element set corresponding to the second-level category is determined. For example, the first-level categories include category element 1, category element 2, category element 3, ..., category element M, and each category element has a corresponding element length. Among them, N is a positive integer, and N is greater than or equal to 2.

[0079] The first material template includes categories associated with business lines, and the categories include business information of different dimensions. Business information can include customer materials. Customer materials can be divided according to feature extraction, and can be divided into categories including enterprise information module, legal person module, etc. There are subcategories under the enterprise module, which can also be called category elements, including enterprise name, enterprise unified credit code, etc. According to the above logic, the enterprise information module is used as a first-level series category, and the following materials in its module are used as second-level elements (when there are third-level elements, they are defined as second-level categories at this moment). By traversing different categories and elements, the material elements corresponding to the original different business lines are associated. Categories and category element sets belonging to categories can be fields.

[0080] For each category in the first material template, there can be at least one semantic expression for the same category semantics of each category. It can be understood that there is a corresponding mapping relationship between the categories in the first material template and the categories expressed for the same category semantics in all business scenarios. For example, the category semantics of the category in the first material template is the enterprise information module, which is enterprise information in business A and enterprise module in business B. However, the semantics of enterprise information and enterprise module are the same and are used to represent the enterprise name. Therefore, there is a mapping between the enterprise information module and the enterprise information and the enterprise module, and the enterprise information module includes two descriptions: enterprise information and enterprise module.

[0081] The first material template can be determined by obtaining the source data of the material history template, performing feature extraction on the source data of the material history template, identifying the categories with the same semantics in the material history template, and determining a new category for the categories with the same semantics. The new category covers the categories with the same semantics in the material history template. Based on this, the data module category is redefined, and a mapping relationship exists between the determined first material template and the material history template. It should be noted that the first material template has the same dimension as the data module category in the material history template, but the first material template can cover the data module categories of all material history templates.

[0082] Step 204, traverse the category set, and for each level of categories, obtain a category element set corresponding to each level of categories.

[0083] It is understandable that data aggregation is performed based on the first material template, and the aggregation is implemented according to the category level. It is necessary to wait until the aggregation of the categories at the same level is completed before performing the aggregation of the categories at the next level. The principle of allocating aggregation threads for each level of categories is the same. For the sake of ease of description, the aggregation of single-level categories is used to expand the description. Among them, for each level of categories, the category element set corresponding to each level of categories is obtained. The method of obtaining the category element set can be determined by existing data analysis methods and is not limited here.

[0084] Exemplarily, the category set is traversed, and the aggregation tasks corresponding to each level of categories are executed in the order of time series nodes to obtain the category element set corresponding to each level of categories.

[0085] Step 206, determining the reference element length of each level of categories according to the length of each category element in the category element set and the number of categories in the category element set, and determining the total number of threads executed to process each level of categories according to the number of categories.

[0086] The reference element length of each level category may refer to the average element length of each level category. The total number of thread executions may be the actual number of threads required to process each level category. However, in actual applications, the thread resources allocated by the thread pool are predetermined.

[0087] Exemplarily, thread pool parameters are defined, and length extraction is performed on each level of categories so that the column length is equal to the total number of thread executions, that is, the total number of thread executions for processing each level of categories is determined according to the number of categories. The length of each category element in the category element set and the number of categories in the category element set are obtained, and the sum of the lengths of each category element is divided by the number of categories to obtain the average element length of each level of categories, and the average element length is determined as the reference element length of each level of categories.

[0088] Furthermore, the category code obtained by the category can be assigned to the thread name to clearly mark the thread resources to which the category belongs, so as to facilitate the query of logic details and abnormal situations.

[0089] Step 208, determining the thread resource requirements for processing each category element based on the relationship between the length of each category element and the length of the reference element.

[0090] When each category element is executed, corresponding thread resources are allocated, and the thread resources required for each category element are associated with the length of the category element. The thread resources required here take into account the processing time of processing the category element based on a given thread.

[0091] Exemplarily, initial resource allocation is performed based on the number of CPU cores of the physical machine, that is, the total number of threads executing to process each level of categories is determined according to the number of categories, and the reference element length of each level of categories is used as the efficiency weight value. The efficiency weight value can be used to characterize that when the length of a category or element exceeds the reference element length, it means that the more resources its execution process occupies. On the contrary, when it is less than the reference element length, the higher the efficiency is and the fewer resources are occupied.

[0092] Step 210, determining the processing priority of each category element according to the total number of thread executions and thread resource requirements, allocating thread resources to each category element based on the processing priority, and completing the aggregation processing of the category element set by executing the allocated thread resources.

[0093] Among them, the purpose of determining the processing priority of each category element according to the total number of thread executions and thread resource requirements is to avoid tasks with large thread resource requirements from being processed, so that the terminal's physical memory maintains a relatively good utilization rate while dynamically adjusting the task granularity to balance resources.

[0094] The processing priority of each category element can be determined by determining the proportion of the total number of thread executions. If the number of category elements in the category element set whose length is greater than the reference element length is greater than the proportion, the task submission of the category element is suspended (i.e., the submission of new high-weight tasks is suspended), and the available thread resources are allocated to the category elements in the category element set whose length is less than the reference element length (i.e., the submission of low-weight tasks). The proportion of the total number of thread executions can be used to ensure that the physical memory reaches the preset utilization rate while dynamically adjusting the task granularity to balance resources.

[0095] Aggregation processing is to perform aggregation calculation on the material history template source data corresponding to the category element set, and store the aggregated data in the database according to the category of the first material template. Each level of category can store data after the aggregation processing is completed.

[0096] The above-mentioned data aggregation method, in the scenario of traversing and aggregating a large amount of data, for each level of categories in the first material template, by determining and obtaining the category element set corresponding to each level of categories, according to the length of each category element in the category element set and the number of categories in the category element set, determines the reference element length of each level of categories and the total number of thread executions for processing each level of categories; for categories of one level, based on the known thread resources of each level, according to the relationship between the length of each category element and the reference element length, determines the thread resource requirements for processing each category element compared to the reference element length, and determines the thread resource requirements of each category element; on this basis, based on the total number of thread executions and the thread resource requirements of each category element, by determining the processing priority of each category element, the granularity of tasks can be adjusted to adapt to the thread resource requirements and processing performance of different category elements, thereby achieving more balanced task allocation and solving the problem of uneven thread resource allocation.

[0097] In view of the inconsistency of data formats and structures in data aggregation, which increases the complexity of data analysis and processing, a method is proposed to collect data features to specify data specifications, establish material templates to unify data model analysis, data format, and coding mapping. Data is stored according to template specifications to achieve data standardization.

[0098] In an exemplary embodiment, all second material templates associated with the business line and the template source data of each second material template are obtained; the template source data of each second material template is traversed, and features are extracted from the template source data, and all categories used to characterize the same attribute in each second material template are identified based on the extracted features; aggregate categories corresponding to all categories characterizing the same attribute are generated, and the aggregate categories are used as categories of the first material template. There is a mapping relationship between the categories of the first material template and the categories of all second material templates.

[0099] Optionally, in response to the material collection process existing in multiple businesses, the material template source data of the second material template used by each existing business line is extracted, each element in the collected business template is traversed, and the business template is used to collect the source data from the material template by extracting the features of the second material template. The data collected by the business template is displayed through the second material template, and the same modules and their subordinate element materials in all the second material templates are mapped and defined to construct a new material template, namely the first material template.

[0100] Constructing a brand new material template can be specifically done by traversing the business template data separately, defining the modules of the business template as first-level categories, and defining its specific elements as second-level categories. If there is still a hierarchical relationship under the elements, the category levels will be further traversed and divided, and the final traversal results will be object mapped to obtain an independent business template category object (i.e., the material template). Create a module pool, aggregate the categories of the same level, arrange the module information of the statistical categories in disorder, and cover the data of the modules with the same name to obtain a brand new and focused first-level category pool (module pool). Give it a new module code and store it in the database. By traversing the first-level category columns obtained, take the element sets (second-level categories or N-level categories) from the business template category objects respectively. Relative to the module set, the complexity of the element set shows an increase of n2 relative to the module set.

[0101] In an exemplary embodiment, Figure 3 As shown, step 208 includes steps 302 to 304. Among them:

[0102] Step 302: If the length of a first category element in the category element set is greater than the reference element length, the thread resource requirement of the first category element is determined to be the first thread resource requirement.

[0103] The task corresponding to the resource demand of the first thread may be a high-weight task, which indicates that the more resources the execution process occupies, the lower the efficiency.

[0104] Step 304: If there is a second category element in the category element set whose length is smaller than the reference element length, the thread resource requirement of the second category element is determined to be the second thread resource requirement.

[0105] The task corresponding to the second thread resource requirement may be a low-weight task, indicating that the higher the efficiency of its execution process, the less resources it occupies. The number of thread resources occupied by the second thread resource requirement is less than the number of thread resources occupied by the first thread resource requirement.

[0106] In this embodiment, the demand for thread resources of each category element in the category element set can be determined according to the reference element length, and the thread resources can be reasonably arranged based on the demand for thread resources to improve the utilization rate of thread resources.

[0107] In an exemplary embodiment, according to the total number of thread executions and thread resource requirements, determining the processing priority of each category element includes: dividing the category element set according to the reference element length to obtain the subcategory element set, and determining the proportion of the total number of thread executions according to the number of elements in the subcategory element set;

[0108] If the number of elements in the first category is greater than or equal to the product of the proportion value and the total number of thread executions, it is determined that the processing priority of the elements in the first category is lower than the processing priority of the elements in the second category; accordingly, if the processing priority of the elements in the first category is lower than the processing priority of the elements in the second category, the available thread resources corresponding to the total number of thread executions are allocated to the elements in the second category, and the tasks of the elements in the first category are processed when the thread resources are sufficient.

[0109] If the number of elements in the first category is less than the product of the proportion value and the total number of thread executions, it is determined that the processing priority of the elements in the first category is greater than the processing priority of the elements in the second category. Correspondingly, if the processing priority of the elements in the first category is greater than the processing priority of the elements in the second category, the available thread resources corresponding to the total number of thread executions are allocated to the elements in the first category.

[0110] Exemplarily, a scheduling object is created, and the thread execution is monitored through the scheduling object, and the number of high-weight threads currently executing (i.e., the number of first-category elements) is calculated in real time. If the ratio of the number of first-category elements to the number of second-category elements is greater than the ratio, new high-weight task submissions are suspended, and resources are allocated to low-weight task submissions. If the number of first-category elements is less than the ratio, it indicates that thread resources are sufficient, and resources can be allocated to high-weight task submissions.

[0111] In the above method, the category element set is divided according to the reference element length, and the obtained sub-category element set is determined according to the number of elements in the sub-category element set. The proportion of the total number of thread executions is determined, and the task granularity is dynamically adjusted based on the proportion and the current thread resource allocation.

[0112] In an exemplary embodiment, Figure 4As shown, a method for determining a percentage of the total number of thread executions is provided, including steps 402 to 404, wherein:

[0113] Step 402, dividing the category element set according to the reference element length to obtain a first category element set and a second category element set; wherein the element length of each first category element in the first category element set is greater than the reference element length, and the element length of each second category element in the second category element set is less than the reference element length.

[0114] Step 404: Determine a percentage of the total number of thread executions according to the number of elements in the first category element set and the number of elements in the second category element set.

[0115] The percentage of the total number of thread executions is a value obtained by dividing the number of elements in the first category element set by the number of elements in the second category element set, and then multiplying it by 1 / 2.

[0116] Alternatively, according to Amdahl's law and Pareto's principle, the higher the degree of parallelism, the lower the performance. In the resource allocation rule, allocating 50% of the resources to high-priority tasks is usually sufficient to handle most of the critical load. Combine the above resources to calculate the scheduling weight value:

[0117] Given the element lengths E1, E2, E3, ...En of the category element sets of each level category, calculate the average length: Avg = (E1+E2+E3+......+En) / n; divide the category element set into two sets: A = {Ei|Ei>Avg}, A is the set of elements exceeding the average length; B = {Ei|Ei≤Avg}, B is the set of elements less than or equal to the average length; the proportion of the total number of thread executions S = |A| / |B|*50%.

[0118] That is to say, the monitoring logic is added to the threads so that the task threads with higher weight values ​​do not exceed half of the total number of threads being executed during the execution process.

[0119] In the above method for determining the proportion of the total number of thread executions, determining the proportion in this way can maintain a relatively good utilization rate of the physical memory while dynamically adjusting the task granularity to balance resources.

[0120] In order to clearly mark the thread resources to which the categories belong so as to facilitate the query of logic details and abnormal situations, optionally, in an exemplary embodiment, for each level of categories, a category code for each level of categories is generated; the category code is determined as the thread identifier of the thread corresponding to the category element, and the thread identifier is used to query and locate the thread processing progress. Figure 5As shown, it is a schematic diagram of the relationship between category code and thread identification in an exemplary embodiment, the category code of the first-level category is code: 001, the code corresponding to element 1 of category 1 is: 001001, the code corresponding to element 2 of category 1 is 001002, and the code corresponding to element 3 of category 1 is: 001003. If a category element has a next-level element, the code of the element is equivalent to the category code. For example, element 3 of category 1 is a second-level category, and there is a next-level element, element 4 of category 2, code: 001003001, then the code of element 3 of category 1 can be used as the category code.

[0121] In actual application scenarios, the thread processing progress can be directly queried and located according to the thread ID, which improves traceability and debugging efficiency.

[0122] Based on the above embodiments, in an exemplary embodiment, Figure 6 As shown, a thread resource allocation diagram is provided. The thread pool includes multiple threads, thread 1, thread 2, ..., thread N. For multiple categories at the same level, the node sequence includes node 1, node 2, ..., node N. Each node has an identification code and the corresponding average length of the subset. The node can be regarded as a category, and the average length of the subset can be regarded as the average element length of each level category.

[0123] The thread pool schedules and monitors the category aggregation (the number of threads is fixed at 4, the length threshold is 8, and the scheduling task weight is controlled to not exceed 1 / 2 of the total threads in execution). If it is detected that two threads are already executing node tasks exceeding the threshold of 8, then the available threads can only select node tasks with an element length less than 8 for execution. Furthermore, if there are multiple node tasks with an element length less than 8 to be executed, the node tasks with an element length less than 8 can be processed based on the pre-set priority and available thread resources. The specific thread scheduling flowchart is as follows: Figure 7 As shown, for category 1, the corresponding category code is 0001, and the element length is n. By creating a scheduling object, the thread resources are scheduled and monitored. Given a total number of threads of 4, the reference element length of the category is 8, and the proportion of the total number of thread executions is 50%. For each category element in category 1, it is determined whether the length of the category element exceeds 8. If it exceeds 8 and the current number of high-weight tasks is less than half of the total number of threads in execution (i.e., 2), the task of the category element is submitted to the thread pool for thread resource allocation. If it is equal to half of the total number of thread executions, the task is suspended and waits for reentry. If the length of the category element does not exceed 8 and the execution length is less than or equal to 8, the task corresponding to the category element is submitted to the thread pool for thread allocation.

[0124] The thread scheduling process corresponding to the implementation of dynamic thread pool scheduling may include: initializing a thread pool, whose parameters are:

[0125] corePoolSize: the number of core threads, which can be set to the number of CPU cores; maximumPoolSize: the maximum number of threads, which can be set to twice the corePoolSize; keepAliveTime: thread idle time, 60 seconds is recommended; workQueue: task queue, LinkedBlockingQueue can be used; threadFactory: thread factory, used to create and name threads; rejectedExecutionHandler: rejection policy, CallerRunsPolicy can be used.

[0126] Declare a scheduling object, whose specific parameters include:

[0127] currentCategoryCode: the name of the currently executing task (category code); weightProportionRatio: weight ratio (S); currentCategoryLength: the total number of currently executing tasks (the total number of threads executed P); aboveWeightThreadArray: an array of threads whose execution length exceeds H; belowWeightThreadArray: an array of threads whose execution length is less than or equal to H; completedTaskCount: the number of executed tasks; exceptionHander: exception handler; errTask: failed task queue; retryCount: the number of retries.

[0128] Construct the scheduling method:

[0129] init(): initialize the scheduling object and set the initial parameters; submitTask(Task task): submit the task to the thread pool; monitorThreads(): monitor thread execution; calculateWeightRatio(): calculate the current weight ratio; balanceTaskExecution(): balance task execution; handleException(Exception e): handle exceptions; retryFailedTasks(): retry failed tasks.

[0130] In an exemplary embodiment, aggregation processing of a category element set is completed by executing allocated thread resources, including: by executing allocated thread resources, according to the mapping relationship between the category of the first material template and the categories of all second material templates, source data corresponding to the category element set is obtained from the template source data of each second material template; and the source data corresponding to the category element set is aggregated.

[0131] Among them, data aggregation processing can refer to the integration, conversion and aggregation of data from different sources or in different formats, and the aggregation operations on data according to actual business needs, such as sum, average, count, maximum, minimum, etc., to extract key information and indicators.

[0132] Optionally, in an exemplary embodiment, a method for data aggregation is provided, wherein the method is applied to Figure 1 The following is an example of a terminal in the example, including:

[0133] Acquire all second material templates associated with the business line and the template source data of each second material template; traverse the template source data of each second material template, extract features from the template source data, and identify all categories in each second material template that are used to characterize the same attribute based on the extracted features; generate aggregate categories corresponding to all categories that characterize the same attribute, and use the aggregate categories as the categories of the first material template. There is a mapping relationship between the categories of the first material template and the categories of all second material templates.

[0134] In response to a data aggregation processing request, a first material template for data aggregation is obtained, the category set is traversed, and for each level of categories, a category element set corresponding to each level of categories is obtained; based on the length of each category element in the category element set and the number of categories in the category element set, a reference element length of each level of categories is determined; based on the number of categories, the total number of threads to be executed to process each level of categories is determined.

[0135] If the length of a first category element in the category element set is greater than the length of a reference element, the thread resource requirement of the first category element is determined to be the first thread resource requirement; if the length of a second category element in the category element set is less than the length of the reference element, the thread resource requirement of the second category element is determined to be the second thread resource requirement.

[0136] The category element set is divided according to the reference element length to obtain a first category element set and a second category element set; wherein the element length of each first category element in the first category element set is greater than the reference element length, and the element length of each second category element in the second category element set is less than the reference element length; according to the number of elements in the first category element set and the number of elements in the second category element set, the proportion of the total number of thread executions is determined. If the number of first category elements is greater than or equal to the product of the proportion and the total number of thread executions, it is determined that the processing priority of the first category elements is less than the processing priority of the second category elements; if the number of first category elements is less than the product of the proportion and the total number of thread executions, it is determined that the processing priority of the first category elements is greater than the processing priority of the second category elements.

[0137] If the processing priority of the first category element is lower than the processing priority of the second category element, the available thread resources corresponding to the total number of thread executions are allocated to the second category element; if the processing priority of the first category element is higher than the processing priority of the second category element, the available thread resources corresponding to the total number of thread executions are allocated to the first category element. By executing the allocated thread resources, according to the mapping relationship between the category of the first material template and the categories of all the second material templates, the source data corresponding to the category element set is obtained from the template source data of each second material template; the source data corresponding to the category element set is aggregated.

[0138] Wherein, for each level of categories, a category code of each level of categories is generated; the category code is determined as a thread identifier of a thread corresponding to a category element, and the thread identifier is used to query and locate thread processing progress.

[0139] In an exemplary embodiment, Figure 8 As shown, a timing diagram of a data aggregation method is provided, including categories, scheduling objects and thread pools. The scheduling object has thread management capabilities and implements granularity adjustment based on the proportion value S. The scheduling object submits executable tasks to the thread pool and monitors the processing progress of the thread pool in real time.

[0140] The following is an example of a business scenario. User xxx is engaged in business A and B at the same time and is user a and b respectively. According to business needs, the material template and user materials are unified to obtain the first material template. The implementation process includes:

[0141] S11: Traverse business templates, build hierarchical category structure, create module pools, aggregate same-level categories, remove duplicates and assign new codes for storage. Details: A business: enterprise module: enterprise name and B business: enterprise information: enterprise name, unified aggregation: N unified material: enterprise information module: enterprise name, association relationship A: code = 123, B: code = abc, N: code = 00010001, select 00010001 format to generate template code, which is easy for computer logic calculation and subsequent expansion, and other material elements are similar.

[0142] S12: Based on the unified user center, the customer material source data is arranged and analyzed in time sequence, the unique identifier is extracted, and the irreversible encryption algorithm is used to generate a unique subject identification KEY.

[0143] S13: Assign a unique subject identification id (such as UUID) to the customer of the first node in the time series, and store the relationship between the user and the id. If the subsequent node customers have the same unique subject identification KEY, they will be assigned the same unique subject identification id.

[0144] S14: Data collection is performed by obtaining materials of two users with the same unique subject identification ID, and materials are aggregated by the material attributes associated with the users: primary categories and secondary elements. Data materials are aggregated based on the above data aggregation method. The same level categories include category elements: enterprise information, legal representative information, ultimate beneficiary information and agent information. The corresponding identifiers of the category elements are 0001, 0002, 0003 and 0004, and the corresponding element lengths are 32, 16, 8 and 8 respectively.

[0145] According to the number of category elements, the total number of thread executions P = 4, the average length of elements H = (32 + 16 + 8 + 8) / 4 = 16; determine the proportion of the total number of thread executions S = (2) / (2) * 50% = 50%; assign tasks to the category elements, when the fixed number of threads in our thread pool is 2. The specific execution process is:

[0146] Task 0001 (length 32) enters and starts: Assign thread 1; Scheduling description: Number of threads assigned to high-weight tasks: 1, number of idle threads: 1;

[0147] When task 0002 (length 16) enters: thread 2 is assigned; scheduling description: high-weight task threads have occupied 50% of resources, no more thread resources are assigned, and task submission is suspended. Number of idle threads: 1;

[0148] Task 0003 (length 8) enters and starts: Assign thread 2; Scheduling description: Number of threads assigned to low-weight tasks: 1, number of idle threads: 0;

[0149] 0001 Task completed, thread 1 assigned; Scheduling description: High-weight task completed, thread releases resources, number of idle threads: 1;

[0150] 0003 Task completed, thread 2 assigned; Scheduling description: low-weight task completed, thread releases resources, number of idle threads: 2;

[0151] 0002 Task re-execution assigned to thread 1; Scheduling description: Number of threads assigned to high-weight tasks: 1, number of idle threads: 1;

[0152] Task 0004 (length 8) enters and starts to allocate thread 2; Scheduling description: The number of threads allocated to the low-weight task: 1, the number of idle threads: 0;

[0153] 0004 Task completed, thread 2 assigned; Scheduling description: low-weight task completed, thread releases resources, number of idle threads: 1;

[0154] 0002 Task completion assigns thread 1; Scheduling description: High-weight tasks are completed, threads release resources, number of idle threads: 2. Based on the above steps, the corresponding aggregation results are obtained.

[0155] It is understandable that the implementation method of the data aggregation method can be implemented in the above-mentioned limited manner, which will not be elaborated here.

[0156] In the above method, by dynamically defining the thread pool parameters, the number of threads is matched with the number of tasks, and the task granularity is predicted according to the number of category elements belonging to each level of the category, a more balanced task distribution is achieved. By assigning business numbers to threads, traceability and debugging efficiency are improved. On this basis, the reentry logic can be executed in multiple dimensions in combination with thread strategies, priorities, and resources, for example: extracting and reentering execution for abnormal thread numbers. In addition, by real-time monitoring of high-weight task threads, the task granularity is dynamically adjusted to ensure balanced resource utilization.

[0157] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0158] Based on the same inventive concept, the embodiment of the present application also provides a device for data aggregation for implementing the data aggregation method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in the one or more data aggregation device embodiments provided below can refer to the limitations of the data aggregation method above, and will not be repeated here.

[0159] In an exemplary embodiment, Fig. 9 As shown, a data aggregation device is provided, including: a data acquisition module 902, a traversal module 904, a data processing module 906 and a resource allocation module 908, wherein:

[0160] The data acquisition module 902 is used to obtain a first material template for data aggregation in response to a data aggregation processing request; wherein the categories in the category set of the first material template have a hierarchical relationship.

[0161] The traversal module 904 is used to traverse the category set, and for each level of categories, obtain a category element set corresponding to each level of categories.

[0162] Data processing module 906 is used to determine the reference element length of each level of categories based on the length of each category element in the category element set and the number of categories in the category element set, and determine the total number of threads executed to process each level of categories based on the number of categories; based on the relationship between the length of each category element and the reference element length, determine the thread resource requirements for processing each category element.

[0163] The resource allocation module 908 is used to determine the processing priority of each category element according to the total number of thread executions and thread resource requirements, allocate thread resources to each category element based on the processing priority, and complete the aggregation processing of the category element set by executing the allocated thread resources.

[0164] The above-mentioned data aggregation device, in the scenario of traversing and aggregating a large amount of data, for each level of categories in the first material template, by determining and obtaining the category element set corresponding to each level of categories, determines the reference element length of each level of categories and the total number of thread executions for processing each level of categories according to the length of each category element in the category element set and the number of categories in the category element set; for categories of one level, based on the known thread resources of each level, according to the relationship between the length of each category element and the reference element length, determines the thread resource requirement for processing each category element compared to the reference element length, and determines the thread resource requirement of each category element; on this basis, based on the total number of thread executions and the thread resource requirement of each category element, by determining the processing priority of each category element, the granularity of tasks can be adjusted to adapt to the thread resource requirements and processing performance of different category elements, thereby achieving more balanced task allocation and solving the problem of uneven thread resource allocation.

[0165] In an exemplary embodiment, the data processing module 906 is further configured to determine the thread resource requirement of the first category element as the first thread resource requirement if the length of the first category element in the category element set is greater than the reference element length;

[0166] If there is a second category element in the category element set whose length is less than the reference element length, determining the thread resource requirement of the second category element as the second thread resource requirement;

[0167] The number of thread resources occupied by the second thread resource requirement is less than the number of thread resources occupied by the first thread resource requirement.

[0168] In an exemplary embodiment, the category element set includes first category elements and second category elements, and the resource allocation module 908 is further used to divide the category element set according to the reference element length, and obtain the sub-category element set, and determine the proportion of the total number of thread executions according to the number of elements in the sub-category element set;

[0169] If the number of elements in the first category is greater than or equal to the product of the proportion value and the total number of thread executions, it is determined that the processing priority of the elements in the first category is lower than the processing priority of the elements in the second category;

[0170] If the number of elements in the first category is less than the product of the proportion value and the total number of thread executions, it is determined that the processing priority of the elements in the first category is greater than the processing priority of the elements in the second category. In an exemplary embodiment, the resource allocation module 908 is further configured to allocate available thread resources corresponding to the total number of thread executions to the elements in the second category if the processing priority of the elements in the first category is less than the processing priority of the elements in the second category;

[0171] If the processing priority of the first category element is greater than the processing priority of the second category element, the available thread resources corresponding to the total number of thread executions are allocated to the first category element.

[0172] In an exemplary embodiment, the resource allocation module 908 is further configured to divide the category element set according to the reference element length to obtain a first category element set and a second category element set; wherein the element length of each first category element in the first category element set is greater than the reference element length, and the element length of each second category element in the second category element set is less than the reference element length;

[0173] The proportion of the total number of thread executions is determined according to the number of elements in the first category element set and the number of elements in the second category element set.

[0174] In an exemplary embodiment, the data aggregation apparatus includes a mapping module, the mapping module being used to obtain all second material templates associated with a business line and template source data of each second material template;

[0175] Traversing the template source data of each second material template, performing feature extraction on the template source data, and identifying all categories in each second material template for representing the same attribute according to the extracted features;

[0176] An aggregate category corresponding to all categories representing the same attribute is generated, and the aggregate category is used as the category of the first material template. There is a mapping relationship between the category of the first material template and the categories of all second material templates.

[0177] In an exemplary embodiment, the data aggregation device includes an aggregation module, which is used to obtain source data corresponding to the category element set from the template source data of each second material template according to the mapping relationship between the category of the first material template and the categories of all second material templates by executing allocated thread resources; and perform aggregation processing on the source data corresponding to the category element set.

[0178] In an exemplary embodiment, the data processing module 906 is further used to generate a category code for each level of category; determine the category code as a thread identifier of a thread corresponding to the category element, and the thread identifier is used to query and locate the thread processing progress.

[0179] Each module in the above data aggregation device can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.

[0180] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Fig.10 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a method of data aggregation is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0181] Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0182] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0183] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0184] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0185] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0186] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0187] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for data aggregation, characterized in that: The method comprises: In response to a data aggregation processing request, obtaining a first material template for data aggregation; wherein the categories in the category set of the first material template have a hierarchical relationship; Traversing the category set, for each level of categories, obtaining a category element set corresponding to each level of categories; Determine the reference element length of each level of categories according to the length of each category element in the category element set and the number of categories in the category element set, and determine the total number of threads executed to process each level of categories according to the number of categories; Determining thread resource requirements for processing each of the category elements according to a relationship between the length of each of the category elements and the length of the reference element; According to the total number of thread executions and the thread resource requirements, the processing priority of each category element is determined, thread resources are allocated to each category element based on the processing priority, and aggregation processing of the category element set is completed by executing the allocated thread resources.

2. The method according to claim 1, characterized in that The reference element length is an average element length of the category element set, and determining the thread resource requirement for processing each of the category elements according to the relationship between the length of each of the category elements and the reference element length includes: If a length of a first category element in the category element set is greater than the reference element length, determining the thread resource requirement of the first category element as the first thread resource requirement; If there is a second category element in the category element set whose length is less than the reference element length, determining the thread resource requirement of the second category element as the second thread resource requirement; The number of thread resources occupied by the second thread resource requirement is smaller than the number of thread resources occupied by the first thread resource requirement.

3. The method according to claim 2, characterized in that The category element set includes the first category element and the second category element, and determining the processing priority of each of the category elements according to the total number of thread executions and the thread resource requirements includes: Dividing the category element set according to the reference element length to obtain subcategory element sets, and determining the proportion of the total number of thread executions according to the number of elements in the subcategory element sets; If the number of the first category elements is greater than or equal to the product of the proportion value and the total number of thread executions, it is determined that the processing priority of the first category elements is lower than the processing priority of the second category elements; If the number of the first category elements is less than the product of the proportion value and the total number of thread executions, it is determined that the processing priority of the first category elements is greater than the processing priority of the second category elements.

4. The method according to claim 3, characterized in that The allocating thread resources to each of the category elements based on the processing priority, and completing the aggregation processing of the category element set by executing the allocated thread resources, includes: If the processing priority of the first category element is lower than the processing priority of the second category element, allocating the available thread resources corresponding to the total number of thread executions to the second category element; If the processing priority of the first category element is greater than the processing priority of the second category element, the available thread resources corresponding to the total number of thread executions are allocated to the first category element.

5. The method according to claim 3, characterized in that: The step of dividing the category element set according to the reference element length to obtain a subcategory element set, and determining the proportion of the total number of thread executions according to the number of elements in the subcategory element set, includes: Dividing the category element set according to the reference element length to obtain a first category element set and a second category element set; wherein the element length of each first category element in the first category element set is greater than the reference element length, and the element length of each second category element in the second category element set is less than the reference element length; The proportion of the total number of thread executions is determined according to the number of elements in the first category element set and the number of elements in the second category element set.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Acquire all second material templates associated with the business line, and template source data of each of the second material templates; Traversing the template source data of each second material template, performing feature extraction on the template source data, and identifying all categories in each second material template for representing the same attribute according to the extracted features; An aggregate category corresponding to all the categories representing the same attribute is generated, and the aggregate category is used as the category of the first material template. There is a mapping relationship between the category of the first material template and the categories of all the second material templates.

7. The method according to claim 6, characterized in that The step of completing the aggregation processing of the category element set by executing the allocated thread resources includes: By executing the allocated thread resources, according to the mapping relationship between the category of the first material template and the categories of all the second material templates, the source data corresponding to the category element set is acquired from the template source data of each of the second material templates; Aggregation processing is performed on the source data corresponding to the category element set.

8. The method according to claim 7, characterized in that The method further comprises: For each level of categories, generating a category code for each level of categories; The category code is determined as a thread identifier of a thread corresponding to the category element, and the thread identifier is used to query and locate thread processing progress.

9. A data aggregation device, characterized in that: The device comprises: A data acquisition module, configured to acquire a first material template for data aggregation in response to a data aggregation processing request; wherein the categories in the category set of the first material template have a hierarchical relationship; A traversal module, used to traverse the category set, and for each level of categories, obtain a category element set corresponding to each level of categories; A data processing module, configured to determine the reference element length of each level of categories according to the length of each category element in the category element set and the number of categories in the category element set, and determine the total number of threads executed to process each level of categories according to the number of categories; Determining thread resource requirements for processing each of the category elements according to a relationship between the length of each of the category elements and the length of the reference element; A resource allocation module is used to determine the processing priority of each of the category elements according to the total number of thread executions and the thread resource requirements, allocate thread resources to each of the category elements based on the processing priority, and complete the aggregation processing of the category element set by executing the allocated thread resources.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.