A data aggregation method, device and medium for industry identification data

By converting data aggregation requests into multiple subtasks, determining dependencies, and generating computational relationships, the problem of manual aggregation being labor-intensive and prone to errors is solved, achieving efficient and accurate aggregation of industry-specific identification data.

CN114511290BActive Publication Date: 2026-05-01浪潮工业互联网股份有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
浪潮工业互联网股份有限公司
Filing Date
2022-01-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, manually compiling industry identification data is labor-intensive and prone to errors, resulting in low compilation efficiency.

Method used

The data aggregation request is converted into multiple subtasks. The dependencies between the subtasks are determined, the operation relationships are generated, and the data is aggregated according to the operation relationships. Duplicate and similar data are removed to ensure the accuracy of the aggregation results.

Benefits of technology

It improves the accuracy and efficiency of data aggregation, avoids redundant calculations and system burden, and ensures the reliability of the aggregation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511290B_ABST
    Figure CN114511290B_ABST
Patent Text Reader

Abstract

The embodiment of the specification discloses a data aggregation method, device and medium for industry identification data, the method comprising: receiving a plurality of industry identification data and a data aggregation request for the plurality of industry identification data, wherein the industry identification data comprises node identification, identification registration data and identification analysis data of the industry identification data; generating a plurality of data aggregation sub-tasks according to the data aggregation request, judging whether there is a dependency relationship between the plurality of data aggregation sub-tasks; if it is determined that there is a dependency relationship between the plurality of data aggregation sub-tasks, generating an operation relationship between the data aggregation sub-tasks according to the dependency relationship between the data aggregation sub-tasks; aggregating the industry identification data to be aggregated in the plurality of data aggregation sub-tasks to generate aggregation sub-data corresponding to the plurality of data aggregation sub-tasks; and according to the operation relationship between the plurality of data aggregation sub-tasks, operating each aggregation sub-data to obtain aggregation data corresponding to the plurality of industry identification data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of identifier resolution technology, and in particular to a data aggregation method, device and medium for industry identifier data. Background Technology

[0002] With the development of industrial internet technology, the application of identifier resolution systems is becoming increasingly widespread. The identifier resolution system is divided into five levels: root node, national top-level node, secondary nodes, enterprise nodes, and public recursive nodes. Industrial internet identifier secondary nodes provide identifier registration and resolution services to industries, serving as the intermediate link in the industrial internet identifier resolution system and directly providing services to industries and enterprises.

[0003] Regulatory authorities can analyze industry development by using the registration and resolution information of secondary nodes. For example, they can analyze the food industry in northern China or the food industry in North China. Because the secondary nodes are relatively dispersed, data aggregation is necessary before analysis. Currently, data aggregation is done manually, which is labor-intensive, error-prone, and inefficient. Summary of the Invention

[0004] This specification provides one or more embodiments of a method, device, and medium for summarizing industry identification data, which addresses the following technical problem: manual summarization is labor-intensive and prone to errors, resulting in low summarization efficiency.

[0005] One or more embodiments of this specification employ the following technical solutions:

[0006] This specification provides one or more embodiments of a method for aggregating industry identification data. The method includes: receiving multiple industry identification data and a data aggregation request for the multiple industry identification data, wherein the industry identification data includes node identifiers, identifier registration data, and identifier resolution data of the industry identification data; generating multiple data aggregation sub-tasks according to the data aggregation request, wherein the data aggregation sub-tasks include industry identification data to be aggregated in the data aggregation sub-tasks; determining whether there is a dependency relationship between the multiple data aggregation sub-tasks; if it is determined that there is a dependency relationship between the multiple data aggregation sub-tasks, generating an operation relationship between the data aggregation sub-tasks according to the dependency relationship between the data aggregation sub-tasks; aggregating the industry identification data to be aggregated in the multiple data aggregation sub-tasks to generate aggregated sub-data corresponding to the multiple data aggregation sub-tasks; and performing operations on each aggregated sub-data according to the operation relationship between the multiple data aggregation sub-tasks to obtain aggregated data corresponding to the multiple industry identification data.

[0007] Furthermore, the data aggregation subtask further includes: multiple node identifiers corresponding to the identifier data to be aggregated; before generating multiple data aggregation subtasks according to the data aggregation request, the method further includes: classifying the multiple industry identifier data according to the node identifiers in each industry identifier data, and storing them respectively in corresponding data tables; setting the task identifier of the data aggregation subtask according to the multiple node identifiers corresponding to the industry identifier data to be aggregated in the data aggregation subtask, and establishing a mapping relationship between the task identifier and the multiple node identifiers; obtaining the task identifier of the current data aggregation subtask, and determining the multiple node identifiers corresponding to the industry identifier data to be aggregated in the current aggregation subtask through the mapping relationship between the task identifier and the multiple node identifiers; obtaining the industry identifier data to be aggregated from the corresponding data table through the multiple node identifiers corresponding to the industry identifier data to be aggregated.

[0008] Furthermore, if it is determined that there is a dependency relationship between the multiple data aggregation subtasks, the method further includes: randomly selecting one data aggregation subtask as a first subtask among the multiple data aggregation subtasks; determining a second subtask based on the dependency relationship of the first subtask, wherein the second subtask is a data aggregation subtask that the first subtask depends on; setting the second subtask before the first subtask to generate an arrangement order of the first subtask and the second subtask; generating a aggregation order consistent with the arrangement order according to the arrangement order, so as to perform data aggregation sequentially according to the aggregation order.

[0009] Further, determining whether there is a dependency relationship between multiple data aggregation subtasks specifically includes: obtaining multiple node identifiers of the industry identifier data to be aggregated in the data aggregation subtasks; if, through the mapping relationship between the task identifier corresponding to the data aggregation subtask and the multiple node identifiers, it is determined that there is a third subtask and a fourth subtask among the multiple data aggregation subtasks, then it is determined that there is a dependency relationship between the third subtask and the fourth subtask, wherein the third subtask and the fourth subtask satisfy one or more of the following conditions: the multiple node identifiers corresponding to the third subtask are the same as some of the node identifiers corresponding to the fourth subtask; the data aggregation result of the third subtask is the industry identifier data to be aggregated by the fourth subtask.

[0010] Furthermore, after obtaining the industry identifier data to be summarized in the corresponding data table, the method further includes: obtaining multiple industry identifier data belonging to the same node identifier in the data table corresponding to the node identifier; calculating the similarity between any two industry identifier data among the multiple industry identifier data; determining the relationship between the similarity and a preset similarity threshold; if the similarity is higher than the preset similarity threshold, then determining that the two industry identifier data are conflicting data, and removing either of the two industry identifier data.

[0011] Furthermore, the process of receiving multiple industry identifier data and requesting data aggregation for the multiple industry identifier data specifically includes: determining the secondary nodes corresponding to the multiple industry identifier data based on the data source of the multiple industry identifier data to be acquired, wherein the secondary nodes are used to store the identifier registration data and the identifier resolution data; determining the multiple data source types corresponding to the multiple industry identifier data through a pre-set correspondence between the secondary nodes and data source types; dividing the multiple industry identifier data into multiple groups of industry identifier data according to the multiple data source types; calling the data interface corresponding to the data source type corresponding to each group of industry identifier data to acquire each group of industry identifier data; and determining the multiple industry identifier data to be acquired based on each group of industry identifier data.

[0012] Furthermore, after determining whether there is a dependency relationship between the multiple data aggregation subtasks, the method further includes: if it is determined that there is no dependency relationship between the multiple data aggregation subtasks, then performing data aggregation on the industry identification data to be aggregated in the multiple data aggregation subtasks to generate aggregated sub-data corresponding to the multiple data aggregation subtasks; and summing the aggregated sub-data corresponding to the multiple data aggregation subtasks to obtain aggregated data corresponding to the multiple industry identification data.

[0013] Further, the step of generating the operational relationship between the data aggregation subtasks based on the dependencies between them specifically includes: if the dependencies between the data aggregation subtasks are specified dependencies, where the specified dependencies are that the data aggregation result of the third subtask is the industry identifier data to be aggregated by the fourth subtask, then the operational relationship between the third subtask and the fourth subtask is determined to be an addition operation; if the dependencies between the data aggregation subtasks are preset dependencies, where the preset dependencies are that multiple node identifiers corresponding to the third subtask are the same as some of the node identifiers corresponding to the fourth subtask, then the aggregation result of the third subtask is set to zero and added to the aggregation result of the fourth subtask.

[0014] This specification provides one or more embodiments of a data aggregation device for industry identification data, including:

[0015] At least one processor; and,

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions executable by the at least one processor. These instructions, when executed by the at least one processor, enable the at least one processor to: receive multiple industry identification data sets and a data aggregation request for the multiple industry identification data sets, wherein the industry identification data sets include node identifiers, identifier registration data, and identifier resolution data; generate multiple data aggregation subtasks based on the data aggregation request, wherein each data aggregation subtask includes industry identification data to be aggregated within the data aggregation subtask; determine whether there is a dependency relationship between the multiple data aggregation subtasks; if a dependency relationship is determined to exist between the multiple data aggregation subtasks, generate an operational relationship between the data aggregation subtasks based on the dependency relationship; aggregate the industry identification data to be aggregated within the multiple data aggregation subtasks to generate aggregated sub-data sets corresponding to the multiple data aggregation subtasks; and perform operations on each aggregated sub-data set according to the operational relationship between the multiple data aggregation subtasks to obtain aggregated data sets corresponding to the multiple industry identification data sets.

[0018] This specification provides one or more embodiments of a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:

[0019] The system receives multiple industry identifier data and a data aggregation request for the multiple industry identifier data, wherein the industry identifier data includes node identifiers, identifier registration data, and identifier resolution data of the industry identifier data; generates multiple data aggregation sub-tasks according to the data aggregation request, wherein each data aggregation sub-task includes the industry identifier data to be aggregated in the data aggregation sub-task; determines whether there is a dependency relationship between the multiple data aggregation sub-tasks; if it is determined that there is a dependency relationship between the multiple data aggregation sub-tasks, generates an operation relationship between the data aggregation sub-tasks according to the dependency relationship between the multiple data aggregation sub-tasks; aggregates the industry identifier data to be aggregated in the multiple data aggregation sub-tasks to generate aggregated sub-data corresponding to the multiple data aggregation sub-tasks; and performs operations on each aggregated sub-data according to the operation relationship between the multiple data aggregation sub-tasks to obtain the aggregated data corresponding to the multiple industry identifier data.

[0020] The above-mentioned at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: converting the data aggregation request into multiple data aggregation sub-tasks, determining the dependency relationship between each data aggregation sub-task, determining the operation relationship between each data aggregation sub-task through the dependency relationship, and obtaining the aggregated data according to the operation relationship, avoiding the situation of duplicate data summation in the aggregated data, and ensuring the accuracy of the data aggregation result. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0022] Figure 1 A flowchart illustrating a data aggregation method for industry identification data provided in an embodiment of this specification;

[0023] Figure 2 This is a schematic diagram of the structure of a data aggregation device for industry identification data provided in an embodiment of this specification. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0025] With the development of industrial internet technology, the application of identifier resolution systems is becoming increasingly widespread. The identifier resolution system is divided into five levels: root node, national top-level node, secondary nodes, enterprise nodes, and public recursive nodes. Industrial internet identifier secondary nodes provide identifier registration and resolution services to industries, serving as the intermediate link in the industrial internet identifier resolution system and directly providing services to industries and enterprises.

[0026] Regulatory authorities can analyze industry development across various sectors by utilizing the registration and resolution information of secondary nodes. For example, they can analyze the food industry in northern China or the food industry in North China. However, due to the relatively dispersed distribution of secondary nodes, data aggregation is necessary before analysis. When multiple data aggregation tasks are involved, these tasks are interconnected, and directly aggregating the data leads to inaccurate results. Furthermore, current technologies rely on manual aggregation of various identifier data, which is labor-intensive, error-prone, and inefficient.

[0027] This specification provides a flowchart illustrating a method for aggregating industry identification data. The executing entity in this embodiment can be a server or any device with data processing capabilities. Figure 1 This is a flowchart illustrating a data aggregation method for industry identification data provided in an embodiment of this specification, as shown below. Figure 1 As shown, the method mainly includes the following steps:

[0028] Step S101: Obtain multiple industry identification data and a data aggregation request for the multiple industry identification data.

[0029] It should be noted that industry identification data includes node identifiers, identifier registration data, and identifier resolution data. Node identifiers store the identification information of the nodes corresponding to the industry identification data. Generally, these nodes are secondary nodes, which provide identifier registration and resolution services to the industry. They are the intermediate link in the industrial internet identifier resolution system, directly providing services to industries and enterprises. Therefore, the secondary nodes in each region store the identifier registration data and identifier resolution data for each industry. This type of industry identification data allows for the analysis of industry development trends.

[0030] In one embodiment of this specification, data aggregation is obtained. For example, the data aggregation request may be to aggregate food industry identification data from various regions in northern China for the purpose of analyzing the development trend of the northern food market. Based on the data aggregation request, multiple industry identification data are obtained. The corresponding secondary nodes are determined by the regions specified in the data aggregation request. For example, secondary nodes for the Beijing region include secondary nodes for Beijing, Jinan, Qingdao, and other regions. The industry identification data in each secondary node is determined by the industry specified in the data aggregation request.

[0031] Specifically, the process involves receiving multiple industry identifier data sets and data aggregation requests for these sets. This includes: determining the secondary nodes corresponding to the multiple industry identifier data sets based on their data sources; determining the multiple data source types corresponding to the multiple industry identifier data sets through a pre-defined correspondence between secondary nodes and data source types; dividing the multiple industry identifier data sets into multiple groups based on the multiple data source types; calling the data interface corresponding to the data source type for each group of industry identifier data sets to obtain each group of industry identifier data sets; and determining the multiple industry identifier data sets to be acquired based on each group of industry identifier data sets.

[0032] In one embodiment of this specification, based on the data source of the multiple industry identifier data to be acquired, the secondary nodes corresponding to the multiple industry identifier data are determined. After determining the secondary nodes and their respective industries for the identifier data to be acquired, a correspondence between the secondary nodes and data types is generated based on the data source type of the secondary nodes corresponding to the identifier data. For example, the secondary nodes in region A support the data source type MySQL, and the secondary nodes in region B support the data source type Oracle. Based on the node identifier of the secondary nodes in each region, a correspondence between the node identifier and the data type is generated. Based on the node identifier and the correspondence, the data source type of the industry identifier data to be acquired is determined. The multiple industry identifier data are divided into multiple groups of industry identifier data according to the data source type. Each group of industry identifier data has the same data source type, which facilitates calling the data interface corresponding to the data source type to acquire each group of industry identifier data. By acquiring multiple groups of industry identifier data according to the above method, the multiple industry identifier data to be acquired is obtained.

[0033] Step S102: Generate multiple data aggregation subtasks based on the data aggregation request.

[0034] The data aggregation subtask includes the industry identifier data to be aggregated in the data aggregation subtask.

[0035] In one embodiment of this specification, a data aggregation request is determined based on the industry and region to be analyzed. This request includes multiple data aggregation sub-tasks. For example, to analyze the development trends of the food industry in various northern regions, it is necessary to aggregate industry identification data for each region, including industry analysis for North China, industry analysis for Northwest China, and industry development analysis for the entire northern region. This generates multiple data aggregation sub-tasks: data aggregation for Northwest China, data aggregation for North China, and data aggregation for the entire northern region.

[0036] Specifically, before generating multiple data aggregation subtasks based on the data aggregation request, the process includes: classifying the multiple industry identifier data according to the node identifiers in each industry identifier data, and storing them in corresponding data tables; setting the task identifier of the data aggregation subtask based on the multiple node identifiers corresponding to the industry identifier data to be aggregated in the data aggregation subtask, and establishing a mapping relationship between the task identifier and the multiple node identifiers; obtaining the task identifier of the current data aggregation subtask, and determining the multiple node identifiers corresponding to the industry identifier data to be aggregated in the current aggregation subtask through the mapping relationship between the task identifier and the multiple node identifiers; and obtaining the industry identifier data to be aggregated from the corresponding data table using the multiple node identifiers corresponding to the industry identifier data to be aggregated.

[0037] In one embodiment of this specification, multiple industry identification data are categorized according to the node identifiers in each industry identification data, and stored in corresponding data tables. For example, data for node A is stored in data table A, and data for node B is stored in data table B. Based on the multiple node identifiers corresponding to the industry identification data to be summarized in the data summarization subtask, a task identifier for the data summarization subtask is set, and a mapping relationship between the task identifier and multiple node identifiers is established. For example, if the task identifier for the first task is set to 1, and the corresponding identification data in this task corresponds to nodes A and B, then a one-to-two mapping relationship is established between task identifier 1 and nodes A and B.

[0038] Obtain the task identifier of the current data aggregation subtask. Using the task identifier and the mapping relationship between the task identifier and multiple node identifiers, determine the multiple node identifiers corresponding to the industry identifier data to be aggregated in the current aggregation subtask. After obtaining the multiple node identifiers in the current data aggregation subtask, retrieve the industry identifier data to be aggregated from the data tables corresponding to each node using these node identifiers.

[0039] Specifically, after obtaining the industry identifier data to be summarized in the corresponding data table, the process also includes: obtaining multiple industry identifier data belonging to the same node identifier in the data table corresponding to the node identifier; calculating the similarity between any two industry identifier data among the multiple industry identifier data; determining the relationship between the similarity and the preset similarity threshold; if the similarity is higher than the preset similarity threshold, then determining that the two industry identifier data are conflicting data and removing any one of the two industry identifier data.

[0040] In the actual data aggregation process, there may be duplicate or similar data. If the data is aggregated directly, the aggregation results will be inaccurate.

[0041] In one embodiment of this specification, since the industry identifier data stored in each data table belongs to the same node, industry identifier data belonging to the same node in the same data table are obtained, and these data are compared for similarity to remove duplicate and similar data. The similarity between any two industry identifier data is calculated. This similarity can be obtained by extracting data features and calculating the similarity of these features. A similarity threshold is set based on the characteristics of the industry identifier data. The relationship between the similarity of any two industry identifier data and the similarity threshold is determined. If the similarity between any two industry identifier data is higher than the similarity threshold, the two industry identifier data are considered conflicting data, and one of them is removed during aggregation. The remaining industry identifier data is then used for data aggregation.

[0042] Step S103: Determine whether there are dependencies between multiple data aggregation subtasks.

[0043] Specifically, multiple node identifiers of the industry identifier data to be summarized in the data summarization subtask are obtained; if, through the mapping relationship between the task identifier and the multiple node identifiers corresponding to the data summarization subtask, it is determined that there is a third subtask and a fourth subtask among the multiple data summarization subtasks, then it is determined that there is a dependency relationship between the third subtask and the fourth subtask, wherein the third subtask and the fourth subtask satisfy one or more of the following conditions: the multiple node identifiers corresponding to the third subtask are the same as some of the node identifiers corresponding to the fourth subtask; the data summarization result of the third subtask is the industry identifier data to be summarized in the fourth subtask.

[0044] In one embodiment of this specification, multiple identifier nodes of the industry identifier data to be summarized in each data summarization subtask are obtained. For example, the node identifier corresponding to the first summarization subtask is node A, and the node identifiers corresponding to the second summarization subtask are node A, node B, and node C. In the two summarization subtasks, the nodes of the first summarization subtask are part of the nodes of the second summarization subtask, indicating that there is a dependency relationship between the first and second summarization subtasks. In addition, the data summarization result of the third subtask is the industry identifier data to be summarized in the fourth subtask, indicating that there is a dependency relationship between the third and fourth subtasks. Continuing with the previous example, the summarization result of the first summarization subtask is the sum of the industry identifier data corresponding to node A, while the second summarization subtask requires the summarization result of the first summarization subtask. Furthermore, if the summarization result of the first summarization subtask is inaccurate, it will also lead to the inaccuracy of the summarization result of the second summarization subtask. In this case, it can also be determined that there is a dependency relationship between the first and second summarization subtasks.

[0045] Specifically, if a dependency relationship is determined between multiple data aggregation subtasks, the process further includes: randomly selecting one data aggregation subtask as the first subtask from among the multiple data aggregation subtasks; determining the second subtask based on the dependency relationship of the first subtask, wherein the second subtask is the data aggregation subtask that the first subtask depends on; setting the second subtask before the first subtask to generate an order of arrangement for the first and second subtasks; and generating a aggregation order consistent with the arrangement order so that data can be aggregated sequentially according to the aggregation order.

[0046] In one embodiment of this specification, if there is a dependency between two summary subtasks, when calculating each summary subtask, the calculation order can be generated according to the dependency, and the summary calculation can be performed sequentially according to the calculation order, which can avoid the problems of repeated calculation and increased system operating burden.

[0047] In multiple summary subtasks, the data is sorted according to dependencies. For example, the dependencies are as follows: the first subtask depends on the second subtask, and the second subtask depends on the third subtask. Based on the dependency order, subtasks that do not depend on other subtasks are calculated first. Therefore, the order is: the third subtask, the second subtask, and the first subtask. The data from the third subtask, the second subtask, and the first subtask are summarized in the above order.

[0048] Specifically, if it is determined that there is no dependency between multiple data aggregation subtasks, then the industry identification data to be aggregated in the multiple data aggregation subtasks is aggregated to generate aggregated sub-data corresponding to the multiple data aggregation subtasks; the aggregated sub-data corresponding to the multiple data aggregation subtasks is summed to obtain the aggregated data corresponding to the multiple industry identification data.

[0049] In one embodiment of this specification, there may be a situation where there are no dependencies between multiple data aggregation subtasks. If there are no dependencies between multiple data aggregation subtasks, it means that the multiple data aggregation subtasks are independent subtasks. Data can be directly aggregated for each subtask to obtain aggregated sub-data. Then, the data aggregated sub-data of multiple data aggregations can be summed to obtain aggregated data of multiple industry identifiers.

[0050] Step S104: If it is determined that there is a dependency relationship between multiple data aggregation subtasks, then the operation relationship between each data aggregation subtask is generated according to the dependency relationship between each data aggregation subtask.

[0051] Specifically, if there are dependencies between multiple data aggregation subtasks, the operation relationship between each data aggregation subtask is generated based on the dependencies between them. Specifically, if the dependencies between the data aggregation subtasks are specified dependencies, where the data aggregation result of the third subtask is the industry identifier data to be aggregated by the fourth subtask, then the operation relationship between the third and fourth subtasks is determined to be an addition operation; if the dependencies between the data aggregation subtasks are preset dependencies, where the preset dependencies are that some node identifiers of multiple node identifiers corresponding to the third subtask are the same as some node identifiers of multiple node identifiers corresponding to the fourth subtask, then the aggregation result of the third subtask is set to zero and added to the aggregation result of the fourth subtask.

[0052] In one embodiment of this specification, if there are dependencies between multiple data aggregation subtasks, it is necessary to determine the type of dependency and set the operational relationships between the multiple aggregation subtasks according to different dependencies. If a third subtask exists, and the data aggregation result of the third subtask is the industry identifier data to be aggregated by the fourth subtask, that is, the output of the third subtask is the input of the fourth subtask, then to calculate the aggregated data of the third and fourth subtasks, the data of the third and fourth subtasks needs to be summed to obtain the aggregated data. If multiple node identifiers corresponding to the third subtask are the same as some node identifiers corresponding to the fourth subtask, that is, there is overlapping data in the data to be calculated in both, if the data of the two are directly summed, data duplication will occur, resulting in the aggregated result including an over-calculation of the part in the third subtask. Therefore, the third subtask can be set to zero, and then summed with the aggregated result of the fourth subtask.

[0053] Step S105: Perform data aggregation on the industry identifier data to be aggregated in multiple data aggregation sub-tasks to generate aggregated sub-data corresponding to multiple data aggregation sub-tasks.

[0054] In one embodiment of this specification, after determining the dependencies between the subtasks and identifying their respective computational methods, it is necessary to aggregate the industry identification data within each data aggregation subtask to generate the aggregation result for each data aggregation subtask, i.e., the aggregated sub-data. During this data aggregation process, conflicting data needs to be removed to ensure the accuracy of the settlement results for each data aggregation subtask.

[0055] Step S106: Based on the operational relationships between multiple data aggregation sub-tasks, perform operations on each aggregation sub-data to obtain the aggregation data corresponding to multiple industry identifier data.

[0056] In one embodiment of this specification, numerical calculations are performed based on the operational relationships between each data aggregation subtask and the aggregation results of each data aggregation subtask to generate aggregated data corresponding to multiple industry-identified data. The aggregated data is then presented intuitively in the form of data reports. Furthermore, the data can be aggregated and displayed in a diversified and dynamic manner using an ECharts plugin, facilitating industry development analysis based on the aggregated data.

[0057] The above technical solution converts data aggregation requests into multiple data aggregation subtasks, determines the dependencies between each subtask, identifies the computational relationships between them, and obtains aggregated data according to these relationships. This avoids duplicate data summation in the aggregated data and ensures the accuracy of the data aggregation results.

[0058] This specification also provides a data aggregation device for industry identification data, such as... Figure 2 As shown, the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: acquire multiple industry identification data and a data aggregation request for the multiple industry identification data, wherein the industry identification data includes node identifiers, identifier registration data, and identifier resolution data of the industry identification data; generate multiple data aggregation sub-tasks according to the data aggregation request, wherein the data aggregation sub-tasks include industry identification data to be aggregated in the data aggregation sub-tasks; determine whether there is a dependency relationship between the multiple data aggregation sub-tasks; if it is determined that there is a dependency relationship between the multiple data aggregation sub-tasks, generate an operation relationship between the data aggregation sub-tasks according to the dependency relationship between the data aggregation sub-tasks; aggregate the industry identification data to be aggregated in the multiple data aggregation sub-tasks to generate aggregated sub-data corresponding to the multiple data aggregation sub-tasks; and perform operations on each aggregated sub-data according to the operation relationship between the multiple data aggregation sub-tasks to obtain aggregated data corresponding to the multiple industry identification data.

[0059] This specification also provides a non-volatile computer storage medium storing computer-executable instructions, which are configured to: receive multiple industry identification data and a data aggregation request for the multiple industry identification data, wherein the industry identification data includes node identifiers, identifier registration data, and identifier resolution data of the industry identification data; generate multiple data aggregation sub-tasks according to the data aggregation request, wherein the data aggregation sub-tasks include industry identification data to be aggregated in the data aggregation sub-tasks; determine whether there is a dependency relationship between the multiple data aggregation sub-tasks; if it is determined that there is a dependency relationship between the multiple data aggregation sub-tasks, generate an operation relationship between the data aggregation sub-tasks according to the dependency relationship between the data aggregation sub-tasks; aggregate the industry identification data to be aggregated in the multiple data aggregation sub-tasks to generate aggregated sub-data corresponding to the multiple data aggregation sub-tasks; and perform operations on each aggregated sub-data according to the operation relationship between the multiple data aggregation sub-tasks to obtain aggregated data corresponding to the multiple industry identification data.

[0060] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0061] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0062] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.

Claims

1. A method for aggregating industry identification data, characterized in that, The method includes: Receive multiple industry identification data and a data aggregation request for the multiple industry identification data, wherein the industry identification data includes the node identifier, identifier registration data and identifier resolution data of the industry identification data; Based on the data aggregation request, multiple data aggregation sub-tasks are generated, wherein the data aggregation sub-tasks include industry identification data to be aggregated in the data aggregation sub-tasks; Determine if there are dependencies between multiple data aggregation subtasks; If it is determined that there is a dependency relationship between the multiple data aggregation subtasks, then the operation relationship between the data aggregation subtasks is generated according to the dependency relationship between the data aggregation subtasks. The industry identification data to be summarized in the multiple data summarization sub-tasks are summarized to generate the summary sub-data corresponding to the multiple data summarization sub-tasks; Based on the operational relationship between the multiple data aggregation sub-tasks, each aggregation sub-data is processed to obtain the aggregation data corresponding to the multiple industry identification data; The data aggregation subtask further includes: multiple node identifiers corresponding to the identifier data to be aggregated; before generating multiple data aggregation subtasks according to the data aggregation request, it also includes: According to the node identifiers in the industry identifier data, the multiple industry identifier data are classified and stored in the corresponding data tables respectively; Based on the multiple node identifiers corresponding to the industry identifier data to be summarized in the data summarization subtask, set the task identifier of the data summarization subtask, and establish a mapping relationship between the task identifier and the multiple node identifiers; Obtain the task identifier of the current data aggregation subtask, and determine the multiple node identifiers corresponding to the industry identifier data to be aggregated in the current aggregation subtask through the mapping relationship between the task identifier and the multiple node identifiers; The industry identification data to be summarized is obtained from the corresponding data table by using multiple node identifiers corresponding to the industry identification data to be summarized. The determination of whether there is a dependency relationship between multiple data aggregation subtasks specifically includes: Obtain multiple node identifiers of the industry identifier data to be summarized in the data summarization subtask; If, through the mapping relationship between the task identifier corresponding to the data aggregation subtask and the multiple node identifiers, it is determined that there are a third subtask and a fourth subtask among the multiple data aggregation subtasks, then it is determined that there is a dependency relationship between the third subtask and the fourth subtask. The third subtask and the fourth subtask satisfy one or more of the following conditions: The third subtask corresponds to multiple node identifiers that are the same as some of the node identifiers in the multiple node identifiers corresponding to the fourth subtask; The data aggregation result of the third subtask is the industry identification data to be aggregated in the fourth subtask; If a dependency relationship is determined to exist between the multiple data aggregation subtasks, the method further includes: Among multiple data aggregation subtasks, one data aggregation subtask is randomly selected as the first subtask. Based on the dependency relationship of the first subtask, the second subtask is determined, wherein the second subtask is the data aggregation subtask that the first subtask depends on. The second subtask is placed before the first subtask, generating the order in which the first and second subtasks are arranged. Generate a summary order consistent with the stated arrangement order, so that data can be summarized sequentially according to the summary order. After obtaining the industry identifier data to be summarized from the corresponding data table, the method further includes: In the data table corresponding to the node identifier, retrieve the data of multiple industry identifiers belonging to the same node identifier; Calculate the similarity between any two industry identifier data among the multiple industry identifier data; The relationship between the similarity and a preset similarity threshold is determined. If the similarity is higher than the preset similarity threshold, the two industry identifier data are determined to be conflicting data, and either of the two industry identifier data is removed. The process of receiving multiple industry identifier data and the data aggregation request for the multiple industry identifier data specifically includes: Based on the data sources of the multiple industry identifier data to be acquired, secondary nodes corresponding to the multiple industry identifier data are determined. The secondary nodes are used to store the identifier registration data and the identifier resolution data. By establishing a pre-defined correspondence between secondary nodes and data source types, the multiple data source types corresponding to the multiple industry identifier data are determined. According to the multiple data source types, the multiple industry identification data are divided into multiple groups of industry identification data; Based on the data source type corresponding to each group of industry identification data, call the data interface corresponding to the data source type to obtain each group of industry identification data; Based on each set of industry identifier data, the multiple industry identifier data to be acquired are determined; The step of generating the computational relationships between the data aggregation subtasks based on their dependencies specifically includes: If the dependency relationship between the data aggregation subtasks is a specified dependency relationship, and the specified dependency relationship is that the data aggregation result of the third subtask is the industry identification data to be aggregated by the fourth subtask, then the operation relationship between the third subtask and the fourth subtask is determined to be an addition operation. If the dependency relationship between the data aggregation subtasks is a preset dependency relationship, wherein the third subtask corresponds to multiple node identifiers that are the same as some of the node identifiers in the multiple node identifiers corresponding to the fourth subtask, then the aggregation result of the third subtask is set to zero and summed with the aggregation result of the fourth subtask.

2. The data aggregation method for industry identification data according to claim 1, characterized in that, After determining whether there is a dependency relationship between multiple data aggregation subtasks, the method further includes: If it is determined that there is no dependency relationship between the multiple data aggregation subtasks, then the industry identification data to be aggregated in the multiple data aggregation subtasks is aggregated to generate aggregated sub-data corresponding to the multiple data aggregation subtasks. The aggregated sub-data corresponding to the multiple data aggregation sub-tasks are summed to obtain the aggregated data corresponding to the multiple industry identification data.

3. A data aggregation device for industry identification data, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-2.

4. A non-volatile computer storage medium storing computer-executable instructions, the computer-executable instructions being configured to perform the method as described in any one of claims 1-2.

Citation Information

Patent Citations

  • Method and device for data summarization

    CN102929929A

  • Data processing method and data processing device

    CN103793349A

  • Spreadsheet data processing method and device, equipment and storage medium

    CN113420537A