A big data aggregation method and system for fast service

Through data heterogeneity, online processing node pre-aggregation and pipeline aggregation, the problem of inefficiency in the big data aggregation method is solved, and the fast service of big data multi-dimensional information processing is realized.

CN115438250BActive Publication Date: 2025-09-02BEIJING CYBER BASS DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211181702.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-09-02
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

In the prior art, big data aggregation methods are difficult to realize multi-dimensional information pre-aggregation and online analysis processing for fast services, resulting in inefficiency.

Method used

Through data heterogeneity, online processing node pre-aggregation, multi-dimensional information bucket aggregation and pipeline aggregation, rapid service heterogeneous multi-dimensional aggregation of big data is achieved.

Benefits of technology

It greatly improves the efficiency and service efficiency of big data aggregation, can quickly process multi-dimensional information, and meet the needs of fast service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438250B_ABST
    Figure CN115438250B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for aggregating big data for quick services, comprising: obtaining multidimensional information of the big data for quick services through data heterogeneity; performing pre-aggregation and online analysis processing on the multidimensional information of the big data through an online processing node to obtain online pre-aggregated multidimensional information; performing big data bucket aggregation and big data metric aggregation on the online pre-aggregated multidimensional information to obtain a first aggregation result of the big data; performing pipeline aggregation on the first aggregation result of the big data to obtain a second aggregation result of the big data, thereby realizing heterogeneous multidimensional aggregation of the big data for quick services; the present invention also discloses a big data aggregation system for quick services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data aggregation, and more specifically, to a method and system for fast service-oriented big data aggregation. Background Art

[0002] At present, big data aggregation of commonly used services is achieved through data isomorphism of low-dimensional information; there are still many problems such as how to perform pre-aggregation online analytical processing, obtain online pre-aggregated multi-dimensional information, perform big data aggregation to obtain big data results, and realize big data aggregation for fast services; therefore, it is necessary to propose a big data aggregation method and system for fast services to at least partially solve the problems existing in the existing technology. Summary of the Invention

[0003] A series of simplified concepts are introduced in the summary of the invention, which will be further explained in detail in the specific implementation method. The summary of the invention does not mean to attempt to limit the key features and necessary technical features of the technical solution for protection, nor does it mean to attempt to determine the scope of protection of the technical solution for protection.

[0004] To at least partially solve the above problems, the present invention provides a method for aggregating big data for fast service, comprising:

[0005] S100, which uses data heterogeneity to obtain multi-dimensional information of fast service big data;

[0006] S200, performing pre-aggregation online analysis processing on the big data multi-dimensional information through the online processing node to obtain online pre-aggregated multi-dimensional information;

[0007] S300, performing big data bucket aggregation and big data metric aggregation on the online pre-aggregated multi-dimensional information to obtain a first big data aggregation result;

[0008] S400, pipeline aggregation is performed on the first aggregation result of the big data to obtain the second aggregation result of the big data, thereby realizing heterogeneous multi-dimensional aggregation of the big data for fast service.

[0009] Preferably, the S100 includes:

[0010] S101, classifying and splitting the big data for fast services into multiple service type data according to service types;

[0011] S102, performing data heterogeneity processing on data of multiple service types to obtain service type heterogeneity processed data;

[0012] S103, upgrading the dimension of the heterogeneous processing data of service types to obtain multi-dimensional information of fast service big data.

[0013] Preferably, the S200 includes:

[0014] S201, distributing the big data multidimensional information to the set range of online processing nodes according to the number of online processing nodes in the set range;

[0015] S202, establishing a pre-aggregated OLAP framework for OLAP nodes within a set range;

[0016] S203: Perform pre-aggregation OLAP according to the aggregate OLAP framework to obtain online pre-aggregation multi-dimensional information.

[0017] Preferably, the S300 includes:

[0018] S301, centrally backing up the online pre-aggregated multi-dimensional information to obtain the pre-aggregated multi-dimensional information centralized backup data;

[0019] S302, performing big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data;

[0020] S303: Perform big data metric aggregation on the bucket-type aggregated data to obtain a first big data aggregation result.

[0021] Preferably, the S400 includes:

[0022] S401, transmitting the first aggregation result of the big data to the data processing pipeline input end of the data pipeline aggregation;

[0023] S402: The data processing pipeline input terminal classifies the first aggregation result of the big data and transmits it to each data processing pipeline for pipeline aggregation processing;

[0024] S403, through pipeline aggregation processing, obtain the second aggregation result of big data, and realize heterogeneous multi-dimensional aggregation of big data for fast service.

[0025] A big data aggregation system for fast service, including:

[0026] The data heterogeneous multi-dimensional construction module converts the big data for fast service into multi-dimensional information through data heterogeneity;

[0027] The pre-aggregation online analysis module performs pre-aggregation online analysis on the multi-dimensional information of the big data through the online processing node to obtain online pre-aggregated multi-dimensional information;

[0028] The multi-dimensional information progressive aggregation module performs big data bucket aggregation and big data metric aggregation on the online pre-aggregated multi-dimensional information to obtain the first big data aggregation result;

[0029] The pipeline aggregation rapid response module performs pipeline aggregation on the first aggregation result of big data to obtain the second aggregation result of big data, realizing heterogeneous multi-dimensional aggregation of big data for rapid service.

[0030] Preferably, the fast service-oriented big data aggregation system according to claim 6 is characterized in that the data heterogeneous multi-dimensional construction module includes:

[0031] The service type classification and splitting sub-module classifies and splits the big data for fast services into multiple service type data according to service types;

[0032] The multi-type data heterogeneous processing submodule performs data heterogeneity processing on multiple service type data to obtain service type heterogeneous processing data; the service type heterogeneous processing of multiple service types also includes: storing the benchmark service type table and the comparison service type table in the character separated value format to be verified in the distributed column-oriented database, the original service type record primary key as the primary key of the distributed column-oriented database table, the non-primary key attribute of the original service type record as a column of the distributed column-oriented database table, different columns belong to different column families, and the column-oriented storage of the distributed column-oriented database is used to improve the response performance when querying a certain column service type; storing the query index table of the verification rule verification field in the distributed column-oriented database In a column-oriented database, the check field serves as the primary key of the distributed column-oriented database query index table, and the original service type record primary key serves as the column name of the query index table. All primary keys belong to the same column family. This service type mode facilitates the addition, deletion, modification, and query of query index table records. The query index table of the service type record timestamp is stored in the distributed column-oriented database, and the service type record timestamp serves as the primary key of the distributed column-oriented database query index table, while the original service type record primary key serves as the column value of the query index table. When the query index table of the check rule check field is stored in the distributed column-oriented database, the query index table is also stored in the index file of the distributed file infrastructure.

[0033] The dimension upgrade multi-dimensional conversion sub-module upgrades the dimension of heterogeneous processing data of service types to obtain multi-dimensional information of fast service big data.

[0034] Preferably, the pre-aggregation online analysis module includes:

[0035] The online processing node distribution submodule distributes the multi-dimensional information of big data to the online processing nodes within the set range according to the number of online processing nodes within the set range;

[0036] The OLAP framework submodule establishes a pre-aggregated OLAP framework for a set range of OLAP nodes;

[0037] The pre-aggregation online analytical processing submodule performs pre-aggregation online analytical processing according to the aggregation online analytical processing framework to obtain online pre-aggregated multi-dimensional information; the pre-aggregation online analytical processing includes: pre-aggregating the source data into pre-aggregated table data represented in the form of a mesh data structure according to the spatial dimension, wherein the level of the mesh data structure represents the spatial hierarchy of the pre-aggregation, and the content of the mesh data structure node includes the spatial area corresponding to the node and the lightweight pre-aggregated data obtained by pre-aggregating the pre-aggregation dimension selected from the data attribute; when generating a report, determining the selected area selected by the user in space, and pre-aggregating the lightweight pre-aggregated data of each mesh data structure node belonging to the selected area according to the query dimension selected by the user to obtain the query result; pre-aggregating the source data into mesh area pre-aggregated table data according to the mesh area division according to the pre-aggregation dimension selected from the data attribute; when generating a report, pre-aggregating the mesh area pre-aggregated table data of the mesh area selected by the user according to the query dimension to obtain the query result.

[0038] Preferably, the multi-dimensional information progressive aggregation module includes:

[0039] The multi-dimensional information centralized backup submodule performs centralized backup of the online pre-aggregated multi-dimensional information to obtain the pre-aggregated multi-dimensional information centralized backup data;

[0040] The big data bucket aggregation submodule performs big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data;

[0041] The big data metric aggregation submodule performs big data metric aggregation on the bucket-type aggregated data to obtain the first big data aggregation result.

[0042] Preferably, the pipeline aggregation rapid response module includes:

[0043] The aggregation transmission pipeline input submodule transmits the first aggregation result of the big data to the data processing pipeline input end of the data pipeline aggregation;

[0044] The classification is input into the pipeline aggregation submodule. The data processing pipeline input end classifies the first aggregation result of the big data and inputs it into each data processing pipeline for pipeline aggregation processing;

[0045] The fast service aggregation implementation submodule obtains the second aggregation result of big data through pipeline aggregation processing, realizing heterogeneous multi-dimensional aggregation of big data for fast services.

[0046] Compared with the prior art, the present invention has at least the following beneficial effects:

[0047] The present invention provides a big data aggregation method for fast service, which obtains fast service big data multidimensional information through data heterogeneity; performs pre-aggregation and online analysis processing on the big data multidimensional information through online processing nodes to obtain online pre-aggregated multidimensional information; performs big data bucket aggregation and big data metric aggregation on the online pre-aggregated multidimensional information to obtain a first big data aggregation result; performs pipeline aggregation on the first big data aggregation result to obtain a second big data aggregation result, thereby realizing heterogeneous multidimensional aggregation of big data for fast service; classifies and splits the big data for fast service into multiple service type data according to service type; performs data heterogeneous processing on the multiple service type data to obtain service type heterogeneous processing data; performs dimension upgrade on the service type heterogeneous processing data to obtain fast service big data multidimensional information; distributes the big data multidimensional information to the set range of online processing nodes according to the number of set range online processing nodes. Processing nodes; establishing a pre-aggregation online analytical processing framework for online processing nodes within a set range; performing pre-aggregation online analytical processing according to the aggregated online analytical processing framework to obtain online pre-aggregated multi-dimensional information; performing centralized back-up of the online pre-aggregated multi-dimensional information to obtain centralized backup data of the pre-aggregated multi-dimensional information; performing big data bucket-type aggregation on the centralized backup data of the pre-aggregated multi-dimensional information to obtain bucket-type aggregated data; performing big data metric aggregation on the bucket-type aggregated data to obtain a first aggregation result of big data; transmitting the first aggregation result of big data to the data processing pipeline input end of the data pipeline aggregation; the data processing pipeline input end classifies the first aggregation result of big data and transmits it to each data processing pipeline for pipeline aggregation processing; obtaining a second aggregation result of big data through pipeline aggregation processing, thereby realizing heterogeneous multi-dimensional aggregation of big data for fast service; being able to perform data heterogeneity to obtain fast service big data multi-dimensional information, thereby greatly improving the efficiency of service and the efficiency of big data aggregation.

[0048] The present invention describes a method and system for aggregating big data for rapid service. Other advantages, objectives, and features of the present invention will be partially reflected in the following description and partially understood by those skilled in the art through research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0050] Figure 1 This is a diagram of a big data aggregation method for fast service described in the present invention.

[0051] Figure 2 FIG1 is an embodiment of a big data aggregation method for fast service according to the present invention.

[0052] Figure 3 This is a block diagram of a big data aggregation system for fast services described in the present invention. DETAILED DESCRIPTION

[0053] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments so that those skilled in the art can implement the invention with reference to the description. Figure 1-3 As shown, the present invention provides a big data aggregation method for fast service, including:

[0054] S100, which uses data heterogeneity to obtain multi-dimensional information of fast service big data;

[0055] S200, performing pre-aggregation online analysis processing on the big data multi-dimensional information through the online processing node to obtain online pre-aggregated multi-dimensional information;

[0056] S300, performing big data bucket aggregation and big data metric aggregation on the online pre-aggregated multi-dimensional information to obtain a first big data aggregation result;

[0057] S400, pipeline aggregation is performed on the first aggregation result of the big data to obtain the second aggregation result of the big data, thereby realizing heterogeneous multi-dimensional aggregation of the big data for fast service.

[0058] The working principle of the above technical solution is: the present invention provides a big data aggregation method for fast service, including: obtaining fast service big data multidimensional information through data heterogeneity; pre-aggregating the big data multidimensional information through online processing nodes and performing online analysis and processing to obtain online pre-aggregated multidimensional information; performing big data bucket aggregation and big data metric aggregation on the online pre-aggregated multidimensional information to obtain a first big data aggregation result; performing pipeline aggregation on the first big data aggregation result to obtain a second big data aggregation result, thereby realizing heterogeneous multi-dimensional aggregation of big data for fast service.

[0059] The beneficial effects of the above technical solution are as follows: the present invention provides a big data aggregation method for fast service, which obtains fast service big data multidimensional information through data heterogeneity; pre-aggregates the big data multidimensional information through online processing nodes for online analysis and processing to obtain online pre-aggregated multidimensional information; performs big data bucket aggregation and big data metric aggregation on the online pre-aggregated multidimensional information to obtain a first aggregation result of big data; performs pipeline aggregation on the first aggregation result of big data to obtain a second aggregation result of big data, thereby realizing heterogeneous multidimensional aggregation of big data for fast service; classifies and splits the big data for fast service into multiple service type data according to service type; performs data heterogeneous processing on multiple service type data to obtain service type heterogeneous processing data; performs dimension upgrade on service type heterogeneous processing data to obtain fast service big data multidimensional information; distributes the big data multidimensional information according to the number of online processing nodes within a set range. To the set range of online processing nodes; establish a pre-aggregation online analysis and processing framework for the set range of online processing nodes; perform pre-aggregation online analysis and processing according to the aggregation online analysis and processing framework to obtain online pre-aggregated multi-dimensional information; centrally back up the online pre-aggregated multi-dimensional information to obtain pre-aggregated multi-dimensional information centralized backup data; perform big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data; perform big data metric aggregation on the bucket-type aggregated data to obtain the first aggregation result of big data; transmit the first aggregation result of big data to the data processing pipeline input end of the data pipeline aggregation; the data processing pipeline input end classifies the first aggregation result of big data and transmits it to each data processing pipeline for pipeline aggregation processing; through pipeline aggregation processing, the second aggregation result of big data is obtained, realizing heterogeneous multi-dimensional aggregation of big data for fast service; it is possible to perform data heterogeneity to obtain fast service big data multi-dimensional information, greatly improving the efficiency of service and the efficiency of big data aggregation.

[0060] In one embodiment, the S100 includes:

[0061] S101, classifying and splitting the big data for fast services into multiple service type data according to service types;

[0062] S102, performing data heterogeneity processing on data of multiple service types to obtain service type heterogeneity processed data;

[0063] S103, upgrading the dimension of the heterogeneous processing data of service types to obtain multi-dimensional information of fast service big data.

[0064] The working principle of the above technical solution is: classify and split the big data for fast service into multiple service type data according to the service type; perform data heterogeneous processing on the multiple service type data to obtain service type heterogeneous processing data; upgrade the dimension of the service type heterogeneous processing data to obtain multi-dimensional information of fast service big data.

[0065] The beneficial effects of the above technical solution are: classifying and splitting the big data for fast services into multiple service type data according to the service type; performing data heterogeneous processing on the multiple service type data to obtain service type heterogeneous processing data; performing dimensionality upgrade on the service type heterogeneous processing data to obtain multidimensional information of fast service big data; and being able to improve the efficiency of heterogeneous and multidimensional construction of data.

[0066] In one embodiment, the S200 includes:

[0067] S201, distributing the big data multidimensional information to the set range of online processing nodes according to the number of online processing nodes in the set range;

[0068] S202, establishing a pre-aggregated OLAP framework for OLAP nodes within a set range;

[0069] S203: Perform pre-aggregation OLAP according to the aggregate OLAP framework to obtain online pre-aggregation multi-dimensional information.

[0070] The working principle of the above technical solution is: distribute the big data multidimensional information to the set range of online processing nodes according to the number of online processing nodes in the set range; establish a pre-aggregation online analysis and processing framework for the set range of online processing nodes; perform pre-aggregation online analysis and processing according to the aggregated online analysis and processing framework to obtain online pre-aggregated multidimensional information.

[0071] The beneficial effects of the above technical solution are: distributing big data multidimensional information to the set range of online processing nodes according to the number of online processing nodes in the set range; establishing a pre-aggregation online analysis and processing framework for the set range of online processing nodes; performing pre-aggregation online analysis and processing according to the aggregated online analysis and processing framework to obtain online pre-aggregated multidimensional information.

[0072] In one embodiment, the S300 includes:

[0073] S301, centrally backing up the online pre-aggregated multi-dimensional information to obtain the pre-aggregated multi-dimensional information centralized backup data;

[0074] S302, performing big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data;

[0075] S303: Perform big data metric aggregation on the bucket-type aggregated data to obtain a first big data aggregation result.

[0076] The working principle of the above technical solution is: centrally back up the online pre-aggregated multidimensional information to obtain the pre-aggregated multi-dimensional information centralized backup data; perform big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data; perform big data metric aggregation on the bucket-type aggregated data to obtain the first aggregation result of big data.

[0077] The beneficial effects of the above technical solution are: online pre-aggregated multidimensional information is centrally backed up to obtain pre-aggregated multidimensional information centralized backup data; pre-aggregated multi-dimensional information centralized backup data is bucket-aggregated into big data to obtain bucket-aggregated data; bucket-aggregated data is metrically aggregated into big data to obtain the first aggregation result of big data.

[0078] In one embodiment, the S400 includes:

[0079] S401, transmitting the first aggregation result of the big data to the data processing pipeline input end of the data pipeline aggregation;

[0080] S402: The data processing pipeline input terminal classifies the first aggregation result of the big data and transmits it to each data processing pipeline for pipeline aggregation processing;

[0081] S403, through pipeline aggregation processing, obtain the second aggregation result of big data, and realize heterogeneous multi-dimensional aggregation of big data for fast service.

[0082] The working principle of the above technical solution is: the first aggregation result of big data is transmitted to the data processing pipeline input end of the data pipeline aggregation; the data processing pipeline input end classifies the first aggregation result of big data and transmits it to each data processing pipeline for pipeline aggregation processing; through pipeline aggregation processing, the second aggregation result of big data is obtained, realizing heterogeneous multi-dimensional aggregation of big data for fast service.

[0083] The beneficial effects of the above technical solution are: the first aggregation result of big data is transmitted to the data processing pipeline input end of the data pipeline aggregation; the data processing pipeline input end classifies the first aggregation result of big data and transmits it to each data processing pipeline for pipeline aggregation processing; through pipeline aggregation processing, the second aggregation result of big data is obtained, realizing heterogeneous multi-dimensional aggregation of big data for fast service.

[0084] The present invention provides a big data aggregation system for fast service, comprising:

[0085] The data heterogeneous multi-dimensional construction module converts the big data for fast service into multi-dimensional information through data heterogeneity;

[0086] The pre-aggregation online analysis module performs pre-aggregation online analysis on the multi-dimensional information of the big data through the online processing node to obtain online pre-aggregated multi-dimensional information;

[0087] The multi-dimensional information progressive aggregation module performs big data bucket aggregation and big data metric aggregation on the online pre-aggregated multi-dimensional information to obtain the first big data aggregation result;

[0088] The pipeline aggregation rapid response module performs pipeline aggregation on the first aggregation result of big data to obtain the second aggregation result of big data, realizing heterogeneous multi-dimensional aggregation of big data for rapid service.

[0089] The working principle of the above technical solution is: a big data aggregation system for fast service, which uses a data heterogeneous multidimensional construction module to obtain fast service big data multidimensional information through data heterogeneity; a pre-aggregation online analysis module, which pre-aggregates the big data multidimensional information through an online processing node and performs online analysis and processing to obtain online pre-aggregated multidimensional information; a multi-dimensional information progressive aggregation module, which performs big data bucket aggregation and big data metric aggregation on the online pre-aggregated multidimensional information to obtain a first aggregation result of big data; a pipeline aggregation rapid response module, which performs pipeline aggregation on the first aggregation result of big data to obtain a second aggregation result of big data, thereby realizing heterogeneous multi-dimensional aggregation of big data for fast service.

[0090] A big data aggregation system for fast service includes a data heterogeneous multidimensional construction module, which obtains fast service big data multidimensional information through data heterogeneity; a pre-aggregation online analysis module, which performs pre-aggregation online analysis and processing on the big data multidimensional information through online processing nodes to obtain online pre-aggregated multidimensional information; a multi-dimensional information progressive aggregation module, which performs big data bucket aggregation and big data metric aggregation on the online pre-aggregated multidimensional information to obtain a first big data aggregation result; a pipeline aggregation fast response module, which performs pipeline aggregation on the first big data aggregation result to obtain a second big data aggregation result, thereby realizing heterogeneous multidimensional aggregation of big data for fast service; classifying and splitting the big data for fast service into multiple service type data according to service type; performing data heterogeneous processing on the multiple service type data to obtain service type heterogeneous processing data; performing dimension upgrade on the service type heterogeneous processing data to obtain fast service big data multidimensional information; and linking the big data multidimensional information according to a set range. The number of machine processing nodes is distributed to the online processing nodes in the set range; a pre-aggregation online analytical processing framework is established for the online processing nodes in the set range; pre-aggregation online analytical processing is performed according to the aggregated online analytical processing framework to obtain online pre-aggregated multi-dimensional information; the online pre-aggregated multi-dimensional information is centrally backed up to obtain pre-aggregated multi-dimensional information centralized backup data; the pre-aggregated multi-dimensional information centralized backup data is subjected to big data bucket aggregation to obtain bucket-type aggregated data; the bucket-type aggregated data is subjected to big data metric aggregation to obtain a first aggregation result of big data; the first aggregation result of big data is transmitted to the data processing pipeline input end of the data pipeline aggregation; the data processing pipeline input end classifies the first aggregation result of big data and transmits it to each data processing pipeline for pipeline aggregation processing; through pipeline aggregation processing, a second aggregation result of big data is obtained, thereby realizing heterogeneous multi-dimensional aggregation of big data for fast service; data heterogeneity can be performed to obtain fast service big data multi-dimensional information, thereby greatly improving the efficiency of service and the efficiency of big data aggregation.

[0091] In one embodiment, the data heterogeneous multi-dimensional construction module includes:

[0092] The service type classification and splitting sub-module classifies and splits the big data for fast services into multiple service type data according to service types;

[0093] The multi-type data heterogeneous processing submodule performs data heterogeneity processing on multiple service type data to obtain service type heterogeneous processing data; the service type heterogeneous processing of multiple service types also includes: storing the benchmark service type table and the comparison service type table in the character separated value format to be verified in the distributed column-oriented database, the original service type record primary key as the primary key of the distributed column-oriented database table, the non-primary key attribute of the original service type record as a column of the distributed column-oriented database table, different columns belong to different column families, and the column-oriented storage of the distributed column-oriented database is used to improve the response performance when querying a certain column service type; storing the query index table of the verification rule verification field in the distributed column-oriented database In a column-oriented database, the check field serves as the primary key of the distributed column-oriented database query index table, and the original service type record primary key serves as the column name of the query index table. All primary keys belong to the same column family. This service type mode facilitates the addition, deletion, modification, and query of query index table records. The query index table of the service type record timestamp is stored in the distributed column-oriented database, and the service type record timestamp serves as the primary key of the distributed column-oriented database query index table, while the original service type record primary key serves as the column value of the query index table. When the query index table of the check rule check field is stored in the distributed column-oriented database, the query index table is also stored in the index file of the distributed file infrastructure.

[0094] The dimension upgrade multi-dimensional conversion sub-module upgrades the dimension of heterogeneous processing data of service types to obtain multi-dimensional information of fast service big data.

[0095] The working principle of the above technical solution is: the data heterogeneous multi-dimensional construction module includes: a service type classification and splitting submodule, which classifies and splits the big data for fast services into multiple service type data according to the service type; a multi-type data heterogeneous processing submodule, which performs data heterogeneous processing on multiple service type data to obtain service type heterogeneous processing data; the service type heterogeneous processing of multiple service types also includes: storing the benchmark service type table and the comparison service type table in the character separated value format to be verified in a distributed column-oriented database, the original service type record primary key as the primary key of the distributed column-oriented database table, and the non-primary key attribute of the original service type record as a column of the distributed column-oriented database table, different columns belong to different column families, and the column-oriented storage of the distributed column-oriented database is used to improve the response performance when querying a certain column service type; the verification rule verification The query index table of the field is stored in a distributed column-oriented database, the check field serves as the primary key of the query index table of the distributed column-oriented database, and the primary key of the original service type record serves as the column name of the query index table. All primary keys belong to the same column family. This service type mode facilitates the addition, deletion, modification and query of query index table records; the query index table of the service type record timestamp is stored in a distributed column-oriented database, the service type record timestamp serves as the primary key of the query index table of the distributed column-oriented database, and the original service type record primary key is stored as the column value of the query index table; when the query index table of the check rule check field is stored in the distributed column-oriented database, the query index table is also stored in the index file of the distributed file infrastructure; the dimension upgrade multidimensional conversion submodule upgrades the dimension of the service type heterogeneous processing data to obtain fast service big data multidimensional information.

[0096] The beneficial effects of the above technical solution are: through the service type classification and splitting sub-module, the big data for fast services is classified and split into multiple service type data according to the service type; the multi-type data heterogeneous processing sub-module performs data heterogeneous processing on multiple service type data to obtain service type heterogeneous processing data; the service type heterogeneous processing of multiple service types also includes: storing the benchmark service type table and the comparison service type table in the character separated value format to be verified in the distributed column-oriented database, the original service type record primary key as the primary key of the distributed column-oriented database table, the non-primary key attribute of the original service type record as a column of the distributed column-oriented database table, different columns belong to different column families, and the column-oriented storage of the distributed column-oriented database is used to improve the response performance when querying a certain column service type; storing the query index table of the verification rule verification field in the distributed column-oriented database In a column-oriented database, the check field serves as the primary key of a distributed column-oriented database query index table, and the original service type record primary key serves as the column name of the query index table. All primary keys belong to the same column family. This service type mode facilitates the addition, deletion, modification, and query of query index table records. The query index table of the service type record timestamp is stored in a distributed column-oriented database, and the service type record timestamp serves as the primary key of the distributed column-oriented database query index table. The original service type record primary key serves as the column value of the query index table. When the query index table of the check rule check field is stored in a distributed column-oriented database, the query index table is also stored in an index file of a distributed file infrastructure. The dimension upgrade multidimensional conversion submodule upgrades the dimension of the service type heterogeneous processing data to obtain multidimensional information of fast service big data. This can improve the efficiency of heterogeneous multidimensional construction of data.

[0097] In one embodiment, the pre-aggregation online analysis module includes:

[0098] The online processing node distribution submodule distributes the multi-dimensional information of big data to the online processing nodes within the set range according to the number of online processing nodes within the set range;

[0099] The OLAP framework submodule establishes a pre-aggregated OLAP framework for a set range of OLAP nodes;

[0100] The pre-aggregation online analytical processing submodule performs pre-aggregation online analytical processing according to the aggregation online analytical processing framework to obtain online pre-aggregated multi-dimensional information; the pre-aggregation online analytical processing includes: pre-aggregating the source data into pre-aggregated table data represented in the form of a mesh data structure according to the spatial dimension, wherein the level of the mesh data structure represents the spatial hierarchy of the pre-aggregation, and the content of the mesh data structure node includes the spatial area corresponding to the node and the lightweight pre-aggregated data obtained by pre-aggregating the pre-aggregation dimension selected from the data attribute; when generating a report, determining the selected area selected by the user in space, and pre-aggregating the lightweight pre-aggregated data of each mesh data structure node belonging to the selected area according to the query dimension selected by the user to obtain the query result; pre-aggregating the source data into mesh area pre-aggregated table data according to the mesh area division according to the pre-aggregation dimension selected from the data attribute; when generating a report, pre-aggregating the mesh area pre-aggregated table data of the mesh area selected by the user according to the query dimension to obtain the query result.

[0101] The working principle of the above technical solution is as follows: the pre-aggregation online analysis module includes: an online processing node distribution submodule, which distributes the big data multidimensional information to the set range of online processing nodes according to the number of online processing nodes in the set range;

[0102] The OLAP framework submodule establishes a pre-aggregation OLAP framework for OLAP nodes within a set range; the pre-aggregation OLAP submodule performs pre-aggregation OLAP according to the aggregation OLAP framework to obtain online pre-aggregated multi-dimensional information; performing pre-aggregation OLAP includes: pre-aggregating source data according to spatial dimensions into pre-aggregated table data represented in the form of a mesh data structure, wherein the levels of the mesh data structure represent the spatial hierarchy of pre-aggregation, and the content of the mesh data structure nodes includes the spatial area corresponding to the node and lightweight pre-aggregated data obtained by pre-aggregating the pre-aggregation dimension selected from the data attributes; determining a selected area selected by the user in space when generating a report, and pre-aggregating the lightweight pre-aggregated data of each mesh data structure node within the selected area according to the query dimension selected by the user to obtain a query result; pre-aggregating the source data according to the mesh area division according to the pre-aggregation dimension selected from the data attributes into mesh area pre-aggregated table data; and pre-aggregating the mesh area pre-aggregated table data of the mesh area selected by the user according to the query dimension when generating a report to obtain a query result.

[0103] The beneficial effects of the above technical solution are as follows: through the online processing node distribution submodule, the big data multidimensional information is distributed to the set range of online processing nodes according to the number of online processing nodes in the set range; the online analysis processing framework submodule establishes a pre-aggregation online analysis processing framework for the set range of online processing nodes; the pre-aggregation online analysis processing submodule performs pre-aggregation online analysis processing according to the aggregation online analysis processing framework to obtain online pre-aggregated multidimensional information; the pre-aggregation online analysis processing includes: pre-aggregating the source data into pre-aggregation table data represented in the form of a mesh data structure according to the spatial dimension, wherein the level of the mesh data structure represents the spatial hierarchy of the pre-aggregation, and the mesh data structure represents the spatial hierarchy of the pre-aggregation. According to the content of the structure node, the spatial area corresponding to the node and the pre-aggregation dimension selected from the data attribute are pre-aggregated to obtain lightweight pre-aggregated data; when generating a report, the selection area selected by the user in space is determined, and according to the query dimension selected by the user, the lightweight pre-aggregated data of each mesh data structure node belonging to the selection area are pre-aggregated to obtain the query result; according to the pre-aggregation dimension selected from the data attribute, the source data are pre-aggregated into mesh area pre-aggregated table data according to the mesh area; when generating a report, the mesh area pre-aggregated table data of the mesh area selected by the user are pre-aggregated according to the query dimension to obtain the query result; and the automatic data processing mode can be further converted.

[0104] In one embodiment, the multi-dimensional information progressive aggregation module includes:

[0105] The multi-dimensional information centralized backup submodule performs centralized backup of the online pre-aggregated multi-dimensional information to obtain the pre-aggregated multi-dimensional information centralized backup data;

[0106] The big data bucket aggregation submodule performs big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data;

[0107] The big data metric aggregation submodule performs big data metric aggregation on the bucket-type aggregated data to obtain the first big data aggregation result.

[0108] The working principle of the above technical solution is: the multidimensional information progressive aggregation module includes: a multidimensional information centralized backup submodule, which performs centralized back-up of the online pre-aggregated multidimensional information to obtain pre-aggregated multidimensional information centralized backup data; a big data bucket-type aggregation submodule, which performs big data bucket-type aggregation on the pre-aggregated multidimensional information centralized backup data to obtain bucket-type aggregated data; a big data measurement aggregation submodule, which performs big data measurement aggregation on the bucket-type aggregated data to obtain the first big data aggregation result.

[0109] The beneficial effects of the above technical solution are: through the multi-dimensional information centralized backup sub-module, the online pre-aggregated multi-dimensional information is centrally backed up to obtain the pre-aggregated multi-dimensional information centralized backup data; the big data bucket aggregation sub-module performs big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data; the big data measurement aggregation sub-module performs big data measurement aggregation on the bucket-type aggregated data to obtain the first big data aggregation result.

[0110] In one embodiment, the pipeline aggregation rapid response module includes:

[0111] The aggregation transmission pipeline input submodule transmits the first aggregation result of the big data to the data processing pipeline input end of the data pipeline aggregation;

[0112] The classification is input into the pipeline aggregation submodule. The data processing pipeline input end classifies the first aggregation result of the big data and inputs it into each data processing pipeline for pipeline aggregation processing;

[0113] The fast service aggregation implementation submodule obtains the second aggregation result of big data through pipeline aggregation processing, realizing heterogeneous multi-dimensional aggregation of big data for fast services.

[0114] The working principle of the above technical solution is: the pipeline aggregation rapid response module includes: an aggregation transmission pipeline input submodule, which transmits the first aggregation result of big data to the data processing pipeline input end of the data pipeline aggregation; a classification input pipeline aggregation submodule, the data processing pipeline input end classifies the first aggregation result of big data and transmits it to each data processing pipeline for pipeline aggregation processing; a rapid service aggregation implementation submodule, which obtains the second aggregation result of big data through pipeline aggregation processing, and realizes heterogeneous multi-dimensional aggregation of big data for rapid service.

[0115] The beneficial effects of the above technical solution are: the pipeline aggregation rapid response module includes: an aggregation transmission pipeline input submodule, which transmits the first aggregation result of big data to the data processing pipeline input end of the data pipeline aggregation; a classification input pipeline aggregation submodule, the data processing pipeline input end classifies the first aggregation result of big data and inputs it into each data processing pipeline for pipeline aggregation processing; a rapid service aggregation implementation submodule, which obtains the second aggregation result of big data through pipeline aggregation processing, and realizes heterogeneous multi-dimensional aggregation of big data for rapid service; greatly improving the efficiency of service and the efficiency of big data aggregation.

[0116] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A big data aggregation method for fast service, characterized in that: include: S100, which uses data heterogeneity to obtain multi-dimensional information of fast service big data; S200, performing pre-aggregation online analysis processing on the big data multi-dimensional information through the online processing node to obtain online pre-aggregated multi-dimensional information; S300, performing big data bucket aggregation and big data metric aggregation on the online pre-aggregated multi-dimensional information to obtain a first big data aggregation result; S400: Pipeline aggregation is performed on the first aggregation result of the big data to obtain a second aggregation result of the big data, thereby realizing heterogeneous multi-dimensional aggregation of big data for fast service. The S100 includes: S101, classifying and splitting the big data for fast services into multiple service type data according to service types; S102, performing data heterogeneity processing on multiple service type data to obtain service type heterogeneity processing data; performing service type heterogeneity processing on multiple service types also includes: storing a benchmark service type table and a comparison service type table in a character-separated value format to be verified in a distributed column-oriented database, using the primary key of the original service type record as the primary key of the distributed column-oriented database table, and using the non-primary key attribute of the original service type record as a column of the distributed column-oriented database table, different columns belong to different column families, and using the column-oriented storage of the distributed column-oriented database to improve the response performance when querying a certain column service type; storing the query index table of the verification rule verification field in the distributed column-oriented database In the database, the check field is used as the primary key of the distributed column-oriented database query index table, and the original service type record primary key is used as the column name of the query index table. All primary keys belong to the same column family. This service type mode facilitates the addition, deletion, modification and query of query index table records. The query index table of the service type record timestamp is stored in the distributed column-oriented database. The service type record timestamp is used as the primary key of the distributed column-oriented database query index table, and the original service type record primary key is stored as the column value of the query index table. When the query index table of the check rule check field is stored in the distributed column-oriented database, the query index table is also stored in the index file of the distributed file infrastructure. S103, upgrading the dimension of the heterogeneous processing data of service types to obtain multi-dimensional information of fast service big data.

2. A method for aggregating big data for fast service according to claim 1, characterized in that: The S200 includes: S201, distributing the big data multidimensional information to the set range of online processing nodes according to the number of online processing nodes in the set range; S202, establishing a pre-aggregated OLAP framework for OLAP nodes within a set range; S203: Perform pre-aggregation OLAP according to the aggregate OLAP framework to obtain online pre-aggregation multi-dimensional information.

3. The method for aggregating big data for fast service according to claim 1, characterized in that: The S300 includes: S301, centrally backing up the online pre-aggregated multi-dimensional information to obtain the pre-aggregated multi-dimensional information centralized backup data; S302, performing big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data; S303: Perform big data metric aggregation on the bucket-type aggregated data to obtain a first big data aggregation result.

4. The method for aggregating big data for fast service according to claim 1, characterized in that: The S400 includes: S401, transmitting the first aggregation result of the big data to the data processing pipeline input end of the data pipeline aggregation; S402: The data processing pipeline input terminal classifies the first aggregation result of the big data and transmits it to each data processing pipeline for pipeline aggregation processing; S403, through pipeline aggregation processing, obtain the second aggregation result of big data, and realize heterogeneous multi-dimensional aggregation of big data for fast service.

5. A big data aggregation system for fast service, characterized by: include: The data heterogeneous multi-dimensional construction module converts the big data for fast service into multi-dimensional information through data heterogeneity; The pre-aggregation online analysis module performs pre-aggregation online analysis on the multi-dimensional information of the big data through the online processing node to obtain online pre-aggregated multi-dimensional information; The multi-dimensional information progressive aggregation module performs big data bucket aggregation and big data metric aggregation on the online pre-aggregated multi-dimensional information to obtain the first big data aggregation result; The pipeline aggregation rapid response module performs pipeline aggregation on the first aggregation result of big data to obtain the second aggregation result of big data, realizing heterogeneous multi-dimensional aggregation of big data for rapid service; The data heterogeneous multi-dimensional building module includes: The service type classification and splitting sub-module classifies and splits the big data for fast services into multiple service type data according to the service type; The multi-type data heterogeneous processing submodule performs data heterogeneity processing on multiple service type data to obtain service type heterogeneous processing data; the service type heterogeneous processing of multiple service types also includes: storing the benchmark service type table and the comparison service type table in the character separated value format to be verified in the distributed column-oriented database, the original service type record primary key as the primary key of the distributed column-oriented database table, the non-primary key attribute of the original service type record as a column of the distributed column-oriented database table, different columns belong to different column families, and the column-oriented storage of the distributed column-oriented database is used to improve the response performance when querying a certain column service type; storing the query index table of the verification rule verification field in the distributed column-oriented database In a column-oriented database, the check field serves as the primary key of the distributed column-oriented database query index table, and the original service type record primary key serves as the column name of the query index table. All primary keys belong to the same column family. This service type mode facilitates the addition, deletion, modification, and query of query index table records. The query index table of the service type record timestamp is stored in the distributed column-oriented database, and the service type record timestamp serves as the primary key of the distributed column-oriented database query index table, while the original service type record primary key serves as the column value of the query index table. When the query index table of the check rule check field is stored in the distributed column-oriented database, the query index table is also stored in the index file of the distributed file infrastructure. The dimension upgrade multi-dimensional conversion sub-module upgrades the dimension of heterogeneous processing data of service types to obtain multi-dimensional information of fast service big data.

6. The fast service-oriented big data aggregation system according to claim 5, characterized in that: The pre-aggregation online analysis module includes: The online processing node distribution submodule distributes the multi-dimensional information of big data to the online processing nodes within the set range according to the number of online processing nodes within the set range; The OLAP framework submodule establishes a pre-aggregated OLAP framework for a set range of OLAP nodes; The pre-aggregation online analytical processing submodule performs pre-aggregation online analytical processing according to the aggregation online analytical processing framework to obtain online pre-aggregated multi-dimensional information; the pre-aggregation online analytical processing includes: pre-aggregating the source data into pre-aggregated table data represented in the form of a mesh data structure according to the spatial dimension, wherein the level of the mesh data structure represents the spatial hierarchy of the pre-aggregation, and the content of the mesh data structure node includes the spatial area corresponding to the node and the lightweight pre-aggregated data obtained by pre-aggregating the pre-aggregation dimension selected from the data attribute; when generating a report, determining the selected area selected by the user in space, and pre-aggregating the lightweight pre-aggregated data of each mesh data structure node belonging to the selected area according to the query dimension selected by the user to obtain the query result; pre-aggregating the source data into mesh area pre-aggregated table data according to the mesh area division according to the pre-aggregation dimension selected from the data attribute; when generating a report, pre-aggregating the mesh area pre-aggregated table data of the mesh area selected by the user according to the query dimension to obtain the query result.

7. The fast service-oriented big data aggregation system according to claim 5, characterized in that: The multi-dimensional information progressive aggregation module includes: The multi-dimensional information centralized backup submodule performs centralized backup of the online pre-aggregated multi-dimensional information to obtain the pre-aggregated multi-dimensional information centralized backup data; The big data bucket aggregation submodule performs big data bucket aggregation on the pre-aggregated multi-dimensional information centralized backup data to obtain bucket-type aggregated data; The big data metric aggregation submodule performs big data metric aggregation on the bucket-type aggregated data to obtain the first big data aggregation result.

8. The fast service-oriented big data aggregation system according to claim 5, characterized in that: The pipeline aggregation rapid response module includes: The aggregation transmission pipeline input submodule transmits the first aggregation result of the big data to the data processing pipeline input end of the data pipeline aggregation; The classification is input into the pipeline aggregation submodule. The data processing pipeline input end classifies the first aggregation result of the big data and inputs it into each data processing pipeline for pipeline aggregation processing; The fast service aggregation implementation submodule obtains the second aggregation result of big data through pipeline aggregation processing, realizing heterogeneous multi-dimensional aggregation of big data for fast services.

Citation Information

Patent Citations

  • Data association analysis method and device

    CN112434022A

  • Multi-source heterogeneous big data fusion system based on large-scale popularization of traditional Chinese medicine knowledge

    CN113111244A