Data analysis methods, models, equipment and storage media

Through the tower multi-layer data model and message routing mechanism, a multi-layer fully connected network is built, solving the problems of insufficient complexity and flexibility of Hadoop's big data solution, and achieving efficient massive data processing and real-time analysis capabilities.

CN113139007BActive Publication Date: 2025-05-13CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010060047.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-01-19
Publication Date
2025-05-13
Estimated Expiration
2040-01-19

AI Technical Summary

Technical Problem

The existing Hadoop big data solutions are complex and have low flexibility, making them difficult to meet the needs of OLAP operation such as flexible combination of query, drilling, and slicing, especially when processing massive data and real-time analysis of scenarios.

Method used

The tower-type multi-layer data model is adopted to aggregate the parsers, counters and aggregators layer by layer, and data analysis and statistical analysis are realized through the message routing mechanism to form a multi-layer fully connected network to realize the flow of data flow, control flow and state flow.

Benefits of technology

It improves data calculation and query efficiency, increases the flexibility of data analysis, can efficiently process massive data and support real-time analysis requirements, and improves performance by 3 to 5 times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113139007B_ABST
    Figure CN113139007B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a data analysis method, model, device and storage medium, wherein the method comprises: each of the M parsers obtains the data to be analyzed from different devices in the data pool, wherein the data to be analyzed includes a device identification; each of the parsers parses the data to be analyzed, assembles the data to be analyzed according to statistical requirements, and obtains an assembled message; each of the routers sends each of the assembled messages to a corresponding counter according to the device identification; each of the N counters aggregates the assembled messages according to a specified calculation logic to obtain first statistical data; the aggregator aggregates the first statistical data outputted from the N counters into second statistical data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of big data technology, and are related to but not limited to a data analysis method, model, device and storage medium. Background Art

[0002] Currently, the most commonly used big data technology solution in the industry is distributed computing (Hadoop). Hadoop big data solution is an ecosystem. In actual use, there are the following problems:

[0003] Problem 1: Complex solution: Hadoop big data solution is very complex, and practice has shown that this complexity has a very large adverse impact on the learning, installation, deployment, development, use, and operation and maintenance of big data solutions.

[0004] Problem 2: Low flexibility: Currently, the application scenarios of data analysis are deterministic data indicator analysis (a scheduled task starts at 0:00 every day) and real-time online analytical processing (OLAP). In the OLAP scenario, operations such as flexible combination of queries, drilling, and slicing in any dimension are required. For abnormal data, the data source needs to be traced back level by level to verify the correctness of the data. Hadoop solutions cannot meet the above requirements well. Summary of the invention

[0005] In view of this, embodiments of the present application provide a data analysis method, model, device and storage medium.

[0006] The technical solution of the embodiment of the present application is implemented as follows:

[0007] In a first aspect, an embodiment of the present application provides a data analysis method, which is applied to a data analysis model, wherein the data analysis model includes M parsers, M routers corresponding to the M parsers, N counters and 1 aggregator, where M and N are integers greater than or equal to 2, and the method includes: each of the M parsers obtains data to be analyzed from different devices in a data pool, wherein the data to be analyzed includes a device identifier; each of the parsers parses the data to be analyzed, and assembles the data to be analyzed according to statistical requirements to obtain an assembled message; each of the routers sends each assembled message to the corresponding counter according to the device identifier; each of the N counters aggregates the assembled message according to a specified calculation logic to obtain first statistical data; and the aggregator aggregates the first statistical data outputted by the N counters into second statistical data.

[0008] In a second aspect, an embodiment of the present application provides a data analysis model, comprising M parsers, M routers corresponding to the M parsers, N counters and 1 aggregator, where M and N are integers greater than or equal to 2; wherein each of the M parsers is used to obtain data to be analyzed from different devices in a data pool, wherein the data to be analyzed includes a device identifier; each of the parsers is used to parse the data to be analyzed, and assemble the data to be analyzed according to statistical requirements to obtain an assembled message; each of the routers is used to send each assembled message to the corresponding counter according to the device identifier; each of the N counters is used to aggregate the assembled message according to a specified calculation logic to obtain a first statistical data; and the aggregator is used to aggregate the first statistical data outputted by the N counters into a second statistical data.

[0009] In a third aspect, an embodiment of the present application provides a data analysis device, comprising: a memory and a processor, wherein the processor implements the steps in the above method by calling a data analysis model to execute the program.

[0010] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above method when executed by a processor.

[0011] In the embodiments of the present application, 1) a tower-type multi-layer data model is adopted to bring together parsers, counters and aggregators, which are jointly responsible for the statistical analysis function, namely "raw data - user granularity summary data - full-dimensional statistical data - combined dimensional data", and each layer of data calculation is generated based on the directly adjacent next level of data. In this way, compared with the traditional data model "raw data - combined dimensional data", this model not only improves the efficiency of data calculation and query, but also increases the flexibility of data analysis. 2) In the tower-type multi-layer data model, multiple parsers form a parsing layer, and multiple counters form a counting layer. The tower-type multi-layer data model also includes a router. The message routing mechanism realizes the message transmission between the data parsing layer and the counting layer through a virtual router. The purpose is to ensure that multiple data of the same device are processed at the same counter computing node, and to ensure the atomicity of the data statistical results, so that the statistical results of multiple counters can be accumulated and aggregated at the result aggregation layer. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1A This is a schematic diagram of the composition structure of the data analysis model of the embodiment of the present application;

[0013] Figure 1B A schematic diagram of an implementation flow of a data analysis method provided in an embodiment of the present application;

[0014] Figure 2AA schematic diagram of the implementation flow of another data analysis method provided in the application embodiment;

[0015] Figure 2B A schematic diagram of a process flow of implementing another data analysis method provided in an embodiment of the application;

[0016] Figure 3 A schematic diagram of a process flow of implementing another data analysis method provided in an embodiment of the application;

[0017] Figure 4 A flowchart of another data analysis method provided in an embodiment of the present application;

[0018] Figure 5A A schematic diagram of a composition structure of a data analysis model according to an embodiment of the present application;

[0019] Figure 5B This is another schematic diagram of the composition structure of the data analysis model of the embodiment of the present application;

[0020] Figure 6 A schematic diagram of a hardware entity of a data analysis device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0022] It should be understood that some embodiments described herein are only used to explain the technical solution of the present application, and are not used to limit the technical scope of the present application.

[0023] This embodiment first provides a data analysis model. Figure 1A Schematic diagram of the composition structure of the data analysis model of the embodiment of the present application, such as Figure 1A As shown, the data analysis model includes: a manager 101 for file distribution, M parsers 102 for file parsing, M routers 103 corresponding to the M parsers, N counters 104 for message counting, an aggregator 105 for result aggregation, and a logger 106 for recording manager, parser, router, counter and aggregator logs, where M and N are integers greater than or equal to 2. Wherein:

[0024] The manager 101 is connected to the M lower-level parsers 102 to distribute data processing tasks to multiple parsers; the parser 102 is responsible for decoding, cleaning and conversion of data files, and each parser is connected to the manager and the corresponding router, as well as the logger; the router 103 is responsible for forwarding each message processed by the parser, and each router is connected to the corresponding parser, all counters and loggers; the counter 104 is connected to all routers, aggregators and loggers, and the counter can be set up in multiple layers, and each layer is counted according to different dimensions; the parser 102 routes the data of the same batch of devices to a counter 104 through the router 103 for statistical analysis; the aggregator 105 is connected to all counters and loggers and is responsible for global data statistics and data storage. The logger 106 is connected to all computing units for logging.

[0025] The technical architecture of the overall data analysis model uses a data model builder (studybuilder) to stack six types of atomic computing nodes layer by layer in the manner of building a model with building blocks. The communication between computing nodes in adjacent layers is achieved through message queues to achieve full mesh connection - similar to a deep learning neural network model, thereby realizing the flow of data flow, control flow, log flow and state flow between computing nodes in each layer. Any data analysis needs can be mapped to a data analysis model built by the data model builder and the specific computing logic of different computing nodes in the model. Since the computing nodes are divided into sufficiently fine parts, each computing node only needs to complete extremely simple tasks. However, through the reasonable division of labor and cooperation of computing nodes at each layer in the data analysis model, various very complex data analysis tasks can be completed efficiently and flexibly.

[0026] The process of building a data analysis model through the data model builder is as follows:

[0027] First, a pool of data files to be processed is generated according to the data processing range specified by the user (assuming there are 10,000 data files), a manager service is generated and waits for the active task application of the parser; the manager and the parser are in a service and client relationship, a manager connects all the subordinate parsers, and the manager is used to distribute data processing tasks to multiple parsers.

[0028] Then, according to the number of parser instances specified by the user, assuming it is M, M parser instances and M routers are generated respectively. Each router corresponds to a parser. The parser is responsible for decoding, cleaning and conversion of data files, and the router is responsible for forwarding each message processed by the parser.

[0029] Secondly, a specified number (assuming N) of counter instances are generated according to the number of counter instances specified by the user; each counter is connected to all the upper-level parsers and their routers, and the parser routes the data of the same batch of devices to a counter through the router for statistical analysis.

[0030] Finally, an aggregator and a logger are generated; the aggregator is responsible for global data statistics and data storage, and the logger is responsible for logging the entire data analysis model; during the entire model operation process, instances of each category send logs directly to the logger.

[0031] based on Figure 1A , a data processing method provided in an embodiment of the present application, referring to Figure 1B As shown, including:

[0032] Step S101: Each of the M parsers obtains data to be analyzed from different devices in a data pool, wherein the data to be analyzed includes a device identifier;

[0033] Step S102: Each of the parsers parses the data to be analyzed, assembles the data to be analyzed according to statistical requirements, and obtains assembled messages;

[0034] After receiving the data to be analyzed from the manager, the parser reads and decodes the data, cleans and converts the original data line by line, and generates assembled data according to business needs and sends it to the corresponding router. Data cleaning is generally a standard process for data analysis, and its purpose is to remove invalid and erroneous data; data conversion is to determine which fields need to be converted in content or format based on the logic of this data analysis. For example, an example of content conversion is: extracting the year of birth from the ID card number and counting by year. The user's data analysis needs can be mapped to the logic of data analysis. For example, this analysis needs to extract 20 fields from all 100 fields, so the parser only extracts these 20 fields for assembly and sends them to the next level counter through the router. The assembled data uses the simplest special characters to separate each field into a string, which can achieve the effect of reducing the transmission volume.

[0035] Step S103: Each of the routers sends each of the assembly messages to a corresponding counter according to the device identifier;

[0036] The router extracts the device ID from the received assembly message and routes the message to the corresponding counter Ni according to the following algorithm, where the value of i ranges from 0 to n-1. The algorithm is: i = hash (device number) % n; where hash (device number) is the hash value of the device ID, % is the modulus operation, and n is the number of counter instances;

[0037] The purpose of this step is to ensure that multiple messages from the same device are counted by the same counter.

[0038] Step S104: Each of the N counters aggregates the assembled messages according to a specified calculation logic to obtain first statistical data;

[0039] The counter generates summary data based on user granularity according to user needs, and updates the summary data after receiving the assembly message, that is, the data of the same device is aggregated according to the specified calculation logic.

[0040] Step S105: The aggregator aggregates the first statistical data outputted by the N counters into second statistical data.

[0041] Aggregate statistics from n counters into global summary data.

[0042] In the embodiments of the present application, 1) a tower-type multi-layer data model is adopted to bring together parsers, counters and aggregators, which are jointly responsible for the statistical analysis function, namely "raw data-user granularity summary data-full-dimensional statistical data-combined dimensional data", and each layer of data calculation is generated based on the directly adjacent next level of data. In this way, compared with the traditional data model "raw data-combined dimensional data", this model not only improves the efficiency of data calculation and query, but also increases the flexibility of data analysis. 2) In the tower-type multi-layer data model, multiple parsers constitute the parsing layer, and multiple counters constitute the technical layer. The tower-type multi-layer data model also includes a router. The message routing mechanism realizes the message transmission between the data parsing layer and the counting layer through a virtual router. The purpose is to ensure that multiple data of the same device are processed at the same counter computing node, and to ensure the atomicity of the data statistical results, so that the statistical results of multiple counters can be accumulated and aggregated at the result aggregation layer.

[0043] The embodiment of the present application provides another data analysis method, which is applied to Figure 1A The data analysis model shown is Figure 2A As shown, the method includes:

[0044] Step S201: When each of the parsers determines that it is idle, it sends an idle notification message to the manager;

[0045] If the parser is idle, it will actively tell the manager, and the manager will take a data name from the waiting data pool and send it to the parser for processing. This active claim mechanism of everyone doing their best ensures maximum processing efficiency.

[0046] Step S202: When the manager checks that the data pool is not empty, it obtains the data to be analyzed from different devices from the data pool, and sends the data to be analyzed to the parser;

[0047] After the model is started, the parser sends a "parser idle notification" message to the manager. After receiving the message, the manager checks whether there is any unprocessed data in the data pool to be processed. If so, the data is sent to a specific parser through a message.

[0048] Step S203: Each of the parsers parses the data to be analyzed, assembles the data to be analyzed according to statistical requirements, and obtains assembled messages;

[0049] Step S204: Each of the routers sends each of the assembly messages to a corresponding counter according to the device identifier;

[0050] Step S205: Each of the N counters aggregates the assembled messages according to a specified calculation logic to obtain first statistical data;

[0051] Step S206: The aggregator aggregates the first statistical data outputted by the N counters into second statistical data.

[0052] Step S207: The logger is responsible for recording the logs output by the manager, each of the parsers, each of the routers, each of the counters and the aggregator.

[0053] In the embodiment of the present application, the manager and the data parser adopt a mechanism of actively claiming tasks from bottom to top. Each parser computing node actively sends an application to the manager computing node when it is idle to obtain a data file for processing. This data distribution mechanism not only ensures the simplicity of the mechanism but also ensures the optimal overall performance. At the same time, it is also compatible with abnormal situations that may occur in the parser computing node. The message transmission between the data parsing layer and the counting layer is realized through a virtual router. The purpose is to ensure that multiple data of the same device are processed in the same counter computing node, to ensure the atomicity of the data statistical results, so that the statistical results of multiple counters can be accumulated and aggregated in the result aggregation layer.

[0054] The embodiment of the present application provides another data analysis method, which is applied to Figure 1A The data analysis model shown is Figure 2B As shown, the method includes:

[0055] Step S211, when each of the parsers determines that it is idle, it sends an idle notification message to the manager;

[0056] Step S212: When the manager finds that the data pool is empty, it sends an end message to the parser;

[0057] If the manager finds that there is no data file in the data pool, it sends an end message to the parser. Figure 2A The embodiment shown.

[0058] Step S213: When each of the resolvers receives the end message, it sends a resolution end message to the corresponding router;

[0059] Step S214: When each of the routers receives the resolution completion message, it sends a routing completion message to the N counters;

[0060] Step S215: When the M routers complete sending the routing end message, the counter aggregates the assembled messages according to the specified calculation logic to obtain first statistical data;

[0061] Step S216: The aggregator aggregates the first statistical data outputted from the N counters into second statistical data;

[0062] Step S217, the logger is responsible for recording the logs output by the manager, each of the parsers, each of the routers, each of the counters and the aggregator.

[0063] In the embodiment of the present application, when the M routers complete sending the routing end message, the counter aggregates the assembled message according to the specified calculation logic to obtain the first statistical data. The aggregator aggregates the first statistical data output from the N counters into the second statistical data. The data statistics can be completed very simply and efficiently.

[0064] The embodiment of the present application provides another data analysis method, which is applied to Figure 1A The data analysis model shown in the figure further includes a logger. When any of the above embodiments is completed, Figure 3 As shown, the method includes:

[0065] Step S301, when each of the parsers parses the data to be analyzed, it records first status data; wherein: the first status data includes at least one of the following: the number of files processed, the number of lines, the error distribution and the parser status of itself;

[0066] The parser records various status data during the parsing process (such as the distribution of various exceptions, the number of files processed, the number of lines, the duration, etc.); each parser instance will record the status data of the parser, such as how many files were parsed, how many lines of data were abnormal, how long it took, etc.

[0067] Step S302: Each of the parsers sends the first status data to the aggregator;

[0068] Step S303: When the M parsers finish sending the first state data to the aggregator, the aggregator aggregates the M first state data into second state data;

[0069] Step S304: output the second state data.

[0070] Step S305: Each of the counters records third status data when aggregating the first statistical data; wherein the second status data includes at least one of the following: recording the number of files processed, the number of lines, the error distribution and the status of the counter;

[0071] The counter records various status data during the counting process (such as the distribution of various exceptions, the number of processed files, the number of lines, the duration, etc.);

[0072] Step S306: Each of the counters sends the third status data to the aggregator.

[0073] Step S307: When the N counters finish sending the third state data to the aggregator, the aggregator aggregates the N third state data into fourth state data;

[0074] Step S308: output the fourth state data.

[0075] Step S309: The aggregator aggregates the second state data and the fourth state data to obtain fifth state data.

[0076] In the embodiment of the present application, a standard data communication interface for different types of computing units is constructed through a message queue, and different types of computing units are organized into a multi-layer fully connected network through the interface to realize the flow of status messages at different computing unit levels, and finally the global status messages are summarized in the aggregator. The black box effect of other big data processing frameworks is avoided. In addition to the final statistical data results, all computing nodes and intermediate status data of the computing level, such as error data distribution, processing time, etc., are summarized to facilitate subsequent other analyses, such as data accuracy backtracking analysis, performance optimization, and data quality analysis.

[0077] At present, the commonly used big data technology solutions in the industry include distributed computing (Hadoop), Spark, GreenPlum, etc. Most of these technologies appeared more than 10 years ago. The background at that time was that the Internet industry had accumulated massive amounts of information and data after nearly 10 years of slow and fast development. The traditional host-based scale-up computing model was increasingly unable to cope with these massive amounts of data. In addition to being expensive, the traditional computing model Symmetrical Multi-Processing (SMP) architecture was difficult to expand and technically difficult to meet the performance index requirements of massive data computing. Therefore, a revolution in computing methods was urgently needed. In response to this demand, new solutions supporting distributed storage and distributed computing theory were proposed one after another. Although these new solutions have their own characteristics, their common feature is that they use horizontal expansion (scale-out) between different hosts instead of upward expansion (scale-up) of the same host, that is, using cheaper x86 servers to form a fault-tolerant distributed computing cluster, and using distributed computing (MapReduce) or similar methods to decompose data computing tasks into smaller-granularity subtasks to achieve distributed statistics and aggregation on parallel computing nodes, thereby meeting users' requirements for massive data analysis and performance indicators. Two of the better representatives are Hadoop and greenplum:

[0078] Hadoop is an open source distributed computing platform of the Apache Foundation. It is centered around the distributed storage file system (HDFS) and distributed computing (MapReduce) algorithm, and provides users with a distributed infrastructure with transparent underlying system details. Hadoop's distributed storage and distributed computing are completed in cluster nodes. Cluster nodes mainly include namenodes and datanodes. Hadoop can be expanded linearly to process larger data sets by adding cluster nodes, and the maximum cluster size can reach tens of thousands. Hadoop can develop applications that process structured or unstructured data. In the Internet field, 95% of data is unstructured, so Hadoop is currently very popular in the Internet field.

[0079] Greenplum is an open source distributed database storage solution that focuses on data warehouses and business intelligence. It is developed based on the popular PostgreSQL. Greenplum's architecture uses massively parallel processing (MPP). The computing nodes are also called SMP nodes. Each SMP node has its own operating system, database, etc. The information interaction between nodes is realized through the node interconnection network. It has very good performance in data storage, high concurrency, high availability, linear expansion, response speed, ease of use and cost performance. However, the cluster scale is rarely thousands, usually dozens or hundreds. Greenplum is a data warehouse solution based on the relational model. It has advantages in processing structured data, especially relational data. It is more suitable for enterprises or organizations such as telecommunications and banks whose data is mainly stored in a structured manner.

[0080] China Mobile currently has nearly 140 million home broadband users nationwide. The smart gateway and the built-in plug-ins (soft probes) of the Magic Box upload more than 10 billion data items every day (more than 300 billion data items per month), with a daily data size of about 10 to 20T and a monthly data size of about 500T. The Network Department requires that the data be subjected to on-line analytical processing (OLAP) multidimensional data analysis in dozens of dimensions at a specific time granularity (5 minutes, hours, days, months or any time), and that data insights be obtained through any combination of dimensions to gain an in-depth understanding of the quality of the data itself, the status of the equipment network, and the development of various businesses; and on the basis of multidimensional analysis of soft probe data, further exploration of applications based on machine learning and deep learning technologies, such as user profiling and precision marketing, be conducted.

[0081] At present, the big data analysis technology solution of soft probe uses Hadoop / Spark parallel computing cluster based on more than 20 machines. However, in the actual application process layer, the following three problems are found:

[0082] 1) Complex solutions

[0083] Hadoop big data solution is an ecosystem, which can be compared to the various tools needed in a kitchen, such as pots, pans, and bowls, each with its own uses and overlaps with each other. Hadoop Distributed File System (HDFS) is used for distributed storage, MapReduce, Tez, and Spark are distributed computing engines, Pig and Hive are higher-level and more abstract language layers used to describe algorithms and data processing processes, and interactive SQL engines such as Impala, Presto, and Drill are used to process SQL tasks more quickly, sacrificing features such as generality and stability. Hive on Tez / Spark and SparkSQL are used in scenarios that require higher data query speeds, and Storm is the most popular stream computing platform. Since Hadoop big data solutions are very complex, practice has proved that this complexity has a very large adverse impact on the learning, installation, deployment, development, use, and operation and maintenance of big data solutions.

[0084] 2) Performance does not meet requirements

[0085] Hadoop itself runs on a cluster of ordinary personal computers (PC) servers to distribute and process big data, and has very high resource requirements. The project has currently increased the total number of cluster servers to more than 20 machines, but it can only guarantee 8 to 10 hours to process the 10 to 20 terabytes of data generated by the soft probe every day, which has a low cost-performance ratio. In addition, the use of Hadoop has performance degradation in specific scenarios, such as the presence of a large number of small files and scenarios that are sensitive to response delays. The actual scenario of this user is precisely the presence of a large number of small files and the need for instant response analysis.

[0086] 3) Low flexibility

[0087] At present, the main user application scenarios are deterministic data indicator analysis (starting scheduled tasks at midnight every day) and real-time OLAP data analysis (sudden user troubleshooting and data analysis tasks). The second scenario requires flexible combination of query, drilling, slicing and other OLAP operations in any dimension. For abnormal data, it is also necessary to trace back the data source level by level to verify the correctness of the data. Hadoop solutions cannot meet the above requirements well.

[0088] The purpose of the embodiments of the present application is to solve the above three technical problems and realize a big data computing framework for building a multi-layer data analysis model based on the building block stacking method, so as to achieve the purpose of simplicity, efficiency and flexibility and meet the user's requirements for functions and performance in big data analysis.

[0089] LEGO is a kind of plastic building block invented by Ole Kirk Christiansen of Denmark. By studying the relevant information of LEGO, its characteristics can be summarized as follows: LEGO has only several basic components, each component is connected to other components through a standard interface. Although each basic component is very simple, any object model can be flexibly built through appropriate combination.

[0090] Since the object model built with Lego blocks is simple, efficient, and flexible, it perfectly meets our needs. Therefore, we can consider borrowing the idea of ​​Lego blocks to perform data analysis. First, each data analysis task can be abstracted into six atomic tasks: task distribution, data cleaning, data forwarding, statistical analysis, result aggregation, and log management. Each atomic task is packaged into a computing unit, which is equivalent to a type of Lego block. Secondly, a standard data communication interface for different types of computing units is constructed through a message queue. Through this interface, different types of computing units are organized into a multi-layer fully connected network to realize the flow of messages at different computing unit levels. Finally, most of each computing unit is fixed, and only the business computing logic interface changes. By overloading this interface, it can adapt to different user analysis needs, while maintaining simplicity and flexibility.

[0091] In the above way, any data analysis needs of users can be mapped into a multi-layer data analysis model in which different types of computing units are stacked in a building block manner. In this model, different types of computing units are organized into a multi-layer fully connected network. Each computing node only focuses on a certain type of simple data computing task. However, through the reasonable collaboration of computing units at different levels, any data analysis task can be completed simply, flexibly and efficiently.

[0092] Figure 4 A flow chart of another data analysis method is provided for the embodiment of the present application, such as Figure 4 As shown, the method includes:

[0093] Step S401, the parser idle notification is sent to the manager;

[0094] The parser sends a Parser Idle Notification message to the manager after the model is started.

[0095] Step S402: The manager checks that the pool is not empty and returns a file in the pool to the parser;

[0096] After receiving the "parser idle notification" message sent by the parser, the manager checks whether there are unprocessed files in the waiting data file pool. If there are, the manager sends the file name to the specific parser through a message.

[0097] Each data analysis task contains multiple files, and each data file manager will be assigned a parser task. The task allocation rule means that if the parser is idle, it will actively tell the manager, and the manager will then take a file name from the pool of files to be processed and send it to the parser for processing. This active claim mechanism of everyone doing their best ensures maximum processing efficiency.

[0098] Step S403: Process the file sent by the manager in the parser;

[0099] First, after receiving the file name from the manager, the parser reads and decodes the file, cleans and converts the original data line by line, and generates assembled data according to business needs; at the same time, various status data are recorded during the parsing process (such as the distribution of various exceptions, the number of processed files, the number of lines, the duration, etc.). Here, the user's data analysis needs can be mapped to the logic of data analysis. For example, this analysis needs to extract 20 fields from all 100 fields, then the parser only extracts these 20 fields for assembly and sends them to the next level counter through the router. The assembled data uses the simplest special characters to separate each field into a string. The assembled data uses the simplest special characters to separate each field into a string.

[0100] Step S404, the parser sends an assembly message to the router;

[0101] Send the assembly message containing the assembly data to the router.

[0102] Step S405: The router receives the assembly message and routes it to a counter according to the hash value of the device ID;

[0103] The router extracts the device number from the received assembly message and routes the message to the corresponding counter Ni according to the following algorithm, where the value of i ranges from 0 to n-1. The algorithm is: i = hash (device number) % n; where hash (device number) is the hash value of the device ID, % is the modulus operation, and n is the number of counter instances;

[0104] The purpose of this step is to ensure that multiple messages from the same device are counted by the same counter.

[0105] Step S406, the counter processes the received data;

[0106] The counter generates summary data based on user granularity according to user needs, and updates the summary data after receiving the assembly message, that is, the data of the same device is aggregated according to the specified calculation logic. The following is the simplest example. Table 1A is aggregated to obtain Table 1B as shown below: the assembled data of home users watching programs are aggregated into summary data with user + day granularity.

[0107] Table 1A

[0108]

[0109]

[0110] Table 1B

[0111] Device Number area time Watching time … 10001 Beijing 2019-1-5 640 10002 Tianjin 2019-1-5 360

[0112] At the same time, various status data (such as the distribution of various exceptions, the number of messages processed, duration, etc.) are recorded during the aggregation process.

[0113] Step S407, repeat the above steps to process the next file until the pool of files to be processed is empty;

[0114] Repeat steps S401-S406 to process the next data file until M parsers have processed all the data files to be processed in the manager. In this case, parser Mi sends a "parser idle notification" message to the manager, because the pool of data files to be processed in the manager is empty, so the manager sends an end command to parser Mi; at the same time, the manager records the parser as completed, and then the manager waits for all parsers to complete and then ends the manager process.

[0115] Step S408: The parser idle notification is sent to the manager;

[0116] In this case, the resolver Mi sends a "Resolver Idle Notification" message to the manager.

[0117] Step S409: The manager checks that the pool is empty and returns an end command to the parser;

[0118] Because the pool of data files to be processed in the manager is empty, the manager sends an end command to the parser Mi.

[0119] Step S410: the parser sends the status data to the aggregator, and then goes to 415;

[0120] After receiving the end message from the manager, the resolver first sends the status data of the resolver to the aggregator.

[0121] Step S411, in the manager, the total number of parsers completed is increased by 1;

[0122] At the same time, the manager records the parser as completed, and then the manager waits for all parsers to complete and then ends the manager process.

[0123] Step S412: the parser sends an end message to the router;

[0124] Step S413: The router routes the end message to all lower-level counters;

[0125] The end message is routed by the router to all counter instances C1…Cn.

[0126] Step S414, add 1 (+1) to the total number of parser completions in the counter;

[0127] Because a router's end message is received, indicating that a resolver has completed the resolution, the total number of resolvers completed in the counter is increased by 1 until all resolvers have completed the resolution.

[0128] Step S415: The aggregator aggregates the status data of all the parsers into global status data;

[0129] Aggregate the status data from M parsers, N counters and its own status data into global status data; for example, each parser will record the number of files processed, the number of lines, the error distribution, etc., and the aggregation by the aggregator will obtain the number of files processed by the entire data model parsing layer, the number of input and output lines, the error distribution, etc.

[0130] Step S416, waiting in the counter for all upper-level parsers to complete processing;

[0131] After receiving the end message, the counter marks a parent parser connected to it as completed, and then waits for all parent parsers to complete processing.

[0132] Step 417: Send the status data of the counter to the aggregator for aggregation;

[0133] Step S418: Aggregate the status data of all counters into global status data in the aggregator;

[0134] In step 419 and step 420, the counter generates full-dimensional statistical data based on the user-granularity summary data according to user needs and sends it to the aggregator. That is, the data of the same dimension is aggregated according to the specified calculation logic. The following is a simple example. The following Table 2A is aggregated to obtain Table 2B as shown below: The user-granularity summary data is aggregated into full-dimensional (only two dimensions of time and region) data.

[0135] Table 2A

[0136] Device Number area time Watching time … 10001 Beijing 2019-1-5 680 10002 Tianjin 2019-1-5 360 10004 Beijing 2019-1-5 680 10005 Tianjin 2019-1-5 360 … Device Number area time Watching time …

[0137] Table 2B

[0138] area time Number of viewers Total viewing time Average viewing time Beijing 2019-1-5 43 680 15.8 Beijing 2019-1-6 12 360 30.0 Tianjin 2019-1-5 22 450 20.5 …

[0139] Step S421: Aggregate all full-dimensional data into global full-dimensional data in the aggregator;

[0140] Aggregate the statistical data from n counters into global summary data; for example, each counter will count the number of viewers in Beijing within the range of the counter, and the global number of viewers in Beijing will be obtained through aggregation by the aggregator.

[0141] Step S422: the counter sends an end message to the aggregator;

[0142] Sending a completion message to the aggregator indicates that the counter has completed all work. Receiving the completion message of the counter marks that a certain upper-level counter connected to it has completed, and then waits for all n upper-level counters to complete processing.

[0143] Step S423, in the aggregator, the total number of the counter is increased by 1;

[0144] Every time a counter sending completion message is received in the aggregator, the total number of the counter is increased by 1.

[0145] Step S424, the aggregator waits for all upper-level counters to complete processing;

[0146] Until the total number of counter completions is equal to the number of counter instances N, it means that all upper-level counter processing is completed.

[0147] Step S425: The aggregator generates combined dimension data based on the full dimension data, summarizes the data at each level, and saves and stores the status data in the database;

[0148] According to user needs, combined dimension data is generated based on full-dimensional data, such as the multi-day viewing situation of Magic 100 Box nationwide is generated based on full-dimensional data as shown in Table 3A, as shown in Table 3B;

[0149] Table 3A

[0150]

[0151]

[0152] Table 3B

[0153] area time Number of viewers Total viewing time Average viewing time Nationwide 2019-1-5 666565656 68099999999 102.2 Nationwide 2019-1-6 56565656 360878888 6.4 Nationwide 2019-1-7 56565666 450666 0.0 …

[0154] The data at each level (user granularity data -> full dimension data -> combined dimension data) and status data are saved in the database for subsequent user inquiries.

[0155] Step S426: The aggregator sends an end message to the logger;

[0156] The data analysis model described in the embodiment of the present application can be deployed on a physical machine. In other embodiments, the data analysis model can also be used in a distributed system including multiple physical machines. Multiple physical machines deployed with the data analysis model can perform data analysis tasks at the same time, and then these physical machines transmit the output analysis results to a physical machine in the distributed system to execute the final converged results, so that the analysis data can achieve the effect of improving data analysis efficiency.

[0157] The key point of the embodiment of the present application is to draw lessons from the idea of ​​building object models with Lego blocks, that is, to abstract the general process of data analysis into six types of atomic computing nodes (file manager, parser, counter, aggregator, logger and virtual router), stack the six types of atomic computing nodes layer by layer, and realize the mesh full connection of the communication between adjacent layers of computing nodes through message queues-similar to the deep learning neural network model, so as to realize the circulation of data flow, control flow, log flow and state flow between computing nodes in each layer. Different categories of computing nodes are organized into a multi-layer fully connected network through standard data interfaces, and each type of computing node is only responsible for very simple tasks, but through the reasonable division of labor and cooperation of different computing nodes, it can be very simple, efficient and flexible to meet any data analysis needs of users. For any data analysis needs, it can be mapped to the data analysis model built by studybuilder and the specific calculation logic of different computing nodes in the model. Since the computing nodes are divided into fine enough, each computing node only needs to complete extremely simple tasks, but through the reasonable division of labor and cooperation of computing nodes in each layer in the data analysis model, various very complex data analysis tasks can be completed efficiently and flexibly.

[0158] The advancedness of the embodiments of the present application is mainly reflected in the following five aspects:

[0159] First, high performance. On a single physical machine (Intel(R) Xeon(R) CPU E5-2640v3@2.60G32 cores 512G Mem), it can process 1.2 billion records per hour (about 2T of data), which is 3 to 5 times faster than Hadoop, another big data processing framework, under the same conditions.

[0160] Second, ease of use. By referring to the use of artificial intelligence (AI) frameworks, the complexity is hidden and the threshold for use is lowered through program templates. Users only need to specify the number of analysis levels and the number of computing nodes involved in each layer, as well as the calculation logic of each layer. More than 80% of data analysis only requires changing dozens or even a few lines of code.

[0161] Third, robustness. The computing framework uses message queues to organically combine six computing node types into a specific computing model in a pipeline mode. The task of each computing node is very simple, with less code and easy maintenance. However, various complex data analysis tasks can be completed through appropriate combinations. This design ensures the robustness of the system.

[0162] Fourth, flexibility. Any data analysis requirement can be mapped to a data analysis model composed of six basic computing nodes and their corresponding computing logic. At the same time, the tower-like multi-layer data model can be adapted to various applications and instantly generate statistical data of any combination of dimensions, ensuring the flexibility of data analysis.

[0163] The fifth aspect is transparency, which avoids the black box effect of other big data processing frameworks. In addition to the final statistical data results, the intermediate state data of all computing nodes and computing levels, such as error data distribution and processing time, are summarized to facilitate subsequent other analyses, such as data accuracy retrospective analysis, performance optimization, and data quality analysis.

[0164] Based on the foregoing embodiments, the embodiments of the present application provide a data analysis model, which includes various parsers, routers, counters and aggregators, etc., which can be implemented by a processor in a data analysis device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.

[0165] Figure 5A A schematic diagram of the composition structure of the data analysis model provided in the embodiment of the present application is shown in FIG. Figure 5A As shown, the model 500 includes a parser 501, a router 502, a counter 503 and an aggregator 504, where M and N are integers greater than or equal to 2; wherein:

[0166] A parser is used to obtain data to be analyzed from different devices in a data pool, wherein the data to be analyzed includes a device identifier; each of the parsers is used to parse the data to be analyzed, assemble the data to be analyzed according to statistical requirements, and obtain an assembled message;

[0167] A router, configured to send each of the assembly messages to a corresponding counter according to a device identifier;

[0168] A counter, used to aggregate the assembled messages according to a specified calculation logic to obtain first statistical data;

[0169] An aggregator is used to aggregate the first statistical data outputted by the N counters into second statistical data.

[0170] Figure 5B A schematic diagram of the composition structure of the data analysis model provided in the embodiment of the present application is shown in FIG. Figure 5B As shown, the model 510 includes a manager 511, a parser 501, a router 502, a counter 503, an aggregator 504 and a logger 515, wherein:

[0171] The manager 511 is used to send the data to be analyzed to the parser;

[0172] The parser 501 is used to parse the data to be analyzed, assemble the data to be analyzed according to statistical requirements, and obtain assembled messages;

[0173] Router 502, used for sending each of the assembly messages to a corresponding counter according to the device identification;

[0174] A counter 503, used to aggregate the assembled messages according to a specified calculation logic to obtain first statistical data;

[0175] an aggregator 504, configured to aggregate the first statistical data outputted from the N counters into second statistical data;

[0176] The logger 515 is used to record the logs output by the manager, each of the parsers, each of the routers, each of the counters and the aggregator.

[0177] Based on the foregoing embodiments, an embodiment of the present application provides a data analysis model, the model comprising a manager, a parser, a router, a counter and an aggregator, wherein:

[0178] The manager is used for sending an idle notification message to the manager when each of the parsers determines that it is idle; when the manager checks that the data pool is not empty, it obtains the data to be analyzed from different devices from the data pool and sends the data to be analyzed to the parser.

[0179] A parser is used to obtain data to be analyzed from different devices in a data pool, wherein the data to be analyzed includes a device identifier; each of the parsers is used to parse the data to be analyzed, assemble the data to be analyzed according to statistical requirements, and obtain an assembled message;

[0180] A router, configured to send each of the assembly messages to a corresponding counter according to a device identifier;

[0181] A counter, used to aggregate the assembled messages according to a specified calculation logic to obtain first statistical data;

[0182] An aggregator is used to aggregate the first statistical data outputted by the N counters into second statistical data.

[0183] Based on the foregoing embodiments, the embodiments of the present application provide a data analysis model, the model including a manager, a parser, a router, a counter, an aggregator and a logger, wherein:

[0184] The manager is used for sending an idle notification message to the manager when each of the parsers determines that it is idle; when the manager checks that the data pool is not empty, it obtains the data to be analyzed from different devices from the data pool and sends the data to be analyzed to the parser.

[0185] A parser is used to obtain data to be analyzed from different devices in a data pool, wherein the data to be analyzed includes a device identifier; each of the parsers is used to parse the data to be analyzed, assemble the data to be analyzed according to statistical requirements, and obtain an assembled message;

[0186] A router, configured to send each of the assembly messages to a corresponding counter according to a device identifier;

[0187] A counter, used to aggregate the assembled messages according to a specified calculation logic to obtain first statistical data;

[0188] An aggregator is used to aggregate the first statistical data outputted by the N counters into second statistical data.

[0189] A logger is used to record logs output by the manager, each of the parsers, each of the routers, each of the counters and the aggregator.

[0190] Based on the foregoing embodiments, the embodiments of the present application provide a data analysis model, the model including a manager, a parser, a router, a counter, an aggregator and a logger, wherein:

[0191] The manager is configured to send an idle notification message to the manager when each of the parsers determines that it is idle; when the manager finds that the data pool is not empty, obtains the data to be analyzed from different devices from the data pool, and sends the data to be analyzed to the parser. When the manager finds that the data pool is empty, it sends an end message to the parser;

[0192] The parser is used to obtain the data to be analyzed from different devices in the data pool, wherein the data to be analyzed includes a device identifier; each of the parsers is used to parse the data to be analyzed, assemble the data to be analyzed according to statistical requirements, and obtain an assembly message; when each of the parsers receives the end message, it sends a parsing end message to the corresponding router;

[0193] A router, configured to send each of the assembly messages to a corresponding counter according to a device identifier; when each of the routers receives the parsing end message, send a routing end message to the N counters;

[0194] A counter, when the M routers complete sending the routing end message, the counter aggregates the assembled messages according to the specified calculation logic to obtain the first statistical data.

[0195] An aggregator is used to aggregate the first statistical data outputted by the N counters into second statistical data.

[0196] A logger is used to record logs output by the manager, each of the parsers, each of the routers, each of the counters and the aggregator.

[0197] Based on the foregoing embodiments, the embodiments of the present application provide a data analysis model, the model including a manager, a parser, a router, a counter, an aggregator and a logger, wherein:

[0198] The manager is used for sending an idle notification message to the manager when each of the parsers determines that it is idle; when the manager checks that the data pool is not empty, it obtains the data to be analyzed from different devices from the data pool and sends the data to be analyzed to the parser.

[0199] A parser is used to obtain data to be analyzed from different devices in a data pool, wherein the data to be analyzed includes a device identifier; each of the parsers is used to parse the data to be analyzed, assemble the data to be analyzed according to statistical requirements, and obtain an assembled message; when each of the parsers parses the data to be analyzed, it records first status data; wherein: the first status data includes at least one of the following: the number of files processed, the number of lines, the error distribution, and the parser status of itself; each of the parsers sends the first status data to the aggregator;

[0200] A router, configured to send each of the assembly messages to a corresponding counter according to a device identifier;

[0201] A counter is used to aggregate the assembled messages according to a specified calculation logic to obtain first statistical data; each of the counters records third status data when aggregating the first statistical data; wherein the second status data includes at least one of the following: recording the number of processed files, the number of lines, the error distribution and the status of the counter; each of the counters sends the third status data to the aggregator;

[0202] An aggregator is used to aggregate the first statistical data outputted from the N counters into second statistical data. When the M parsers finish sending the first status data to the aggregator, the aggregator aggregates the M first status data into second status data and outputs the second status data; when the N counters finish sending the third status data to the aggregator, the aggregator aggregates the N third status data into fourth status data and outputs the fourth status data. The aggregator aggregates the second status data and the fourth status data to obtain fifth status data;

[0203] A logger is used to record logs output by the manager, each of the parsers, each of the routers, each of the counters and the aggregator.

[0204] Therefore, the description of the above model embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the model embodiment of this application, please refer to the description of the method embodiment of this application for understanding.

[0205] It should be noted that in the embodiment of the present application, if the above-mentioned data analysis method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer or multiple computer-readable storage media. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions to enable an electronic device (which can be a mobile phone, a tablet computer, a laptop computer, a desktop computer, a robot, a drone, etc.) to perform all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0206] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the data analysis method provided in the above embodiment are implemented.

[0207] Correspondingly, an embodiment of the present application provides a data analysis device, Figure 6 A schematic diagram of a hardware entity of a data analysis device according to an embodiment of the present application is shown in FIG. Figure 6As shown, the hardware entity of the device 600 includes: a memory 601 and a processor 602, wherein the memory 601 stores a computer program that can be run on the processor 602, and the processor 602 implements the steps in the data analysis method provided in the above embodiment when executing the program.

[0208] The memory 601 is configured to store instructions and applications executable by the processor 602, and can also cache data to be processed or processed by the processor 602 and various modules in the data analysis device 600 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0209] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0210] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0211] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, or product including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, or product. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, or product including the element.

[0212] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0213] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0214] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0215] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.

[0216] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly reflected in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a data analysis device (which can be a mobile phone, tablet computer, laptop computer, desktop computer, robot, drone, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0217] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0218] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0219] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0220] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A data analysis method, characterized in that: Applied to a data analysis model, the data analysis model includes M parsers, M routers corresponding to the M parsers, N counters, 1 aggregator, and 1 manager, M and N are integers greater than or equal to 2, and the method includes: Each of the M parsers obtains data to be analyzed from different devices in the data pool, wherein the data to be analyzed includes a device identifier; Each of the parsers parses the data to be analyzed, assembles the data to be analyzed according to statistical requirements, and obtains assembled messages; Each of the routers sends each of the assembly messages to a corresponding counter according to the device identifier; Each of the N counters aggregates the assembled messages according to a specified calculation logic to obtain first statistical data; The aggregator aggregates the first statistical data outputted from the N counters into second statistical data; Each of the N counters aggregates the assembled messages according to a specified calculation logic to obtain first statistical data, including: When the manager finds that the data pool is empty, it sends an end message to the parser; When each of the parsers receives the end message, it sends a parsing end message to the corresponding router; When each of the routers receives the resolution completion message, it sends a routing completion message to the N counters; When the M routers complete sending the routing completion messages, the counter aggregates the assembled messages according to the specified calculation logic to obtain the first statistical data.

2. The method according to claim 1, characterized in that Each of the M parsers obtains the data to be analyzed from different devices in the data pool, including: When each of the parsers determines that it is idle, it sends an idle notification message to the manager; When the manager checks that the data pool is not empty, it obtains the data to be analyzed from different devices from the data pool, and sends the data to be analyzed to the parser.

3. The method according to claim 1 or 2, characterized in that: The data analysis model also includes a logger. Correspondingly, the method also includes: The logger is responsible for recording the logs output by the manager, each of the parsers, each of the routers, each of the counters and the aggregator.

4. The method according to any one of claims 1 to 2, characterized in that: The method further comprises: Each of the parsers records first status data when parsing the data to be analyzed; wherein: the first status data includes at least one of the following: the number of files processed, the number of lines, the error distribution and the parser status of itself; Each of the parsers sends the first status data to the aggregator; When the M parsers finish sending the first state data to the aggregator, the aggregator aggregates the M first state data into second state data; The second state data is output.

5. The method according to claim 4, characterized in that The method further comprises: Each of the counters records third status data when aggregating the first statistical data; wherein: the second status data includes at least one of the following: recording the number of files processed, the number of lines, the error distribution and the status of the counter; Each of the counters sends the third status data to the aggregator; When the N counters finish sending the third state data to the aggregator, the aggregator aggregates the N third state data into fourth state data; The fourth state data is output.

6. The method according to claim 5, characterized in that The method further comprises: The aggregator aggregates the second state data and the fourth state data to obtain fifth state data.

7. A data analysis model, characterized in that: It includes M resolvers, M routers corresponding to the M resolvers, N counters, 1 aggregator and 1 manager, where M and N are integers greater than or equal to 2; wherein, Each of the M parsers is used to obtain data to be analyzed from different devices in the data pool, wherein the data to be analyzed includes a device identifier; Each of the parsers is used to parse the data to be analyzed, assemble the data to be analyzed according to statistical requirements, and obtain assembled messages; Each of the routers is used to send each of the assembly messages to a corresponding counter according to a device identifier; Each of the N counters is used to aggregate the assembled messages according to a specified calculation logic to obtain first statistical data; The aggregator is used to aggregate the first statistical data outputted from the N counters into second statistical data; The manager is used to send an end message to the parser when checking that the data pool is empty; Each of the parsers is further configured to send a parsing end message to a corresponding router upon receiving the end message; Each of the routers is further configured to send a routing completion message to the N counters upon receiving the resolution completion message; Each of the N counters is also used for aggregating the assembled messages according to a specified calculation logic to obtain the first statistical data when the M routers complete sending the routing end message.

8. A data analysis device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: The processor implements the steps in the method of any one of claims 1 to 6 when executing the program by calling the data analysis model.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Big data query method, system, computer and storage medium

    CN108009236A

  • Network health data aggregation service

    CN110036600A