Data analysis method, system, terminal and storage medium

By using load balancing strategies in the ClickHouse cluster to analyze user behavior data, the problems of low computing efficiency and instability in the analysis of massive user behavior data are solved, and efficient and stable data analysis is achieved and maintenance costs are reduced.

CN115048466BActive Publication Date: 2025-05-06ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210504417.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-10
Publication Date
2025-05-06
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

The existing technology has problems with low computing efficiency, high maintenance costs and data privacy in the analysis of massive user behavior data, especially when the ClickHouse cluster reaches its computing peak.

Method used

By obtaining the behavioral data of the target object and storing it in the ClickHouse cluster according to preset rules, a variety of load balancing strategies are used to schedule data analysis tasks to ensure load balancing of each node and improve the stability and efficiency of data analysis.

Benefits of technology

It has achieved the improvement of computing performance of user behavior data analysis under ultra-large data scale, improved the stability and efficiency of data analysis, reduced maintenance costs, and enhanced the accuracy of data analysis products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115048466B_ABST
    Figure CN115048466B_ABST
Patent Text Reader

Abstract

The present application relates to a data analysis method, system, terminal and storage medium, wherein the data analysis method includes: obtaining the behavior data of the target object, and storing the behavior data in the ClickHouse cluster according to preset rules; obtaining data analysis task information, and scheduling tasks for each node of the ClickHouse cluster according to the data analysis task information to balance the load of each node; analyzing the behavior data of each node according to the task scheduling arrangement to generate a data analysis product. The data analysis method, system, terminal and storage medium provided in the present application use the ClickHouse cluster to store behavior data, and use a variety of load balancing strategies to schedule data analysis tasks, which can meet the analysis needs of user behavior data under ultra-large data scale, improve the stability and efficiency of data analysis, and improve the accuracy of data analysis products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data analysis technology, and in particular, relates to a data analysis method, system, terminal and storage medium. Background Art

[0002] In the field of analyzing massive user behavior data, native big data computing engines such as hive, spark, presto, impala, elasticsearch, etc. seem helpless. Well-known companies in the industry often conduct secondary development of big data computing engines to achieve efficient analysis of massive user behavior data, but the cost of secondary development is extremely high and the computing efficiency is often unsatisfactory, and the subsequent maintenance cost is also very high. Therefore, most companies generally purchase commercial products to make up for the shortcomings in the field of massive user behavior data analysis, but expensive commercial products and data privacy issues have also become potential risks for corporate development.

[0003] Existing technologies, such as patent CN202011006169.X, provide a method for implementing OLAP analysis based on ClickHouse. It provides a detailed description of table construction specifications, data writing, SQL queries, etc., but does not involve how to create stable and efficient data products based on ClickHouse in the face of massive user behavior data. In addition, the computing performance bottleneck of user behavior data analysis under ultra-large data scale, and the instability of ClickHouse clusters when reaching computing peaks, etc., still exist. Summary of the invention

[0004] In response to the above technical problems, the present application provides a data analysis method, system, terminal and storage medium to meet the analysis needs of user behavior data under ultra-large data scale, improve the stability and efficiency of data analysis, and enhance the accuracy of data analysis products.

[0005] The present application provides a data analysis method, including: obtaining behavior data of a target object, and storing the behavior data in a ClickHouse cluster according to preset rules; obtaining data analysis task information, and scheduling tasks for each node of the ClickHouse cluster according to the data analysis task information to balance the load of each node; analyzing the behavior data of each node according to the task scheduling arrangement to generate a data analysis product.

[0006] In one embodiment, the behavior data is stored in a ClickHouse cluster according to preset rules, including: hashing the behavior data according to the identity identification number of the target object, and writing the behavior data of each target object into the corresponding node of the ClickHouse cluster; storing the behavior data of each target object according to a preset storage mode.

[0007] In one embodiment, the step of storing the behavior data of each target object according to a preset storage mode includes: pre-sorting the behavior data of each target object according to a three-level index order; wherein the first-level index is the event number of the behavior data; the second-level index is the identity identification number of the target object to which the behavior data belongs; and the third-level index is the log time of the behavior data.

[0008] In one embodiment, the data analysis task information includes the task type of the data analysis task; wherein the task type includes at least one of event statistics, portrait analysis, funnel analysis, behavior path analysis, table structure change, and clearing expired data.

[0009] In one embodiment, task scheduling is performed on each node of the ClickHouse cluster, including at least one of the following: executing different types of data analysis tasks in sequence according to task execution priority; and adopting corresponding load balancing strategies according to the task type of the data analysis task, wherein the load balancing strategies include random, polling, and minimum load.

[0010] In one embodiment, according to the task type of the data analysis task, a corresponding load balancing strategy is executed, including: if the task type is event statistics and / or portrait analysis and / or funnel analysis and / or behavior path analysis, a minimum load strategy is adopted; if the task type is table structure change, a random strategy is adopted; if the task type is to clean up expired data, a polling strategy is adopted.

[0011] In one embodiment, task scheduling is performed on each node of the ClickHouse cluster, and the further step includes: obtaining the number of read rows and the execution time of the data analysis task; and stopping the data analysis task when the number of read rows of the data analysis task exceeds the maximum number of read rows and / or the execution time of the data analysis task exceeds the maximum execution time.

[0012] The present application also provides a data analysis system, which includes a data writing module, a data storage module, a task scheduling module, and a data analysis module; the data writing module is used to obtain the behavior data of the target object, and write the behavior data into the ClickHouse cluster according to a first preset rule; the data storage module is used to store the behavior data written into the ClickHouse cluster according to a second preset rule; the task scheduling module is used to obtain data analysis task information, and according to the data analysis task information, perform task scheduling on each node of the ClickHouse cluster to balance the load of each node; the data analysis module is used to analyze the behavior data of each node according to the task scheduling arrangement to generate a data analysis product.

[0013] The present application also provides a terminal, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned analysis method when executing the computer program.

[0014] The present application also provides a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned analysis method are implemented.

[0015] The present application provides a data analysis method, system, terminal and storage medium, which use ClickHouse cluster to store behavior data and adopt multiple load balancing strategies to schedule data analysis tasks. It can meet the analysis needs of user behavior data under ultra-large data scale, improve the stability and efficiency of data analysis, and enhance the accuracy of data analysis products. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a flowchart of the data analysis method provided in Example 1 of the present application;

[0017] Figure 2 is a structural diagram of a data analysis system provided in Example 2 of the present application;

[0018] Figure 3 It is a schematic diagram of the structure of the terminal provided in Example 3 of the present application. DETAILED DESCRIPTION

[0019] The technical solution of the present application is further elaborated in detail below in conjunction with the accompanying drawings and specific embodiments of the specification. Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as those commonly understood by technicians in the technical field of this application. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. "And / or" used herein includes any and all combinations of one or more related listed items.

[0020] Figure 1 Schematic diagram of the data analysis method provided in Example 1 of the present application. Figure 1 As shown, the data analysis method of the present application may include the following steps:

[0021] Step S101: Obtain the behavior data of the target object, and store the behavior data in the ClickHouse cluster according to preset rules;

[0022] In one embodiment, the behavior data is stored in the ClickHouse cluster according to preset rules, including:

[0023] According to the target object's identification number, the behavior data is hashed and sharded, and the behavior data of each target object is written to the corresponding node of the ClickHouse cluster;

[0024] The behavior data of each target object is stored according to the preset storage mode.

[0025] Optionally, the open source stream processing framework (Flink) service is used to insert data into the ClickHouse cluster in the form of Java database connection (jdbc); further, in order to ensure the computing efficiency during query and reduce the writing pressure as much as possible, the data is hashed according to the identity identification number (user ID) of the target object and directly inserted into the local table of the ClickHouse cluster to avoid the problem of distributed table write amplification; finally, all behavioral data of the same user are stored on the same machine to avoid a large amount of network read and write (IO) transmission during calculation, thereby realizing localized calculation.

[0026] In one embodiment, the step of storing the behavior data of each target object according to a preset storage mode includes:

[0027] Pre-sort the behavior data of each target object according to the three-level index;

[0028] Among them, the first-level index is the event number of the behavior data; the second-level index is the identity identification number of the target object to which the behavior data belongs; and the third-level index is the log time of the behavior data.

[0029] Optionally, in terms of primary key selection, to ensure efficient retrieval, the event number (event_id) of the behavior data is selected as the primary index; to ensure efficient data query, the identity number (xxHash32 (distinct_id) of the target object to which the behavior data belongs is selected as the secondary index, and the log time (log_time) of the behavior data is selected as the tertiary index, so that all data are pre-sorted in advance in the storage layer according to the order of the tertiary index.

[0030] In one embodiment, the step of storing the behavior data of each target object according to a preset storage mode further includes:

[0031] In terms of engine selection, the mergeTree engine is selected as the distributed computing engine;

[0032] In terms of partition selection, the generation date of the behavior data is selected as the partition field to achieve horizontal table partitioning;

[0033] In terms of sampling field selection, in order to ensure the millisecond-level response of the estimated scenario, the target object's identity identification number (distinct_id) is selected as the sampling field, and the sampling result obtained by the hash function (xxHash32(distinct_id)) is used as the sampling benchmark;

[0034] In terms of data survival time (TTL), considering local storage limitations and data usage frequency, data within a preset time is defined as hot data, such as the past two months, and data outside the preset time is defined as cold data, such as two months ago. Data analysis services are provided for hot data, and cold data can be analyzed through the Hadoop distributed computing platform;

[0035] In terms of data granularity, the official recommendation is adopted, that is, an index is generated every 8192 rows.

[0036] Step S102: Obtain data analysis task information, and schedule tasks for each node of the ClickHouse cluster according to the data analysis task information to balance the load of each node;

[0037] Optionally, the data analysis task information includes task type and task configuration information; wherein the task type includes at least one of event statistics, portrait analysis, funnel analysis, behavior path analysis, table structure change, and clearing expired data; the task configuration information is the analysis dimension or analysis scope of each type of task, such as the configuration information of the portrait analysis task is region, education level, gender, age, etc.; the configuration information of the funnel analysis task is the hierarchical relationship of users from login - ordering - payment.

[0038] Exemplarily, event statistics are to count the number of users who perform operations on an application or web page, such as logging in, placing an order, and making a payment; portrait analysis is to analyze the portrait information of registered users, such as the number of male registered users in the Hangzhou area; funnel analysis is to analyze the conversion rate of the number of users from logging in, placing an order to paying, such as 1 million users log in, 800,000 users place an order, and 500,000 users pay, then the conversion rate from logging in to placing an order is 80%, and the conversion rate from placing an order to payment is 62.5%; behavioral path analysis is to analyze the user's operation trajectory from the first moment to the nth moment, such as the click trajectory or browsing trajectory of the application or web page content; table structure changes include adding and deleting tables, adding and deleting columns in tables, adding, deleting, and modifying fields in tables, etc., and cleaning up expired data is to clean up data that has exceeded the data lifetime.

[0039] In one embodiment, the task scheduling for each node of the ClickHouse cluster includes at least one of the following:

[0040] Execute different types of data analysis tasks in sequence according to task execution priority;

[0041] According to the task type of the data analysis task, the corresponding load balancing strategy is adopted, where the load balancing strategy includes random, polling, and minimum load.

[0042] Optionally, according to the task type of the data analysis task, a corresponding load balancing strategy is executed, including:

[0043] If the task type is event statistics and / or portrait analysis and / or funnel analysis and / or behavior path analysis, the minimum load strategy is adopted;

[0044] If the task type is table structure change, a random strategy is adopted;

[0045] If the task type is to clean up expired data, a polling strategy is adopted.

[0046] Optionally, since event statistics, portrait analysis, funnel analysis, and behavior path analysis occupy more computing resources, a minimum load strategy is adopted. Optionally, event statistics, portrait analysis, funnel analysis, and behavior path analysis tasks are identified in advance through custom task query IDs, and event statistics, portrait analysis, funnel analysis, and behavior path analysis tasks are assigned to nodes with fewer computing tasks in the ClickHouse cluster to reduce the load pressure on the ClickHouse cluster and achieve minimum load; since table structure changes occupy fewer computing resources, a random allocation strategy is adopted for table structure change tasks; since cleaning up expired data requires table construction and analysis on each node of the ClickHouse cluster, a polling strategy is adopted.

[0047] In one embodiment, performing task scheduling on each node of the ClickHouse cluster also includes:

[0048] Get the number of rows read and execution time of the data analysis task;

[0049] When the number of rows read by the data analysis task exceeds the maximum number of rows read, and / or the execution time of the data analysis task exceeds the maximum execution time, the data analysis task is stopped.

[0050] Optionally, the number of rows read by the data analysis task is the sum of the number of rows read by tasks at each node; the execution time of the data analysis task is the sum of the time for the entire process of sending the data analysis task to the master node, the master node sending the data analysis task to the child nodes, and then aggregating the behavior data of the child nodes to the master node; wherein the maximum number of rows read and the maximum execution time of the data analysis task are both preset values.

[0051] Step S103: Analyze the behavior data of each node according to the task scheduling arrangement and generate data analysis products.

[0052] Among them, the data analysis product is the data analysis result obtained by analyzing the behavior data of each node according to the task scheduling arrangement.

[0053] Optionally, task keywords are generated according to the data analysis task information, and the task keywords are associated with the generated data analysis products, so that other data analysis tasks can hit the existing data analysis products through the task keywords, thereby reducing a large amount of repeated calculations;

[0054] Optionally, set a validity period for the data analysis task, and the task will automatically become invalid when it expires;

[0055] Optionally, the data analysis task interacts with the ClickHouse cluster asynchronously to avoid thread blocking during the data analysis process.

[0056] The data analysis method provided in Example 1 of the present application adopts ClickHouse cluster as the basic computing service, and all data is stored in the local disk; at the data storage layer, the data is sharded according to the user ID, and all behavior data of the same user in the user dimension are stored on the same machine, thereby realizing localized computing in the distributed cluster mode; an independent scheduling service is introduced for the ClickHouse cluster, and all data analysis tasks are subject to the scheduling service, so as to effectively control the concurrency of the data analysis tasks; in terms of a single task, a maximum read row number limit and a maximum execution time limit are introduced to effectively control the maximum resources that a single task can occupy; a variety of load balancing strategies are also introduced to ensure that the pressure of each node in the cluster is balanced; the computing performance bottleneck of user behavior data analysis under ultra-large data scale and the instability problem of ClickHouse cluster when reaching the computing peak are solved, so as to meet the analysis needs of user behavior data under ultra-large data scale, effectively improve the stability and efficiency of data analysis, and improve the accuracy of data analysis products.

[0057] Figure 2 Schematic diagram of the structure of the data analysis system provided by the second embodiment of this application. Figure 2 As shown, the analysis system of the present application includes a data writing module 11, a data storage module 12, a task scheduling module 13, and a data analysis module 14;

[0058] Among them, the data writing module 11 is used to obtain the behavior data of the target object and write the behavior data into the ClickHouse cluster according to the first preset rule;

[0059] The data storage module 12 is used to store the behavior data written into the ClickHouse cluster according to the second preset rule;

[0060] The task scheduling module 13 is used to obtain data analysis task information, and according to the data analysis task information, perform task scheduling on each node of the ClickHouse cluster to balance the load of each node;

[0061] The data analysis module 14 is used to analyze the behavior data of each node according to the task scheduling arrangement and generate data analysis products.

[0062] Optionally, the task scheduling module 13 includes a distributed queue storage module 130, a task scheduler 131, and a ClickHouse client module 132;

[0063] Among them, the distributed queue storage module 130 uses a relational database management system (MySQL) to store the distributed data analysis task queue;

[0064] The task scheduler 131 is used to execute different types of data analysis tasks in sequence according to the task execution priority, ensure that the core tasks are executed first, and effectively control the task concurrency of the ClickHouse cluster;

[0065] The ClickHouse client module 132 is used to adopt a corresponding load balancing strategy according to the task type of the data analysis task, wherein the load balancing strategy includes random, polling, and minimum load;

[0066] The ClickHouse client module 132 is also used to obtain the number of read rows and execution time of the data analysis task, and stop the data analysis task when the number of read rows of the data analysis task exceeds the maximum number of read rows and / or the execution time of the data analysis task exceeds the maximum execution time.

[0067] The specific implementation process of this embodiment refers to Embodiment 1 and will not be repeated here.

[0068] The data analysis system provided in the second embodiment of the present application can meet the analysis requirements of user behavior data under ultra-large data scale through the interaction between the data writing module, the data storage module, the task scheduling module, and the data analysis module, effectively improves the stability and efficiency of data analysis, and improves the accuracy of data analysis products. In addition, the availability rate of the behavior data analysis system of the present application reaches 99.99%, and the data analysis efficiency can be improved from minutes to seconds.

[0069] Figure 3 2 is a schematic diagram of the structure of the terminal provided in the third embodiment of the present application. The terminal of the present application includes: a processor 210, a memory 211, and a computer program 212 stored in the memory 211 and executable on the processor 210. When the processor 210 executes the computer program 212, the steps in the above-mentioned data analysis method embodiments are implemented, for example Figure 1 Steps S101 to S103 are shown.

[0070] The terminal may include, but is not limited to, a processor 210 and a memory 211. Those skilled in the art will appreciate that Figure 3 It is only an example of a terminal and does not constitute a limitation on the terminal. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal may also include input and output devices, network access devices, buses, etc.

[0071] The processor 210 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0072] The memory 211 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. The memory 211 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Further, the memory 211 may also include both an internal storage unit and an external storage device of the terminal. The memory 211 is used to store the computer program and other programs and data required by the terminal. The memory 211 may also be used to temporarily store data that has been output or is to be output.

[0073] The present application also provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the data analysis method described above are implemented.

[0074] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0075] In this document, the terms "comprises," "comprising," or any other variations thereof, are intended to cover a non-exclusive inclusion of elements other than those listed and may also include additional elements not expressly listed.

[0076] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A data analysis method, characterized in that: include: Obtain the behavior data of the target object and store the behavior data in the ClickHouse cluster according to preset rules; Obtain data analysis task information, and schedule tasks for each node of the ClickHouse cluster according to the data analysis task information to balance the load of each node; According to the task scheduling arrangement, the behavior data of each node is analyzed to generate data analysis products; The step of performing task scheduling on each node of the ClickHouse cluster to balance the load of each node includes: Use a relational database management system to store distributed data analysis task queues; Execute different types of data analysis tasks in sequence according to task execution priority; According to the task type of the data analysis task, a corresponding load balancing strategy is adopted, wherein the load balancing strategy includes random, polling, and minimum load; Get the number of rows read and execution time of the data analysis task; When the number of rows read by the data analysis task exceeds the maximum number of rows read, and / or the execution time of the data analysis task exceeds the maximum execution time, the data analysis task is stopped.

2. The analysis method according to claim 1, characterized in that The behavior data is stored in the ClickHouse cluster according to the preset rules, including: According to the identity identification number of the target object, the behavior data is hashed and the behavior data of each target object is written to the corresponding node of the ClickHouse cluster; The behavior data of each target object is stored according to a preset storage mode.

3. The analysis method according to claim 2, characterized in that The step of storing the behavior data of each target object according to a preset storage mode includes: Pre-sorting the behavior data of each target object according to the three-level index order; Among them, the first-level index is the event number of the behavior data; the second-level index is the identity identification number of the target object to which the behavior data belongs; and the third-level index is the log time of the behavior data.

4. The analysis method according to claim 1, characterized in that The data analysis task information includes the task type of the data analysis task; wherein, the task type includes at least one of event statistics, portrait analysis, funnel analysis, behavior path analysis, table structure change, and clearing expired data.

5. The analysis method according to claim 1, characterized in that According to the task type of the data analysis task, the corresponding load balancing strategy is executed, including: If the task type is event statistics and / or portrait analysis and / or funnel analysis and / or behavior path analysis, the minimum load strategy is adopted; If the task type is table structure change, a random strategy is adopted; If the task type is to clean up expired data, a polling strategy is adopted.

6. A data analysis system, characterized in that: The analysis system includes a data writing module, a data storage module, a task scheduling module, and a data analysis module; The data writing module is used to obtain the behavior data of the target object and write the behavior data into the ClickHouse cluster according to the first preset rule; The data storage module is used to store the behavior data written into the ClickHouse cluster according to a second preset rule; The task scheduling module is used to obtain data analysis task information, and according to the data analysis task information, perform task scheduling on each node of the ClickHouse cluster to balance the load of each node; The data analysis module is used to analyze the behavior data of each node according to the task scheduling arrangement and generate data analysis products; The task scheduling module includes a distributed queue storage module, a task scheduler and a ClickHouse client module. The distributed queue storage module is used to store the distributed data analysis task queue using a relational database management system; The task scheduler is used to execute different types of data analysis tasks in order according to the task execution priority; according to the task type of the data analysis task, a corresponding load balancing strategy is adopted, wherein the load balancing strategy includes random, polling, and minimum load; The ClickHouse client module is used to obtain the number of read rows and execution time of the data analysis task; when the number of read rows of the data analysis task exceeds the maximum number of read rows, and / or the execution time of the data analysis task exceeds the maximum execution time, the data analysis task is stopped.

7. A terminal, characterized in that: The terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the analysis method according to any one of claims 1 to 5 are implemented.

8. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the analysis method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method and device for realizing OLAP analysis based on ClickHouse

    CN112163048A

  • User behavior analysis system and method, storage medium and computing equipment

    CN111488261A

  • Data query method and device, equipment, storage medium and program product

    CN113111083A