Data Processing Method and Apparatus

By aggregating and compressing user behavior detailed data and generating user path data, the problem of large data storage space and low analysis efficiency is solved, and efficient data query and user experience are achieved.

CN115098029BActive Publication Date: 2025-05-27SHANGHAI BILIBILI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210758030.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-05-27
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

As the amount of user behavior detailed data continues to increase, it occupies a large storage space, affects the efficiency of data analysis and affects the user experience.

Method used

By obtaining user behavior detailed data, aggregating processing is performed based on user identification, dividing event attribute information into divided time intervals, format conversion is used to use target data compression structure to generate user path data.

Benefits of technology

Reduces data storage space, improves data query efficiency, supports associated tags and crowd analysis, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115098029B_ABST
    Figure CN115098029B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method and apparatus. The data processing method includes: obtaining user behavior detail data for a target object within a preset historical time interval; aggregating the user behavior detail data based on the user identifiers in the user behavior detail data to obtain initial aggregated data, where the initial aggregated data includes user attribute information and event attribute information; dividing the event attribute information in the initial aggregated data according to a preset divided time interval to obtain target aggregated data, where the target aggregated data includes user attribute information and an event identifier set; performing format conversion on the target aggregated data based on a target data compression structure to obtain user path data of the target object. By processing a large amount of user behavior detail data in this way, the data storage space can be reduced, and the query efficiency can also be improved when querying the compressed data, thereby completing the analysis of user behavior data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a data processing method. This application also relates to a data processing device, a computing device, and a computer-readable storage medium. Background Art

[0002] With the continuous development of computer technology, users have more and more demands for data processing. In order to facilitate relevant technical personnel to understand the usage situation of users in the application, it is possible to perform data analysis on the detailed data of user behaviors collected in the application, that is, to perform path analysis on different behavioral data of users in the application. Furthermore, according to the path analysis results, tasks such as population selection in data analysis can be completed. However, with the continuous increase in the amount of user behavior detailed data, it not only occupies a large amount of storage space, but also affects the efficiency of data analysis and the user experience.

[0003] Therefore, how to reduce the data storage space and improve the analysis efficiency of user behavior data has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0004] In view of this, the embodiments of this application provide a data processing method. This application also relates to a data processing device, a computing device, and a computer-readable storage medium to solve the problem in the prior art that the data storage space is large, which affects the data analysis efficiency.

[0005] According to the first aspect of the embodiments of this application, a data processing method is provided, including:

[0006] Obtain the user behavior detailed data of the target object within a preset historical time interval;

[0007] Based on the user identifier in the user behavior detailed data, perform aggregation processing on the user behavior detailed data to obtain initial aggregation data, where the initial aggregation data includes user attribute information and event attribute information;

[0008] Divide the event attribute information in the initial aggregation data according to a preset divided time interval to obtain target aggregation data, where the target aggregation data includes user attribute information and an event identifier set;

[0009] Based on the target data compression structure, perform format conversion on the target aggregation data to obtain the user path data of the target object.

[0010] According to the second aspect of the embodiments of this application, a data processing device is provided, including:

[0011] A data acquisition module, configured to acquire user behavior detail data for a target object within a preset historical time interval;

[0012] An initial aggregation module, configured to aggregate the user behavior detail data based on the user identifiers in the user behavior detail data to obtain initial aggregation data, where the initial aggregation data includes user attribute information and event attribute information;

[0013] A target aggregation module, configured to divide the event attribute information in the initial aggregation data according to a preset divided time interval to obtain target aggregation data, where the target aggregation data includes user attribute information and an event identifier set;

[0014] A data conversion module, configured to perform format conversion on the target aggregation data based on a target data compression structure to obtain the user path data of the target object.

[0015] According to a third aspect of the embodiments of the present application, a computing device is provided, including a memory, a processor, and computer instructions stored on the memory and executable on the processor. When the processor executes the computer instructions, the steps of the data processing method are implemented.

[0016] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores computer instructions. When the computer instructions are executed by a processor, the steps of the data processing method are implemented.

[0017] The data processing method provided by the present application acquires user behavior detail data for a target object within a preset historical time interval; aggregates the user behavior detail data based on the user identifiers in the user behavior detail data to obtain initial aggregation data, where the initial aggregation data includes user attribute information and event attribute information; divides the event attribute information in the initial aggregation data according to a preset divided time interval to obtain target aggregation data, where the target aggregation data includes user attribute information and an event identifier set; performs format conversion on the target aggregation data based on a target data compression structure to obtain the user path data of the target object.

[0018] In an embodiment of the present application, by aggregating the user behavior detail data of a target object within a preset historical time interval and dividing the user event attribute information in the aggregated data according to a divided time interval to determine target aggregation data, and then performing format conversion on the target aggregation data according to a target data compression structure to determine the user path data of the target object. Through this processing of a large amount of user behavior detail data, the data storage space is reduced. Furthermore, when querying the compressed data, the query efficiency can also be improved, and the analysis of user behavior data is completed. Description of the Drawings

[0019] Figure 1 FIG. 1 is a schematic structural diagram of a system in which a data processing method provided by an embodiment of the present application is applied to a data processing system;

[0020] Figure 2 FIG. 2 is a flowchart of a data processing method provided by an embodiment of the present application;

[0021] Figure 3 FIG. 3 is a schematic diagram of a display interface of a label user path diagram in a data processing method provided by an embodiment of the present application;

[0022] Figure 4 FIG. 4 is a processing flowchart of a data processing method applied to a user path analysis scenario provided by an embodiment of the present application;

[0023] Figure 5 FIG. 5 is a schematic structural diagram of a data processing device provided by an embodiment of the present application;

[0024] Figure 6 FIG. 6 is a structural block diagram of a computing device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0025] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0026] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "said" used in one or more embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to any and all possible combinations of one or more of the associated listed items.

[0027] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0028] First, the noun terms involved in one or more embodiments of this application are explained.

[0029] Traffic funnel and path analysis: Analyze the end point, passing points, and maximum event time interval of the user behavior path for a given expectation, count the number of users for each path, and sort the paths in descending order according to the number of users.

[0030] Funnel: Specifically used to analyze certain user behaviors, starting from its initial behavior and analyzing the behavior of jumping to certain pages. Path analysis is also similar to this kind of requirement; the funnel can also be understood as path analysis, and the funnel is a targeted path analysis for a certain user.

[0031] DWD: (Data Warehouse Detail), also known as the ODS layer, is the isolation layer between the business layer and the data warehouse.

[0032] DWB: (Data Warehouse Base), stores objective data, generally used as an intermediate layer, and can be considered as a data layer with a large number of metrics.

[0033] DWS: (Data Warehouse Service), based on the basic data on DWB, integrates and summarizes the service data for analyzing a certain subject domain, generally a wide table.

[0034] Hive: A data warehouse tool that can map structured data files into a database table and provide SQL query functions. It can transform SQL statements into MapReduce (MapReduce is a high-performance parallel computing platform based on a cluster) tasks for execution.

[0035] ClickHouse (columnar storage database): The full name is Click Stream Data Warehouse; it is a columnar storage database (DBMS: Database Management System) for online analytical processing queries (OLAP: Online Analytical Processing) in the MPP architecture, and can generate analytical data reports in real time using SQL queries.

[0036] Higher-order function: A query method built into the ClickHouse database.

[0037] BitMap technology: Can be understood as a data structure that stores specific data through a bit array; since bit is the smallest unit of data, this data structure is often very space-saving.

[0038] RBM (Roaring Bitmap) Data Structure: Roaring BitMaps (abbreviated as RBM) is a compression algorithm. Bitmap is a commonly used data structure. Bitmap indexes are widely used in databases and search engines and can quickly locate whether a value exists. It is an efficient data compression algorithm that can significantly speed up querying. However, BitMap still occupies a large amount of memory (grows linearly), so generally, BitMap needs to be compressed to reduce memory occupancy and improve efficiency.

[0039] ClickHouse Materialized View: A materialized view is a persistent storage of a query result set, which is completely different from a normal view and is very similar to a table. The implementation of ClickHouse's materialized view is more like a trigger. If an aggregation function is predefined in the view, then (without specifying the populate keyword) the aggregation function only applies to newly inserted data. Changes to the source table data do not change the materialized view, which is one of the unique features of ClickHouse.

[0040] Dictionary Mapping: A dictionary is the only built-in mapping type, and any immutable object can be used as a dictionary key (such as strings, numbers, tuples, etc.).

[0041] Tag: The full name is the user portrait tag. User tags are the core factors that make up the user portrait and are adjectives with differential characteristics generated by analyzing and refining the behavioral data generated by users within the platform.

[0042] Population: That is, the user population. The user population is a user cluster with differential characteristics generated by analyzing and refining the behavioral data generated by users within the platform.

[0043] Population Selection: Based on the tag portrait, a group of people with common user behaviors are selected to facilitate subsequent data analysis.

[0044] To facilitate relevant technical personnel to understand the usage situation of users in the application, it is possible to collect and record the detailed user behavior data generated by users during the use of the application in real time, thereby facilitating subsequent data analysis of user behavior; in the traffic business analysis scenario, according to the detailed user behavior data, the path flow information of all users on the client or web page can be viewed. As the business grows, the quantity of detailed user behavior data is also increasing. Furthermore, the demands for refined analysis of user funnels and paths will gradually increase.

[0045] At present, most data analysis platforms will add the function of funnel analysis. In the industry, when dealing with such scenarios, ClickHouse database is commonly used for funnel analysis. This database can provide a windowFunnel function to implement funnel analysis on detailed data. Path analysis techniques can generally be divided into two types. One is simple path analysis on detailed data, and the other is complex path analysis, also known as intelligent path analysis, which can be performed through high-order functions provided by the ClickHouse database. Although the query performance of the ClickHouse database is very excellent and high-order functions can provide analysis support for most funnel and path analyses, current traffic funnel and path analyses are all based on detailed data. Therefore, in the application program, the following pain points will occur: (1) High consumption of storage resources; with hundreds of billions of incremental data per day, the daily storage volume increases by dozens of terabytes; (2) Slow analysis and query; analysis and query based on detailed data are only accurate to the minute level, resulting in slow data query and poor user experience; (3) Limited functions; only support simple funnel and path analyses, do not support associated tags and populations, and even less support related conversion analysis functions.

[0046] Based on this, a data processing method provided by an embodiment of the present application is a new type of funnel and path analysis. Through technologies such as offline modeling stratification, pre-aggregation at the user path granularity, and storage as an RBM (data compression algorithm) materialized view, hundreds of billions of data per day can be compressed to billions, which optimizes the data query efficiency from the minute level to the second level. In addition, it can support associated tags and populations, facilitating subsequent provision of various conversion query analyses. Specifically, the new funnel and path analysis processes hundreds of billions of detailed data through modeling stratification, then aggregates using the user identifier through dimension pruning. User identifiers on the same path are aggregated together, and then the RBM data type is introduced to store the aggregated devices, and finally it is stored in a hive table (database table). Import the hive data into ClickHouse. At this stage, from the perspective of table structure design, the materialized view and RBM data structure of ClickHouse are adopted, greatly reducing the storage. The query also achieves second-level query, and functions such as label profiling and population selection are also realized.

[0047] In the present application, a data processing method is provided. The present application also relates to a data processing device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.

[0048] See Figure 1 , Figure 1 FIG. shows a schematic system structure diagram of the data processing method provided by an embodiment of the present application applied to a data processing system.

[0049] Figure 1In the middle is the data processing system 100 to which the data processing method provided by the embodiment of the present application is applied. Among them, the data processing system 100 includes a data warehouse 102 and a database 104. The data warehouse 102 includes a data detail layer, a data basic layer, and a data service layer.

[0050] It should be noted that the data warehouse can collect the user behavior detail data in the application program or web page, and through the hierarchical structure in each data warehouse, layer-process the user behavior detail data, and store the processed data in the database; the database can store the data output by the data warehouse and support the data query function. At the same time, in this embodiment, the type of the database is not specifically limited, including but not limited to the ClickHouse database.

[0051] In practical applications, the data processing system 100 can be understood as the server corresponding to the data analysis platform, and use the data warehouse 102 to preprocess the user behavior detail data in the application program, and store the processed user behavior data in the database 104, so that when the data processing system 100 receives a data query request subsequently, it can directly implement a fast data query operation in the database 104; specifically, the data detail layer in the data warehouse 102 can obtain all the user behavior detail data of the application program within a preset historical time interval, and input it to the data basic layer of the data warehouse 102, and use the data basic layer to aggregate the user behavior detail data, and then input the aggregated user behavior detail data to the data service layer to complete the summary of the user path data of the user behavior detail data. At the same time, the user path data can also be compressed, and then the compressed user path data can be stored in the database 104.

[0052] In summary, the data processing method provided by this embodiment can not only reduce the storage space of user behavior data by aggregating, summarizing, and compressing the user behavior detail data of the application program, but also improve the data query efficiency by querying the compressed stored data.

[0053] Figure 2 The flowchart of a data processing method according to an embodiment of the present application is shown, which specifically includes the following steps:

[0054] Step 202: Obtain the user behavior detail data of the target object within a preset historical time interval.

[0055] Among them, the target object can be understood as the object that outputs the detailed user behavior data. For example, application software, web pages, etc.; the preset historical time interval can be understood as the time interval preset within the historical period. For example, the preset historical time interval is from June 5th to June 6th, 2021; the detailed user behavior data refers to all user behavior data generated by the user using the target object. For example, the behavior data of the user browsing the page on the application program, the behavior data of clicking on the page link, etc.

[0056] In practical applications, the DWD layer in the data warehouse can obtain the detailed user behavior data in the application program within 24 hours from 00:00 on June 1st, 2021 to 24:00 on June 1st, 2021. The detailed user behavior data includes the behavior data of user 1 browsing the shopping page, the behavior data of user 2 clicking on the commodity order, etc.

[0057] Through the data detail layer in the data warehouse, obtain the detailed user behavior data of the target object within the preset historical time interval, which is convenient for further aggregation processing of the detailed user behavior data within the preset historical time interval.

[0058] Step 204: Based on the user identifiers in the detailed user behavior data, perform aggregation processing on the detailed user behavior data to obtain initial aggregation data, where the initial aggregation data includes user attribute information and event attribute information.

[0059] Among them, the initial aggregation data can be understood as the aggregation data after the data basic layer in the data warehouse performs aggregation processing on the detailed user behavior data in the data detail layer, including two types of aggregation data: user attribute information and event attribute information.

[0060] User attribute information can be understood as user behavior data related to the user, including user identifier data, user device usage data, system type data used by the user device, and user device model data, etc. For example, [user 1, device 1, system 1, model 2], [user 2, device 2, system 1, model 1]; event attribute information can be understood as user behavior data related to the events that occur to the user in the application program, including event identifier data of the events that occur to the user in the application program and time data triggered by the events, etc. For example, [event 1, time 1], [event 2, time 2]; the specific attribute fields included in the user attribute information and event attribute information in the embodiments of the present application are not specifically limited, but in this embodiment, the above-mentioned included data content is used as an example to illustrate the solution.

[0061] In practical applications, since a large amount of user behavior detail data will occupy a large storage space, the DWB layer in the data warehouse can aggregate a large amount of user behavior detail data to obtain a lightly summarized table for user path analysis. Specifically, the user identification in the user behavior detail data can be used as the aggregation granularity to aggregate the user behavior detail data, obtain the initial aggregated data and store it in a Hive table. At the same time, the initial aggregated data can include two types of data, namely user attribute information and event attribute information.

[0062] It should be noted that in the data processing method provided in the embodiments of the present application, the user attribute information includes at least one of user identification information, user device information, device system information, and device model information.

[0063] For example, the user identification information can be the user's ID number, such as 001, 002, etc.; the user device information can be mobile devices, computer devices, tablet devices, etc.; the device system information can be IOS system, Android system, etc.; the device model information can be model A, model B, etc.

[0064] Furthermore, based on the user identification in the user behavior detail data, aggregating the user behavior detail data to obtain the initial aggregated data includes:

[0065] Determine the initial event attribute information corresponding to the event type in the user behavior detail data;

[0066] Trim the initial event attribute information to obtain the target user behavior detail data;

[0067] Based on the user identification in the user behavior detail data, aggregate the target user behavior detail data to obtain the initial aggregated data.

[0068] Among them, the initial event attribute information can be understood as all user behavior data corresponding to the event type in the user behavior detail data, including the user's browsing event identifier, browsing event time, browsing event execution times, exposure event identifier, exposure event time, exposure event execution times, etc.

[0069] In practical applications, since it is necessary to perform path analysis on user behavior data, dimensionality reduction can be performed on the user behavior detail data, thereby reducing the user detail data irrelevant to path analysis and reducing the data storage space. Specifically, first determine the initial event attribute information corresponding to the event type in the user behavior detail data. The initial event attribute information is the detail data corresponding to all attribute fields strongly related to the event type, and perform a reduction operation on the initial event attribute information to obtain the target user behavior detail data. Furthermore, taking the user identifier in the user behavior detail data as the aggregation granularity, perform an aggregation process on the target user behavior detail data to obtain the initial aggregation data.

[0070] Furthermore, when performing the reduction on the initial event attribute information, it is necessary to reduce the detail data that changes relatively frequently in the user behavior detail data because in the user path analysis, the detail data that changes relatively frequently may not be representative of user behavior. Specifically, reducing the initial event attribute information to obtain the target user behavior detail data includes:

[0071] Determine the event attribute information and the attribute information to be reduced in the initial event attribute information;

[0072] Retain the event attribute information in the user behavior detail data, and reduce the attribute information to be reduced in the user behavior detail data to obtain the target user behavior detail data.

[0073] Among them, the attribute information to be reduced can be understood as the detail data corresponding to the attribute fields that need to be dimensionally reduced in the user behavior detail data related to the event type, except for the event identification information and the event time information.

[0074] In practical applications, the DWB layer of the data warehouse reduces the user behavior detail data, retains the event attribute information in the user behavior detail data according to the user identifier granularity, and reduces the attribute information to be reduced in the user behavior detail data to obtain the target user behavior detail data.

[0075] Furthermore, determining the event attribute information and the attribute information to be reduced in the initial event attribute information includes:

[0076] Determine the event identification information and the event time information in the initial event attribute information as the event attribute information;

[0077] Determine the other event attribute information in the initial event attribute information except the event identification information and the event time information as the attribute information to be reduced.

[0078] In practical applications, functions can be called in the database to perform some clipping on query dimensions, discard the detailed data corresponding to some attribute information with relatively frequent changes, and retain the detailed data corresponding to the attribute information at the user granularity. Specifically, when implementing, the event identification information and event time information in the initial event attribute information are retained as the event attribute information, while the detailed data corresponding to other attribute information except the event identification information and event time information in the initial time attribute information is used as the attribute information to be clipped, and the data clipping operation is performed.

[0079] Continuing with the above example, if the initial event attribute information includes the user's browsing event identification, browsing event time, browsing event execution times, exposure event identification, exposure event time, and exposure event execution times, then the user's browsing event identification, browsing event time, exposure event identification, and exposure event time can be used as the event attribute information; the browsing event execution times and exposure event execution times are used as the attribute information to be clipped (the reported number parameters may change relatively frequently. For example, the user may repeatedly operate the browsing or exposure event multiple times in a short period).

[0080] It should be noted that the attribute information to be clipped can be understood as some attribute information with relatively frequent changes. In the above embodiments, the event execution times are used as an example of the attribute information to be clipped for the description of the clipping operation, but no limitation is made thereto.

[0081] After the DWB layer of the data warehouse performs dimension clipping, it can aggregate the target user behavior detailed data after dimension clipping according to the user identification granularity to reduce the data volume of the detailed data. Specifically, based on the user identification in the user behavior detailed data, the target user behavior detailed data is aggregated to obtain the initial aggregated data, including:

[0082] Determine the target user identification in the user behavior detailed data;

[0083] Determine the target user behavior sub-data in the target user behavior detailed data according to the target user identification;

[0084] Perform deduplication processing on the user attribute information in the target user behavior sub-data to obtain the user attribute information corresponding to the target user identification;

[0085] Perform aggregation processing on the event attribute information according to the user attribute information to obtain the event attribute information corresponding to the user attribute information;

[0086] Concatenate the user attribute information and the event attribute information to obtain the initial aggregated data corresponding to the target user identification.

[0087] In practical applications, the DWB layer of the data warehouse can determine the target user identifier from the user behavior detail data, and based on this target user identifier, determine the user target behavior sub-data corresponding to the target user identifier in the target user behavior detail data. This process can be understood as an operation process of screening the detail data corresponding to the target user identifier. After screening with each target user identifier as the granularity, duplicate removal processing can be performed on the user attribute information in the target user behavior sub-data, and the duplicate user attribute information is deleted to ensure that each piece of user attribute information is different. For example, three pieces of user attribute information are respectively the first one: [User 1, Device 1, System 1, Model 2], the second one: [User 1, Device 2, System 1, Model 1], and the third one: [User 1, Device 1, System 1, Model 2]. After performing the duplicate removal operation on these three pieces of user attribute information, two pieces of user attribute information can be obtained, namely [User 1, Device 1, System 1, Model 2] and [User 1, Device 2, System 1, Model 1]. Therefore, these two pieces of user attribute information are the user attribute information corresponding to User 1.

[0088] Furthermore, in the target user behavior sub-data, in addition to the various user attribute information corresponding to the target user identifier, it also includes the corresponding event attribute information. Therefore, taking the user attribute information as the granularity, aggregation processing is performed on the event attribute information to obtain the event attribute information corresponding to each piece of user attribute information. For example, for the first piece of user attribute information [User 1, Device 1, System 1, Model 2], the corresponding event attribute information is [Event 1|Time 1, Event 2|Time 2, Event 2|Time 3, Event 2|Time 3, Event 3|Time 4, Event 1|Time 2]. Finally, a concatenation operation can also be performed on the user attribute information and the event attribute information to obtain the initial aggregated data corresponding to the target user identifier. For example, [User 1, Device 1, System 1, Model 2], [Event 1|Time 1, Event 2|Time 2, Event 2|Time 3, Event 2|Time 3, Event 3|Time 4, Event 1|Time 2].

[0089] It should be noted that in this embodiment, the aggregation steps are described by taking one user identifier as an example. Therefore, for the detail data corresponding to each user identifier in the user behavior detail data, the above data aggregation method is referred to.

[0090] The data processing method provided by the embodiment of the present application completes the reprocessing of the user behavior detail data by performing duplicate removal on the user attribute information in the detail data corresponding to each user identifier and aggregating the event attribute information with the user attribute information as the granularity, reducing the data volume and retaining the detail data that can complete the user path analysis.

[0091] In addition, when aggregating event attribute information, event identification information can also be concatenated according to the event time information in the event attribute information, so as to obtain the event attribute information corresponding to the user attribute information. Specifically, the event attribute information includes event identification information and event time information;

[0092] Performing an aggregation process on the event attribute information according to the user attribute information to obtain the event attribute information corresponding to the user attribute information includes:

[0093] Concatenating the event identification information based on the event time information to obtain the event attribute information corresponding to the user attribute information.

[0094] Continuing with the above example, the event attribute information is [Event 1|Time 1, Event 2|Time 2, Event 2|Time 3, Event 2|Time 3, Event 3|Time 4, Event 1|Time 2]. Furthermore, according to the time sequence of the event time information, the event identifications of the corresponding user behavior events can be rearranged. For example, if the event time sequence is [Time 1 - Time 2 - Time 3 - Time 4], then the event identifications are correspondingly adjusted to [Event 1 - Event 2, Event 1 - Event 2, Event 2 - Event 3], that is, the event attribute information is [Event 1|Time 1 - Event 2, Event 1|Time 2 - Event 2, Event 2|Time 3 - Event 3|Time 4].

[0095] Furthermore, the initial aggregated data corresponding to the target user identification can be [User 1, Device 1, System 1, Model 2], [Event 1|Time 1 - Event 2, Event 1|Time 2 - Event 2, Event 2|Time 3 - Event 3|Time 4].

[0096] In addition, when determining the initial aggregated data corresponding to the target user identification, interference filtering processing can also be performed on events repeatedly operated by a certain user at the same time to delete event data that interferes with user behavior. Specifically, concatenating the user attribute information and the event attribute information to obtain the initial aggregated data corresponding to the target user identification includes:

[0097] Determining the event identification that satisfies the preset interference condition in the event attribute information as the interference event identification, where the preset interference condition is the condition of behavior events repeatedly occurring by the user within a preset time interval;

[0098] Deleting the interference event identification in the event attribute information and the event time corresponding to the interference event identification;

[0099] Concatenating the event attribute information after the deletion operation with the user attribute information to obtain the initial aggregated data corresponding to the target user identification.

[0100] In practical applications, the data warehouse can also delete the event identifiers corresponding to interference events in the event attribute information, as well as delete the event times corresponding to the interference event identifiers. It should be noted that a behavior event that occurs repeatedly within a preset time interval for an event corresponding to user behavior can be understood as an interference event. For example, if a user performs a click operation three times at the same time, then an event that repeats within a short period can be considered an interference event, and interference event filtering can be performed through means such as deduplication; further, the event attribute information after the deletion operation for performing interference event filtering is spliced with the user attribute information to obtain the initial aggregated data corresponding to the final target user identifier; Continuing with the above example, the initial aggregated data can be, [User 1, Device 1, System 1, Model 2], [Event 1|Time 1 - Event 2, Event 1|Time 2 - Event 2|Time 3 - Event 3|Time 4], that is, one of the two Event 2s corresponding to Time 3 is deleted, and only one Event 2 is retained.

[0101] It should be noted that interference event filtering can be performed at the DWB layer of the data warehouse or at the DWS layer of the data warehouse. The embodiments of the present application do not make specific limitations in this regard.

[0102] Step 206: Divide the event attribute information in the initial aggregated data according to a preset divided time interval to obtain target aggregated data, where the target aggregated data includes user attribute information and an event identifier set.

[0103] Further, the target aggregated data is aggregated data obtained by reprocessing the initial aggregated data.

[0104] Among them, the event identifier set can be understood as a set of event identifiers corresponding to user behavior data within a preset divided time interval. For example, the event identifier set is [Event 1, Event 2, Event 4].

[0105] In practical applications, the DWS layer of the data warehouse can divide the event attribute information in the initial aggregated data according to a preset divided time interval to obtain target aggregated data, where the preset divided time interval can be understood as 30 min, 1 h, 2 h, etc. In this embodiment, no specific limitation is made on the specific divided time interval, but it is associated with the time interval for querying user path data in the front-end application.

[0106] Further, dividing the event attribute information in the initial aggregated data according to a preset divided time interval to obtain target aggregated data includes:

[0107] According to a preset divided time interval, divide the event time information in the initial aggregated data to obtain at least one event time interval;

[0108] Determine the event identifiers corresponding to each event time interval as an event identifier set, and generate target aggregation data based on the event identifier set and the user attribute information.

[0109] In practical applications, divide the event time information in the initial aggregation data. For example, if the event time information is [time1 - time2 - time3 - time4], divide the event time information according to a preset division time interval, then each event time interval is [time1 - time2], [time3 - time4]; further, the event identifiers corresponding to each event time interval can be determined as an event identifier set, that is, the event 1 corresponding to time1, the event 2 corresponding to time2, and event 1 are determined as the event identifier set [event1, event2, event1]; according to the event 2 corresponding to time3 and the event 3 corresponding to time4, determine the event identifier set as [event2, event3].

[0110] Further, according to the event identifier set and the user attribute information, target aggregation data can be generated; that is, [event1, event2, event1], [user1, device1, system1, model2]; [-1, -1, event2, event3], [user1, device1, system1, model2]. It should be noted that padding operations can be performed on the event identifier set to process the event identifier set for subsequent data query execution.

[0111] It should be noted that the processing process of the initial aggregation data and the processing process of the target aggregation data mentioned above in the embodiments of the present application can both be understood as the pre-aggregation process of the user behavior detail data. In this embodiment, in order to achieve data storage compression and improve the efficiency of subsequent data query at the same time, the user behavior detail data to be stored can be pre-aggregated, and in the face of a large amount of user behavior detail data, the pre-aggregation method is also the key point reflected in the embodiments of the present application.

[0112] Step 208: Perform format conversion on the target aggregation data based on the target data compression structure to obtain the user path data of the target object.

[0113] Among them, the target data compression structure refers to the data result that can compress data. For example, the BitMap data structure; format conversion refers to converting the aggregation data into a format corresponding to the target data compression structure; user event data refers to the data obtained after performing format conversion on the aggregation data based on the target data compression structure.

[0114] In practical applications, performing format conversion on the target aggregation data based on the target data compression structure to obtain the user path data of the target object includes:

[0115] Convert the user identifier in the user attribute information of the target aggregated data based on the target data compression structure to obtain the user path data of the target object.

[0116] Specifically, determine the user identifier in the user attribute information of the aggregated data; convert the user identifier based on the target data compression structure to obtain the user identifier of the target data compression structure; and form the user path data corresponding to the target object from the user identifier of the target data compression structure, the user attribute information other than the user identifier in the user attribute information, and the event attribute information.

[0117] In a specific embodiment of the present application, determine the aggregated data K and the target data compression structure BitMap; determine the data dictionary corresponding to the target data compression structure BitMap, and convert the user identifier in the user attribute information in the aggregated data K into a BitMap data structure based on the data dictionary; splice the user identifier of the BitMap data structure and the user attribute information and event attribute information other than the user identifier in the aggregated data K into the user path data corresponding to the application program within a preset time interval.

[0118] By converting the user identifier in the aggregated data into the target data compression structure and then generating user event data based on the user identifier of the target data compression structure, further compression of the aggregated data is achieved, thereby further reducing the user behavior detail data.

[0119] In practical applications, in order to facilitate data analysis based on tags, after converting the format of the target aggregated data based on the target data compression structure to obtain the user path data of the target object, it is also possible to obtain the tag data of the target object, thereby facilitating subsequent data analysis. The specific method includes:

[0120] Obtain the user identifier of the target object and the attribute tags corresponding to the user identifier;

[0121] Convert the format of the user identifier based on the target data compression structure, and determine the user tag data of the target object based on the converted user identifier and the attribute tags.

[0122] Among them, the attribute tag refers to the tag field corresponding to the user identifier. For example, the tags selected by user A when registering the application program are "animation" and "entertainment", that is, the attribute tags of user A are "animation" and "entertainment"; the user tag data refers to the data composed of the attribute tags and the converted user identifier.

[0123] In a specific embodiment of the present application, the user identifier in the application program and the attribute label corresponding to each user identifier are obtained. Specifically, the user identifier "s1" and the attribute label "animation, movie, food" corresponding to the user identifier "s1" are obtained; the data dictionary corresponding to the BitMap data structure, and the user identifier "s1" is mapped to the BitMap data structure based on the data dictionary; the user label data is composed of the attribute label "animation, movie, food" corresponding to the user identifier "s1" and the user identifier "s1" of the BitMap data structure.

[0124] By determining the user identifier of the target object and obtaining the attribute label corresponding to the user identifier; performing format conversion on the user identifier, and generating user label data based on the converted user identifier and attribute label, thereby enriching the data for data analysis.

[0125] In practical applications, after performing format conversion on the target aggregated data based on the target data compression structure to obtain the user path data of the target object, it further includes:

[0126] Storing the user path data of the target object in a database;

[0127] Correspondingly, after determining the user label data of the target object based on the converted user identifier and the attribute label, it further includes:

[0128] Storing the user label data of the target object in a database.

[0129] Among them, the database refers to a database that can store the target data compression structure. For example, the ClickHouse database; specifically, in the case where the database is the ClickHouse database, the data query result executed in the ClickHouse database can be stored based on the ClickHouse materialized view, thereby improving the data query efficiency.

[0130] By storing the compressed user label data and user path data in the database instead of the user detail data, the database storage space is saved, and due to the reduction of the data volume, the subsequent data analysis efficiency is improved.

[0131] After storing the user path data and user label data in the database, the function in the database can be called to complete the query of the user path data; specifically, in a data processing method provided by an embodiment of the present application, after storing the user path data of the target object in the database, it further includes:

[0132] Receiving a user path data query request for the target object, where the user path data query request carries a basic configuration query condition;

[0133] Query corresponding user path data in the user path data in the database based on the basic configuration query conditions, where the basic configuration query conditions include at least one of an event time condition, a central event condition, and a user device condition;

[0134] Generate a user path map based on the user path data and send the user path map to the user path map display interface of the target object.

[0135] Among them, a user path data query request refers to a request to query user path data that meets the query conditions in the database; a basic configuration query condition refers to a condition for querying user path data in the database, and the basic configuration query conditions include at least one of an event time condition, a central event condition, and a user device condition; a user path map refers to a user path map generated based on user path data.

[0136] In practical applications, after receiving a user path data query request for a target object, determine the basic configuration query conditions in the user path data query request; filter the user path data that meets the basic configuration query conditions in the user path data of the database according to the basic configuration query conditions; generate a user path map based on the user path data and send the user path map to the user path map display interface of the application program, such as the display interface of a computer device to display the user path map.

[0137] In a specific embodiment of the present application, the server receives a user path data query request with a central event of "target play page view"; based on the event time condition, user device condition, etc. in the basic configuration query conditions, query the target user path data corresponding to the basic configuration query conditions in the user path data of the database; generate a user path map based on the target user path data and send the user path map to the user path map display interface.

[0138] By querying data in the user path data of the database based on the basic configuration query conditions, since the data stored in the database is compressed user path data and the number of data is less than that of user behavior detail data, the processing efficiency of the query request can be improved.

[0139] Specifically, for the data processing method provided in another embodiment of the present application, after storing the user label data of the target object in the database, it further includes:

[0140] Receive a user path data query request for a target object, where the user path data query request carries basic configuration query conditions and label data query conditions;

[0141] Determine the to-be-processed user path data from the user path data in the database based on the basic configuration query condition, and determine the to-be-processed user label data from the user label data in the database based on the label data query condition;

[0142] Process the to-be-processed user path data and the to-be-processed user label data according to a preset data processing method to obtain labeled user path data;

[0143] Generate a labeled user path graph based on the labeled user path data and send the labeled user path graph to the user path graph display interface of the target object.

[0144] Among them, the user path data query request refers to a request to query user path data and user label data that meet the query conditions in the database; the label data query condition refers to the condition for querying user label data in the database. For example, corresponding label data such as post-90s, female, anime, and entertainment can be input in the label query condition box on the front-end interface of the application.

[0145] In practical applications, after the database receives a user path data query request for a certain application, the query request carries a basic configuration query condition and a label data query condition; furthermore, based on the basic configuration query condition, determine the to-be-processed user path data from the user path data in the database, where the to-be-processed user path data is the user path data filtered in the database according to the basic configuration query condition, facilitating subsequent intersection and union calculation processing of the to-be-processed user path data; then, based on the label data query condition, determine the to-be-processed user label data from the labeled user path data in the database, facilitating subsequent intersection and union calculation processing of the to-be-processed user label data; it should be noted that the user path data and user label data stored in the database, since their user identifiers have been converted to the RBM storage structure, therefore, subsequently, according to the preset data processing method, process the to-be-processed user path data and the to-be-processed user label data to obtain labeled user path data, and then, generate a labeled user path graph based on the labeled user path data, and send the labeled user path graph to the user path graph display interface of the application.

[0146] See Figure 3 , Figure 3 shows a schematic diagram of the display interface of the labeled user path graph in the data processing method provided by the embodiment of the present application.

[0147] Figure 3 The part of the selection path analysis event and configuration conditions in [[ ]] can be understood as the input and selection boxes for the basic configuration query condition and the label data query condition. After the user determines these two query conditions, they can click Figure 3The "Query" button in the interface, and then it can display for the user Figure 3 The user path diagram in the lower half of Figure 3 . Among them, this user path diagram can be understood as the user path diagram queried in the database according to the above query conditions, which is convenient for directly providing the basis for data analysis for relevant personnel according to this user path diagram in the follow-up. In addition, the above two query conditions can also be stored by clicking the "Save" button, which is convenient for quickly querying the corresponding query conditions in the follow-up.

[0148] Furthermore, according to the preset data processing method, the to-be-processed user path data and the to-be-processed user label data are processed to obtain labeled user path data, including:

[0149] Determine the association relationship between the basic configuration query conditions and the label data query conditions, and determine the preset data processing method based on the association relationship;

[0150] Process the to-be-processed user path data and the to-be-processed user label data based on the preset data processing method to obtain labeled user path data.

[0151] Among them, the preset data processing method refers to the way of mutual calculation of data determined according to the two query conditions, such as intersection calculation, union calculation, etc.

[0152] In practical applications, the association relationship between the basic configuration query conditions and the label data query conditions can be determined, and the processing method executed on the data filtered by these two conditions can be determined, that is, the intersection calculation method or the union calculation method of the data set. Furthermore, according to the determined preset data processing method, intersection calculation or union calculation, etc. is performed on the to-be-processed user path data and the to-be-processed user label data to obtain labeled user path data.

[0153] In summary, the data processing method provided in the embodiment of the present application processes a large amount of user behavior detail data to obtain data in the RBM data structure, which is pre-stored in the database. It can not only compress storage and reduce memory space, but also improve data query efficiency; at the same time, fusing user behavior data and user label data helps to determine the user path data of the population corresponding to the label according to the query label data in the follow-up, thereby realizing accurate population selection.

[0154] See Figure 4 , Figure 4 shows a processing flow chart of a data processing method applied to the user path analysis scenario provided by an embodiment of the present application, which specifically includes the following steps:

[0155] Step 402: The offline processing application APP in the DWD layer of the data warehouse processes the user behavior APP-side data detail table of one hundred billion (120 billion); for specific detail data, seeFigure 4 The detailed data illustrated by the DWD layer in the middle.

[0156] Step 404: The DWB layer of the data warehouse performs dimensionality clipping on the detailed data and aggregates the data according to the user identifier; for the aggregated data, please refer to Figure 4 the aggregated data illustrated by the DWB layer in the middle.

[0157] It should be noted that the detailed data, i.e., traffic data, can be divided into private parameters (detailed data corresponding to the attribute fields associated with the event type) and public parameters (detailed data corresponding to the attribute fields associated with the user). Among them, the public parameters do not change frequently at the user granularity. Since functions in the hive table can be used to perform query dimensionality clipping, some private parameters that change frequently are discarded, and the public parameters at the user granularity are retained. And through the buvid (user identifier) granularity for aggregation, all events with the same buvid are concatenated and aggregated into one field according to the timeline. After aggregation, the data forms the DWB layer and is landed in the hive table.

[0158] Step 406: The DWS layer of the data warehouse filters the aggregated data for interference events and implements compressed storage according to the RBM storage structure type; for the data processed after interference event filtering, please refer to Figure 4 the aggregated data illustrated by the DWS layer in the middle.

[0159] Based on the summary of the paths of the data in the DWB layer, the buvids of the same path are summarized and aggregated into an array structure. Many interference events occur in this process. For example, some paths will appear frequently and out of order, interfering with the real user behavior. Therefore, interference event filtering can be performed through means such as deduplication, and the aggregated data illustrated by the DWS layer in the middle can be obtained; among them, Figure 4 the event identifier set of [Event 1, Event 2, Event 4] in the middle can be understood as the event string aggregated for the events corresponding to the same user granularity within the preset event time interval; Figure 4 the "-1" in the event string in the middle can be understood as an operation for storage padding, and no further limitations are imposed on this. Figure 4

[0160] In addition, the RBM data storage structure is introduced. The aggregated user path data is converted in format according to the RBM storage structure and finally lands in the hive table. The entire process is implemented through spark scripts using code and algorithms, and no further explanations and limitations are made in this embodiment.

[0161] ​Step 408: The DWS layer of the data warehouse can use the outbound script to import hive data into ClickHouse for the user path data. Optimization processing is also carried out at this stage. In terms of the ClickHouse table structure design, the materialized view technology and RBM data structure of ClickHouse are adopted, and the storage is greatly compressed by using the array to materialize RBM.

[0162] Step 410: After the database receives the query request for the user path data, since the data volume of hundreds of billions of detailed data has been aggregated and compressed to billions, furthermore, the ClickHouse query engine can achieve query in seconds.

[0163] In addition, when the query request for the user path data also carries a tag query condition, conversion analysis functions such as tag portrait and population selection can be realized through the intersection and union calculation of Bitmap. It should be noted that since both the user path data and the user tag data can introduce the RBM data storage structure, the intersection and union calculation of the Bitmap storage structure can be realized.

[0164] It should be noted that the user tag data can obtain the tag data that can realize the intersection and union calculation with the user path data by obtaining the user identifier and the tag data corresponding to the user identifier in the application program and introducing the RBM data storage structure.

[0165] In summary, the new funnel and path analysis provided by the embodiments of the present application, by modeling and stratifying the data, compared with the previous processing of hundreds of billions of detailed data, realizes the compression of data offline in the DWB and DWS layers, pre-stores the data summarized by the data warehouse in the database to replace the original detailed data, reduces the data storage space, and improves the data query efficiency.

[0166] Corresponding to the above method embodiments, the present application also provides an embodiment of a data processing device. Figure 5 The structural schematic diagram of a data processing device provided by an embodiment of the present application is shown. As Figure 5 shown, the device includes:

[0167] A data acquisition module 502, configured to acquire user behavior detailed data of a target object within a preset historical time period;

[0168] An initial aggregation module 504, configured to perform aggregation processing on the user behavior detailed data based on the user identifier in the user behavior detailed data to obtain initial aggregation data, where the initial aggregation data includes user attribute information and event attribute information;

[0169] The target aggregation module 506 is configured to divide the event attribute information in the initial aggregation data according to a preset divided time interval to obtain target aggregation data, where the target aggregation data includes user attribute information and an event identifier set;

[0170] The data conversion module 508 is configured to perform format conversion on the target aggregation data based on the target data compression structure to obtain the user path data of the target object.

[0171] Optionally, the initial aggregation module 504 is further configured to:

[0172] Determine the initial event attribute information corresponding to the event type in the user behavior detail data;

[0173] Trim the initial event attribute information to obtain target user behavior detail data;

[0174] Perform an aggregation process on the target user behavior detail data based on the user identifier in the user behavior detail data to obtain initial aggregation data.

[0175] Optionally, the initial aggregation module 504 is further configured to:

[0176] Determine event attribute information and attribute information to be trimmed in the initial event attribute information;

[0177] Retain the event attribute information in the user behavior detail data, and trim the attribute information to be trimmed in the user behavior detail data to obtain target user behavior detail data.

[0178] Optionally, the initial aggregation module 504 is further configured to:

[0179] Determine the event identifier information and event time information in the initial event attribute information as event attribute information;

[0180] Determine other event attribute information in the initial event attribute information except the event identifier information and the event time information as the attribute information to be trimmed.

[0181] Optionally, the initial aggregation module 504 is further configured to:

[0182] Determine a target user identifier in the user behavior detail data;

[0183] Determine target user behavior sub-data in the target user behavior detail data according to the target user identifier;

[0184] Perform deduplication processing on the user attribute information in the target user behavior sub-data to obtain the user attribute information corresponding to the target user identifier;

[0185] Perform aggregation processing on the event attribute information according to the user attribute information to obtain the event attribute information corresponding to the user attribute information;

[0186] Concatenate the user attribute information and the event attribute information to obtain the initial aggregation data corresponding to the target user identifier.

[0187] Optionally, the event attribute information includes event identifier information and event time information;

[0188] Performing aggregation processing on the event attribute information according to the user attribute information to obtain the event attribute information corresponding to the user attribute information includes:

[0189] Concatenate the event identifier information based on the event time information to obtain the event attribute information corresponding to the user attribute information.

[0190] Optionally, the initial aggregation module 504 is further configured to:

[0191] Determine that the event identifier that satisfies the preset interference condition in the event attribute information is an interference event identifier, where the preset interference condition is the condition of behavior events that occur repeatedly by the user within a preset time interval;

[0192] Delete the interference event identifier in the event attribute information and the event time corresponding to the interference event identifier;

[0193] Concatenate the event attribute information after the deletion operation with the user attribute information to obtain the initial aggregation data corresponding to the target user identifier.

[0194] Optionally, the target aggregation module 506 is further configured to:

[0195] Divide the event time information in the initial aggregation data according to a preset divided time interval to obtain at least one event time interval;

[0196] Determine the event identifier corresponding to each event time interval as an event identifier set, and generate target aggregation data based on the event identifier set and the user attribute information.

[0197] Optionally, the data conversion module 508 is further configured to:

[0198] Based on the target data compression structure, perform format conversion on the user identifier in the user attribute information of the target aggregation data to obtain the user path data of the target object.

[0199] Optionally, the device further includes:

[0200] A label data determination module, configured to obtain a user identifier for a target object and an attribute label corresponding to the user identifier;

[0201] Perform format conversion on the user identifier based on a target data compression structure, and determine user label data of the target object based on the converted user identifier and the attribute label.

[0202] Optionally, the device further includes:

[0203] A data storage module, configured to store the user path data of the target object in a database;

[0204] Optionally, the data storage module is further configured to:

[0205] Store the user label data of the target object in a database.

[0206] Optionally, the device further includes:

[0207] A data query module, configured to receive a user path data query request for a target object, where a basic configuration query condition is carried in the user path data query request;

[0208] Query corresponding user path data from the user path data in the database based on the basic configuration query condition, where the basic configuration query condition includes at least one of an event time condition, a central event condition, and a user device condition;

[0209] Generate a user path map based on the user path data, and send the user path map to a user path map display interface of the target object.

[0210] Optionally, the data query module is further configured to:

[0211] Receive a user path data query request for a target object, where a basic configuration query condition and a label data query condition are carried in the user path data query request;

[0212] Determine to-be-processed user path data from the user path data in the database based on the basic configuration query condition, and determine to-be-processed user label data from the user label data in the database based on the label data query condition;

[0213] Process the to-be-processed user path data and the to-be-processed user label data according to a preset data processing method to obtain labeled user path data;

[0214] Generate a labeled user path graph based on the labeled user path data and send the labeled user path graph to the user path display interface of the target object.

[0215] Optionally, the data query module is further configured to:

[0216] Determine the association relationship between the basic configuration query condition and the labeled data query condition, and determine a preset data processing method based on the association relationship;

[0217] Process the to-be-processed user path data and the to-be-processed user labeled data based on the preset data processing method to obtain labeled user path data.

[0218] Optionally, the user attribute information includes at least one of user identification information, user device information, device system information, and device model information.

[0219] Optionally, the target aggregation data is aggregation data obtained by reprocessing the initial aggregation data.

[0220] In summary, the data processing device provided by the embodiments of the present application processes a large amount of user behavior detail data to obtain data in the RBM data structure, which is pre-stored in the database. This can not only compress storage and reduce memory space, but also improve data query efficiency. At the same time, fusing user behavior data with user labeled data helps to subsequently determine the user path data of the population corresponding to the label according to the queried labeled data, thereby achieving accurate population selection.

[0221] The above is a schematic solution of a data processing device according to an embodiment of the present application. It should be noted that the technical solution of the data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.

[0222] Figure 6 FIG. shows a structural block diagram of a computing device 600 according to an embodiment of the present application. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 through a bus 630, and a database 650 is used to store data.

[0223] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0224] In one embodiment of the present application, the above components of the computing device 600 and Figure 6 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 6 the block diagram of the computing device shown is only for illustrative purposes and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.

[0225] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 600 can also be a mobile or stationary server.

[0226] Wherein, when the processor 620 executes the computer instructions, the steps of the data processing method are implemented.

[0227] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above data processing method belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above data processing method.

[0228] One embodiment of the present application also provides a computer-readable storage medium, which stores computer instructions, and when the computer instructions are executed by a processor, the steps of the data processing method as described above are implemented.

[0229] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above data processing method belong to the same concept. For the details not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above data processing method.

[0230] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0231] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0232] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0233] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0234] The preferred embodiments of the present application disclosed above are only used to help illustrate the present application. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the present application. These embodiments are selected and specifically described in the present application to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is only limited by the claims and their full scope and equivalents.

Claims

1. A data processing method, characterized in that, comprising: obtaining user behavior detail data of a target object within a preset historical time interval; aggregating the user behavior detail data based on the user identifiers in the user behavior detail data to obtain initial aggregated data, wherein the initial aggregated data includes user attribute information and event attribute information; dividing the event attribute information in the initial aggregated data according to a preset divided time interval to obtain target aggregated data, wherein the target aggregated data includes user attribute information and an event identifier set; performing format conversion on the target aggregated data based on a target data compression structure to obtain user path data of the target object; wherein, aggregating the user behavior detail data based on the user identifiers in the user behavior detail data to obtain initial aggregated data, includes: determining initial event attribute information corresponding to an event type in the user behavior detail data; clipping the initial event attribute information to obtain target user behavior detail data; aggregating the target user behavior detail data based on the user identifiers in the user behavior detail data to obtain initial aggregated data.

2. The method according to claim 1, characterized in that, clipping the initial event attribute information to obtain target user behavior detail data, includes: determining event attribute information and attribute information to be clipped in the initial event attribute information; retaining the event attribute information in the user behavior detail data, and clipping the attribute information to be clipped in the user behavior detail data to obtain target user behavior detail data.

3. The method according to claim 2, characterized in that, determining event attribute information and attribute information to be clipped in the initial event attribute information, includes: determining event identifier information and event time information in the initial event attribute information as event attribute information; determining other event attribute information in the initial event attribute information except the event identifier information and the event time information as attribute information to be clipped.

4. The method according to claim 2, characterized in that, aggregating the target user behavior detail data based on the user identifiers in the user behavior detail data to obtain initial aggregated data, includes: determining a target user identifier in the user behavior detail data; determining target user behavior sub-data according to the target user identifier in the target user behavior detail data; performing deduplication processing on the user attribute information in the target user behavior sub-data to obtain user attribute information corresponding to the target user identifier; performing aggregation processing on the event attribute information according to the user attribute information to obtain event attribute information corresponding to the user attribute information; concatenating the user attribute information and the event attribute information to obtain initial aggregated data corresponding to the target user identifier.

5. The method according to claim 4, characterized in that, performing aggregation processing on the event attribute information according to the user attribute information to obtain event attribute information corresponding to the user attribute information, includes: Splice the event identification information based on the event time information to obtain the event attribute information corresponding to the user attribute information.

6. The method according to claim 4, wherein, Splicing the user attribute information and the event attribute information to obtain the initial aggregated data corresponding to the target user identifier includes: Determine that the event identifier that satisfies the preset interference condition in the event attribute information is an interference event identifier, where the preset interference condition is a behavioral event condition that the user repeatedly occurs within a preset time interval; Delete the interference event identifier in the event attribute information and the event time corresponding to the interference event identifier; Splice the event attribute information after the deletion operation with the user attribute information to obtain the initial aggregated data corresponding to the target user identifier.

7. The method according to claim 1, wherein, Dividing the event attribute information in the initial aggregated data according to a preset divided time interval to obtain target aggregated data includes: Dividing the event time information in the initial aggregated data according to a preset divided time interval to obtain at least one event time interval; Determine the event identifier corresponding to each event time interval as an event identifier set, and generate target aggregated data based on the event identifier set and the user attribute information.

8. The method according to claim 1, wherein, Performing format conversion on the target aggregated data based on the target data compression structure to obtain the user path data of the target object, including: Performing format conversion on the user identifier in the user attribute information of the target aggregated data based on the target data compression structure to obtain the user path data of the target object.

9. The method according to claim 1, wherein, After performing format conversion on the target aggregated data based on the target data compression structure to obtain the user path data of the target object, it further includes: Obtain the user identifier of the target object and the attribute label corresponding to the user identifier; Perform format conversion on the user identifier based on the target data compression structure, and determine the user label data of the target object based on the converted user identifier and the attribute label.

10. The method according to claim 9, wherein, After performing format conversion on the target aggregated data based on the target data compression structure to obtain the user path data of the target object, it further includes: Store the user path data of the target object in the database; Correspondingly, after determining the user label data of the target object based on the converted user identifier and the attribute label, it further includes: Store the user label data of the target object in the database.

11. The method according to claim 10, wherein, After storing the user path data of the target object in the database, it further includes: Receive a user path data query request for the target object, where the user path data query request carries a basic configuration query condition; Query corresponding user path data from the user path data in the database based on the basic configuration query conditions, where the basic configuration query conditions include at least one of an event time condition, a central event condition, and a user device condition; Generate a user path map based on the user path data and send the user path map to the user path map display interface of the target object.

12. The method according to claim 10, wherein, after storing the user label data of the target object in the database, it further includes: Receiving a query request for user path data of a target object, where the user path data query request carries a basic configuration query condition and a label data query condition; Determine the to-be-processed user path data from the user path data in the database based on the basic configuration query condition, and determine the to-be-processed user label data from the user label data in the database based on the label data query condition; Process the to-be-processed user path data and the to-be-processed user label data according to a preset data processing method to obtain labeled user path data; Generate a labeled user path map based on the labeled user path data and send the labeled user path map to the user path map display interface of the target object.

13. The method according to claim 12, wherein, Processing the to-be-processed user path data and the to-be-processed user label data according to a preset data processing method to obtain labeled user path data, including: Determine the association relationship between the basic configuration query condition and the label data query condition, and determine the preset data processing method based on the association relationship; Process the to-be-processed user path data and the to-be-processed user label data based on the preset data processing method to obtain labeled user path data.

14. The method as claimed in claim 1, wherein, The user attribute information includes at least one of user identification information, user device information, device system information, and device model information.

15. The method as claimed in claim 1, wherein, The target aggregation data is the aggregation data obtained by reprocessing the initial aggregation data.

16. A data processing device, wherein, comprises: A data acquisition module configured to acquire user behavior detail data of a target object within a preset historical time period; An initial aggregation module configured to aggregate the user behavior detail data based on the user identification in the user behavior detail data to obtain initial aggregation data, where the initial aggregation data includes user attribute information and event attribute information; A target aggregation module configured to divide the event attribute information in the initial aggregation data according to a preset divided time period to obtain target aggregation data, where the target aggregation data includes user attribute information and an event identifier set; A data conversion module configured to perform format conversion on the target aggregation data based on a target data compression structure to obtain the user path data of the target object; Among them, based on the user identifier in the user behavior detail data, the user behavior detail data is aggregated to obtain initial aggregated data, including: Determine the initial event attribute information corresponding to the event type in the user behavior detail data; Crop the initial event attribute information to obtain target user behavior detail data; Based on the user identifier in the user behavior detail data, aggregate the target user behavior detail data to obtain initial aggregated data.

17. A computing device, comprising a memory, a processor, and computer instructions stored on the memory and executable on the processor, wherein, when the processor executes the computer instructions, the steps of the method according to any one of claims 1-15 are implemented.

18. A computer-readable storage medium storing computer instructions, wherein, when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1-15 are implemented.

19. A computer program product comprising computer instructions, wherein, when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1-15 are implemented.

Citation Information

Patent Citations

  • User behavior analysis system and method, storage medium and computing equipment

    CN111488261A

  • Method and apparatus of user clustering, computer device and medium

    US20210191958A1