Data analysis method and device
By generating user model data and hierarchical data, the problem of high time and labor costs in user operation analysis is solved, and personalized marketing strategies and system performance improvements are achieved.
Patent Information
- Application Number
- CN202110524867.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-13
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-05-13
AI Technical Summary
Existing technologies consume a lot of time and manpower in user operation analysis, making it difficult to implement personalized marketing strategies.
By obtaining user original data and model configuration information, using the model configuration information to generate user model data, and combining the stratification rule information to determine the user stratification data, stratification data suitable for different users is automatically generated.
It saves time and labor costs, can set personalized marketing strategies for different users, adapt to the different needs of multiple users, reduce duplicate processing, and improve system performance.
Smart Images

Figure CN113239083B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a data analysis method and device. Background Art
[0002] To better serve customers, companies need to develop diverse marketing strategies. This typically involves manually filtering out a list of target users from a large user base, implementing marketing strategies for those users, and conducting relevant user operations analysis. This approach consumes significant time and labor costs, hindering the implementation of personalized marketing strategies for different users. Summary of the Invention
[0003] In view of this, an embodiment of the present invention provides a data analysis method and device, which can save time and manpower costs required for user operation analysis.
[0004] In a first aspect, an embodiment of the present invention provides a data analysis method, comprising:
[0005] Obtain user original data and model configuration information;
[0006] Using the model configuration information, processing the user original data to generate model data of at least one user model, the model data including: user identification and statistical indicator values;
[0007] Acquire stratification rule information, wherein the stratification rule information includes: stratification identifiers and filtering rules of statistical indicators;
[0008] Determine hierarchical data of at least one user layer based on the model data and the hierarchical rule information, and use the hierarchical data for user operation analysis.
[0009] Optionally, the model configuration information includes: model identification, target table, statistical fields and statistical methods;
[0010] The processing of the user original data using the model configuration information to generate model data of at least one user model includes:
[0011] Performing statistical processing on the statistical fields in the target table according to the statistical method to generate statistical indicator values corresponding to the model identifiers;
[0012] Model data corresponding to the model identifier is generated according to the statistical indicator value corresponding to the model identifier.
[0013] Optionally, the user model information includes: a model identifier, multiple target tables, an aggregation method, a statistical field, and a statistical method;
[0014] The processing of the user original data using the model configuration information to generate model data of at least one user model includes:
[0015] performing an aggregation operation on the multiple target tables according to the aggregation method;
[0016] Performing statistical processing on the statistical fields in the aggregated table according to the statistical method to generate statistical indicator values corresponding to the model identifier;
[0017] Model data corresponding to the model identifier is generated according to the statistical indicator value corresponding to the model identifier.
[0018] Optionally, before processing the user original data using the model configuration information to generate model data of at least one user model, the method further includes:
[0019] Receiving model input information input by a user, wherein the model input information includes: at least one target table, a statistical field, and a statistical method;
[0020] The model configuration information is generated according to the target table, the statistical fields and the statistical method.
[0021] Optionally, the receiving model input information input by the user includes:
[0022] Get user permissions;
[0023] Display at least one original data table corresponding to the user authority;
[0024] receiving a selection operation for the at least one original data table;
[0025] According to the selection operation, the target table is determined.
[0026] Optionally, the hierarchical rule information further includes: a model identifier;
[0027] The determining, based on the model data and the stratification rule information, stratification data of at least one user stratification includes:
[0028] Determine the target model according to the model identifier in the hierarchical rule information;
[0029] According to the filtering rule, the hierarchical data corresponding to the hierarchical identifier is filtered out from the model data corresponding to the target model.
[0030] Optionally, the hierarchical rule information further includes: filtering rules of multiple model identifiers, association methods, and multiple statistical indicators;
[0031] The determining, based on the model data and the stratification rule information, stratification data of at least one user stratification includes:
[0032] determining a plurality of target models according to the model identifier;
[0033] Using the filtering rules, respectively determining the data sets of the target models;
[0034] According to the association method, the plurality of data sets are processed to generate hierarchical data corresponding to the hierarchical identifier.
[0035] Optionally, before determining the hierarchical data of at least one user hierarchy according to the model data and the hierarchical rule information, the method further includes:
[0036] Receiving hierarchical input information input by a user, the hierarchical input information including: a model identifier and a filtering rule of a statistical indicator, the model identifier corresponding to the statistical indicator;
[0037] The hierarchical rule information is generated according to the model identifier and the filtering rule.
[0038] Optionally, the hierarchical data includes: time of occurrence;
[0039] After determining the hierarchical data of at least one user hierarchy according to the model data and the hierarchical rule information, the method further includes:
[0040] Filtering the current data set of the current period from the hierarchical data of the target layer;
[0041] Filtering out a previous data set of a previous period from the hierarchical data of the target hierarchical layer;
[0042] The target layered changed user data is generated according to the current data set and the previous data set.
[0043] Optionally, the hierarchical data includes: time of occurrence;
[0044] After determining the hierarchical data of at least one user hierarchy according to the model data and the hierarchical rule information, the method further includes:
[0045] Acquire the flow information of the target layer, wherein the flow information includes: an outgoing layer identifier;
[0046] Filtering a first data set of a first time period from the hierarchical data of the target hierarchical layer;
[0047] Determining a second time period corresponding to the outflow stratification according to the circulation cycle and the first time period;
[0048] Filtering a second data set of a second time period from the hierarchical data corresponding to the outflow hierarchical identifier;
[0049] The outflow user data of the target layer in the first time period is determined based on the first data set and the second data set.
[0050] Optionally, the hierarchical data includes: time of occurrence;
[0051] After determining the hierarchical data of at least one user hierarchy according to the model data and the hierarchical rule information, the method further includes:
[0052] Acquire flow information of a target layer, the flow information including: an inflow layer identifier;
[0053] Filtering a third data set of a third time period from the hierarchical data of the target hierarchical layer;
[0054] Determining a fourth time period corresponding to the inflow layer according to the circulation cycle and the third time period;
[0055] Filtering a fourth data set of a fourth time period from the hierarchical data corresponding to the inflow hierarchical identifier;
[0056] Inflow user data of the target layer in the third time period is determined based on the third data set and the fourth data set.
[0057] In a second aspect, an embodiment of the present invention provides a data analysis device, comprising:
[0058] The first acquisition module is used to obtain user original data and model configuration information;
[0059] A model generation module, configured to process the user original data using the model configuration information to generate model data of at least one user model, wherein the model data includes: a user identifier and a statistical indicator value;
[0060] The second acquisition module is used to obtain stratification rule information, wherein the stratification rule information includes: stratification identification and filtering rules of statistical indicators;
[0061] The layer determination module is used to determine layer data of at least one user layer based on the corresponding relationship and the model data, and the layer data is used for user operation analysis.
[0062] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0063] one or more processors;
[0064] a storage device for storing one or more programs,
[0065] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of the above embodiments.
[0066] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium having a computer program stored thereon, which implements the method described in any of the above embodiments when the program is executed by a processor.
[0067] One embodiment of the above invention has the following advantages or beneficial effects: Different users can configure different model configuration information and stratification rule information according to their needs. The system automatically generates stratified data for different user strata using the model configuration information and stratification rule information. Compared to manually screening target users, this saves time and labor costs. Furthermore, different marketing strategies can be set for different user strata, facilitating the implementation of personalized marketing strategies for different users.
[0068] Furthermore, when processing user data according to a specific model, since the model is fixed, it often cannot adapt to the different needs of multiple users. In the embodiments of the present application, users can refer to and utilize the existing model configuration information and layering rule information in the system, or configure different model configuration information and layering rule information according to their own needs. Therefore, it can meet the different needs of multiple users.
[0069] In addition, the model data in the same user model can be shared by multiple users in layers, reducing the repeated processing of the same data and improving system performance.
[0070] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings are provided for a better understanding of the present invention and are not intended to limit the present invention.
[0072] Figure 1 is a schematic diagram of an exemplary application scenario in which embodiments of the present invention can be applied;
[0073] Figure 2 This is a schematic diagram of a process of a data analysis method provided by an embodiment of the present invention;
[0074] Figure 3 is a schematic diagram of the process of another data analysis method provided by one embodiment of the present invention;
[0075] Figure 4 is a schematic diagram of a process of another data analysis method provided by an embodiment of the present invention;
[0076] Figure 5 This is a schematic diagram of a data flow sequence of a data analysis system provided by one embodiment of the present invention;
[0077] Figure 6 This is a flowchart of a task analysis provided by an embodiment of the present invention;
[0078] Figure 7 This is a schematic diagram of a process for processing user detailed data provided by an embodiment of the present invention;
[0079] Figure 8 This is a schematic diagram of a user data analysis process provided by an embodiment of the present invention;
[0080] Figure 9 This is a schematic structural diagram of a data analysis device provided by one embodiment of the present invention;
[0081] Figure 10 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION
[0082] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, in which various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0083] Figure 1 Schematic diagram of an exemplary application scenario in which the embodiment of the present invention can be applied. Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is used to provide a medium for communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0084] The client or browser of the data analysis system can be deployed in the terminal devices 101, 102, and 103. The terminal devices 101, 102, and 103 can use the client or browser to interact with the server 105. The terminal devices 101, 102, and 103 can be mobile phones, notebooks, tablet computers, laptop computers, etc.
[0085] The terminal devices 101 , 102 , and 103 interact with the server 105 via the network 104 to receive or send messages, etc. The terminal devices 101 , 102 , and 103 send stored videos to the server 105 via the network 104 .
[0086] The user sends the model configuration information and the layering rule information to the server 105 through the terminal devices 101, 102, and 103. The server 105 processes the user's original data according to the model configuration information and the layering rule information to generate different user layered data.
[0087] It should be noted that the data analysis method provided in the embodiment of the present invention is generally executed by the server 105 , and accordingly, the data analysis device is generally set in the server 105 .
[0088] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0089] Figure 2 FIG. 1 is a schematic diagram of a data analysis method according to an embodiment of the present invention. Figure 2 As shown, the method includes:
[0090] Step 201: Obtain user original data and model configuration information.
[0091] User raw data can come from underlying data tables in production systems or from data tables in data marts or data warehouses. User raw data can be uploaded to the system using Excel, CSV, or other file formats. User raw data can also be extracted into the system using ETL (Extract, Transform, and Load) tools.
[0092] Multiple user models can be set up in the system. The model data in the user model is between the user's original data and the user's hierarchical data. Multiple user models can be set up according to different statistical indicators, production systems, and user groups.
[0093] Step 202: Process the user original data using the model configuration information to generate model data of at least one user model.
[0094] The model data for a user model may include: model ID, user ID, statistical indicator value 1, statistical indicator value 2, etc., occurrence time, etc. The statistical indicator value provides the basis for stratification data used to subsequently determine user stratification. Statistical indicators can be set based on specific business needs, such as purchase amount, number of purchases, number of favorites, etc.
[0095] The model configuration information can be used to process the user's original data and generate the model data of the user model. The model configuration information may include: model identification, target table, statistical fields and statistical methods.
[0096] Step 203: Acquire layering rule information, which includes layering identifiers and filtering rules of statistical indicators.
[0097] User stratification is a classification of users into multiple tiers based on their status, behavior data, and attributes. This facilitates operational analysis of users. User stratification can be customized based on specific business needs. For example, users can be categorized into multiple tiers, such as regular users, active users, and contributing users.
[0098] By utilizing the filtering rules of the statistical indicators and processing the model data of the user model, the hierarchical data of the user stratification can be determined. For example, if the stratification identifier in stratification rule information 1 corresponds to an ordinary user, the statistical indicator corresponds to the purchase amount, and the filtering rule indicates that the purchase amount is less than 1,000, then the stratification rule information 1 indicates that users with a purchase amount less than 1,000 are classified into the user stratum of ordinary users. If the stratification identifier in stratification rule information 2 corresponds to an active user, the statistical indicator corresponds to the purchase amount, and the filtering rule indicates that the purchase amount is between 1,000 and 3,000, then the stratification rule information 2 indicates that users with a purchase amount between 1,000 and 3,000 are classified into the user stratum of active users.
[0099] Step 204: Determine hierarchical data of at least one user layer based on the model data and the hierarchical rule information. The hierarchical data is used for user operation analysis.
[0100] The model data may be screened using the filtering rules of the statistical indicators in the stratification rule information and the statistical indicator values in the model data to determine the stratified data of the user stratification.
[0101] It should be noted that different users can configure different user tiers as needed. For example, users can be categorized into free users, active users, paying users, and high-paying users based on purchase amount. Users can also be categorized into new users, registered users, active users, and paying users based on purchase amount and number of purchases.
[0102] In this embodiment of the present invention, different users can configure different model configuration information and stratification rule information based on their needs. The system automatically generates stratified data for different user strata using this information. Compared to manually screening target users, this method saves time and labor costs. Furthermore, different marketing strategies can be set for different user strata, facilitating the implementation of personalized marketing strategies for different users.
[0103] Furthermore, different users can configure different model configuration information and stratification rule information based on their own needs. This allows for flexible and convenient access to different user-stratified data. This solves the problem of using fixed models for data processing, which often fails to adapt to the diverse needs of multiple users.
[0104] In addition, the model data in the same user model can be shared by multiple users in layers, reducing the repeated processing of the same data and improving system performance.
[0105] Figure 3 FIG. 1 is a schematic diagram of a data analysis method according to an embodiment of the present invention. Figure 3 As shown, the method includes:
[0106] Step 301: Obtain user original data and model configuration information, where the model configuration information includes: model identifier, target table, statistical fields, and statistical methods.
[0107] Statistical fields are fields in the target table that correspond to statistical indicators. Statistical indicators correspond to filtering rules in the stratification rule information.
[0108] The statistical method is used to calculate the statistical indicator value based on the statistical field. Statistical methods can include: counting the number of times, summing, etc.
[0109] Step 302: Perform statistical processing on the statistical fields in the target table according to a statistical method to generate model data corresponding to the model identifier.
[0110] Step 303: The statistical fields in the target table are grouped by user ID and statistically processed in a statistical manner to generate statistical index values in the model data.
[0111] Step 304: Generate model data corresponding to the model identifier according to the statistical indicator values in the model data.
[0112] Step 305: Determine hierarchical data of at least one user hierarchy based on the model data and hierarchical rule information.
[0113] Based on the model configuration information, operations can be performed on a single target table to generate corresponding model data. For example, the target table includes the following fields: user ID, purchase amount, and occurrence time. The statistical field in the model configuration information corresponds to the purchase amount, and the statistical method in the model configuration information is summation. Based on this model configuration information, the purchase amounts in the target table are grouped by user ID and summed to obtain the model data corresponding to this model configuration information.
[0114] As one possible implementation, the model configuration information can also include a time condition. Using the above example, the model configuration information also includes a time condition representing the statistical indicator value within seven days. Based on this model configuration information, the user ID is grouped, and records in the target table with a time of occurrence within seven days are filtered. The purchase amounts are summed to obtain the model data corresponding to this model configuration information.
[0115] In the stratification rule information, the stratification identifier corresponds to the contributing user, the statistical indicator corresponds to the purchase amount, and the filtering rule represents that the purchase amount is greater than 3,000. Then, according to the stratification rule information, user records whose total purchase amount is greater than 3,000 are filtered out from the above model data, and these user records are divided into the user stratum of contributing users.
[0116] In one embodiment of the present invention, user model information includes: a model identifier, multiple target tables, an aggregation method, statistical fields, and a statistical method. Model data for at least one user model can also be generated by: performing an aggregation operation on the multiple target tables according to the aggregation method; grouping the statistical fields in the aggregated tables by user identifier and performing statistical processing according to the statistical method to generate a statistical indicator value corresponding to the model identifier; and generating model data corresponding to the model identifier based on the statistical indicator value corresponding to the model identifier.
[0117] Aggregation methods involve joining multiple target tables based on user IDs. Aggregation methods include inner join, left join, right join, and full join.
[0118] Statistical fields are the fields in the aggregated table that correspond to statistical indicators. Statistical indicators correspond to the filtering rules in the stratification rule information. Statistical methods are used to calculate statistical indicator values based on statistical fields. Statistical methods include counting the number of times and summing the values.
[0119] Based on the model configuration information, operations can be performed on multiple target tables to generate corresponding model data. Multiple target tables are aggregated according to the aggregation method. Statistical fields in the aggregated tables are selected and statistically processed according to the statistical method to determine the model data corresponding to the model configuration information.
[0120] In one embodiment of the present invention, before processing user raw data using model configuration information to generate model data for at least one user model, the process further includes: receiving model input information from the user, the model input information including at least one target table, statistical fields, and statistical methods; and generating model configuration information based on the target table, statistical fields, and statistical methods. Users can configure specific model configuration information based on their needs to generate model data that meets their business analysis requirements, flexibly adapting to the diverse needs of multiple users.
[0121] In one embodiment of the present invention, receiving model input information from a user includes: obtaining user permissions; displaying at least one raw data table corresponding to the user permissions; receiving a selection operation for the at least one raw data table; and determining a target table based on the selection operation. Users can only operate on tables corresponding to their permissions. By setting user permissions, the security of user data is ensured.
[0122] In one embodiment of the present invention, before determining stratified data for at least one user stratum based on the model data and stratification rule information, the method further includes: receiving stratification input information input by a user, the stratification input information including a model identifier and filtering rules for statistical indicators, wherein the model identifier corresponds to the statistical indicator; and generating stratification rule information based on the model identifier and the filtering rules. Users can configure specific stratification rule information based on their needs to generate user stratification data that meets their business analysis requirements, thereby flexibly adapting to the diverse needs of multiple users.
[0123] In one embodiment of the present invention, the hierarchical rule information also includes: a model identifier; determining the hierarchical data of at least one user layer based on the model data and the hierarchical rule information, including: determining the target model based on the model identifier in the hierarchical rule information; and filtering out the hierarchical data corresponding to the hierarchical identifier from the model data corresponding to the target model based on the filtering rules.
[0124] Based on the filtering rules in the stratification rule information, operations can be performed on individual model data to generate corresponding stratified data. The stratification rule information includes stratification rule information 3 and stratification rule information 4. The statistical indicators corresponding to stratification rule information 3 and stratification rule information 4 are both the number of purchases. Stratification rule information 3 corresponds to ordinary users, and the filtering rule indicates that the number of purchases is less than 10. Based on this filtering rule, users with less than 10 purchases are classified into the ordinary user stratum.
[0125] The stratification rule information 2 corresponds to active users, and the filtering rule indicates that the number of purchases is between 10 and 50. Based on this filtering rule, users with the number of purchases between 10 and 50 are classified into the user stratum of active users.
[0126] Figure 4 FIG. 1 is a schematic diagram of a process flow of another data analysis method provided by an embodiment of the present invention. Figure 4 As shown, the method includes:
[0127] Step 401: Obtain user original data and model configuration information.
[0128] Step 402: Process the user original data using the model configuration information to generate model data of at least one user model. The model data includes: a model identifier, a user identifier, and a statistical indicator value.
[0129] Step 403: Acquire layering rule information, which includes: a layering identifier, multiple model identifiers, an association method, and filtering rules for multiple statistical indicators.
[0130] The association method refers to the processing method between different model data, such as intersection operation, OR operation, merging operation, etc.
[0131] Step 404: Determine multiple target models according to the model identifiers in the layering rule information.
[0132] Step 405: Using the filtering rules, determine the data sets of each target model respectively.
[0133] Step 406: Process the multiple data sets according to the association method to generate hierarchical data corresponding to the hierarchical identifiers.
[0134] The filtering rules for multiple statistical indicators refer to filtering rules set for different model data. According to the filtering rules for multiple statistical indicators in the layering rule information, operations can be performed on multiple model data to generate corresponding layered data.
[0135] For example, consider hierarchical rule information for active users. For model data 1, filter rule 1 specifies that clicks are greater than 10. For model data 2, filter rule 1 specifies that purchases are greater than 1000. The association method is an intersection operation. Based on this hierarchical rule information, the user data set for model data 1 with clicks greater than 10 is determined, while the data set for model data 2 with purchases greater than 1000 is determined. The intersection of these two sets of user data is then used as the hierarchical data for active users.
[0136] In one embodiment of the present invention, the hierarchical data includes: occurrence time; after determining the hierarchical data of at least one user layer based on model data and hierarchical rule information, it also includes: filtering out the current data set of the current time period from the hierarchical data of the target layer; filtering out the previous data set of the previous time period from the hierarchical data of the target layer; generating the changed user data of the target layer based on the current data set and the previous data set.
[0137] The current time period can be set based on specific needs. The current time period can be the statistical day, statistical week, statistical month, etc. The previous time period can be the day, week, or month before the current time period, etc. The current layer data is compared with the previous layer data. Specifically, the current data set can be fully joined with the previous data set through operations such as full join, left join, or right join. Through this comparison, the changed user data for the target layer is determined. This changed user data includes both newly added user data and outgoing user data.
[0138] In one embodiment of the present invention, the layered data includes: occurrence time; after determining the layered data of at least one user layer based on model data and layered rule information, it also includes: obtaining flow information of the target layer, the flow information includes: an outflow layer identifier; filtering out a first data set of a first time period from the layered data of the target layer; determining a second time period corresponding to the outflow layer based on the flow cycle and the first time period; filtering out a second data set of the second time period from the layered data corresponding to the outflow layer identifier; and determining the outflow user data of the target layer in the first time period based on the first data set and the second data set.
[0139] For example, if the flow order of users in the user layer is: ordinary users -> active users -> contributing users, then for the active user layer, ordinary users are its inflow layer, and contributing users are its outflow layer.
[0140] The first time period can be set according to specific needs. The first time period can be a statistical day, statistical week, statistical month, etc. The turnover cycle is used to characterize the time required for a user to transition from the current user tier to the outflow user tier. The turnover cycle can be determined based on user behavior analysis or empirically. The second time period can be determined based on the turnover cycle and the first time period. For example, if the first time period is from time point 1 to time point 2, the second time period can be set from time point 1 to time point 3. Time point 3 can be obtained by adding the turnover cycle to time point 2.
[0141] The first data set and the second data set are compared. Specifically, the first data set and the second data set can be subjected to full association, left association, or right association operations. Through this comparison, the outflow user data of the target layer in the first time period can be determined, and ultimately, the number of users who are converting to the direction of marketing benefits can be analyzed.
[0142] In one embodiment of the present invention, the layered data includes: occurrence time; after determining the layered data of at least one user layer based on model data and layered rule information, it also includes: obtaining flow information of the target layer, the flow information includes: an inflow layer identifier; filtering out a third data set of a third time period from the layered data of the target layer; determining a fourth time period corresponding to the inflow layer based on the flow cycle and the third time period; filtering out a fourth data set of the fourth time period from the layered data corresponding to the inflow layer identifier; and determining the inflow user data of the target layer in the third time period based on the third data set and the fourth data set.
[0143] The third time period can be set according to specific needs. The third time period can be a statistical day, statistical week, statistical month, etc. The turnover cycle is used to characterize the time required for a user to transition from the incoming user tier to the current user tier. The turnover cycle can be determined based on user behavior analysis or through experience. The fourth time period can be determined based on the turnover cycle and the third time period. For example, if the third time period is from time point 4 to time point 5, the fourth time period can be set from time point 6 to time point 5, where time point 6 is obtained by subtracting the turnover cycle from time point 4.
[0144] The third data set is compared with the fourth data set. Specifically, a full join operation, a left join operation, or a right join operation can be performed on the third data set and the fourth data set. Through this comparison, the inflow user data of the target layer in the third time period can be determined, and ultimately, the number of users who are converting to the direction of favorable marketing can be analyzed.
[0145] To make the solution of the embodiment of the present invention easier to understand, a data analysis system is used as a specific embodiment for explanation below. Figure 5 FIG. 1 is a data flow sequence diagram of a data analysis system provided by an embodiment of the present invention. Figure 5 As shown, the system includes: a user layer, a user operation analysis system, a big data computing platform, and a device storage layer. The user layer uploads user element data to the big data computing platform. The user layer inputs model configuration information, statistical fields of each model, user stratification information, stratification rule information, etc. into the user operation analysis system and stores it in the storage device. The big data computing platform calculates the user analysis results based on the model configuration information and stratification rule information, and stores the user analysis results in the storage device. Users can issue a viewing instruction to the user operation analysis system to view the user analysis results. The data processing flow of the system includes the following steps:
[0146] S01: Upload of original data. First, to analyze user data, it is necessary to obtain the original data of user behavior, such as order details table, user click exposure table and other data upload or connection. Upload is a file upload method provided by the system, which can support data upload in formats such as excel and csv. The connection method can automatically synchronize user behavior data to the HIVE table of the big data platform that the analysis system relies on by configuring the business system database (such as mysql, elasticSearch, etc.) in the system, so as to perform big data analysis. For table fields, there are no strict requirements for the format, but at least it needs to be able to identify the user identification field, and can support multiple user identifications, including but not limited to IMEI (International Mobile Equipment Identity), device number, MAC (Media Access Control Address), mobile phone number, email address, etc.
[0147] S02: Configuration of model configuration information. First, through the big data platform service interface, according to the permission control, obtain the tables that users can operate and analyze for users to operate. It can support multiple tables to be associated and aggregated according to the user ID (identifier). The association and aggregation method is also determined by configuration, selecting the user ID field of each table.
[0148] After association and aggregation, configure the model configuration information. You can select the fields of the table after association and aggregation to determine. For example, specify the click field of Table 1, which corresponds to the statistical indicator. Set the statistical method, which mainly includes two types of statistical methods: sum (sum) and count (count), and finally determine the model. For example, you can use information to set the model configuration information of the click model: count the number of click fields in Table 1 according to the last 14 days.
[0149] S03: Configure the hierarchical information of user behavior, and the hierarchical rule information corresponding to each layer. The hierarchical information includes the attribute information of different layers. For example, users are stratified into a1, a2, a3, and a4, which represent users at different levels. The configuration of the hierarchical rule information is based on the click model defined in step S02, by setting the statistical indicator values in the model data, and performing "intersection (and)" or "or (or)" operations between the models. For example, the filtering rules in the hierarchical rule information corresponding to the a2 layer are set as follows: the number of clicks is greater than 10 and the purchase amount is greater than 1000, and the corresponding SQL expression can be generated to implement the hierarchical rule information.
[0150] S04: After the above configuration, the operator successfully created an analysis task. Figure 6 FIG. 1 is a flowchart of a task analysis process provided by an embodiment of the present invention. Figure 6 As shown, the system will automatically establish big data tasks on the big data platform through pipelines and task executors to process data and obtain user stratification and flow information.
[0151] S05: The first thing that is obtained by executing the big data computing task is the user's detailed data. Figure 7 This is a flowchart of a user detailed data processing method provided by an embodiment of the present invention. Figure 7 As shown, feature summarization is used to summarize user detailed data. Statistical fields in user detailed data can be used as data features. Feature collection utilizes user model information to process user raw data and generate user model data. Since a single user model can generate statistical values for a single metric, result set union (merging) and aggregation are used to merge multiple user model data sets using user identifiers, generating data in the following format: user ID, time, and values for multiple statistical metrics. Rule organization is used to store stratification rule information, including stratification identifiers and their corresponding filtering rules. Rule result evaluation is used to obtain statistical metric values from model data and determine whether these statistical metric values meet the filtering rules for that stratification according to the filtering rules. The same user may hit multiple strata simultaneously, so the stratification results corresponding to the same model data may be a set. If multiple strata are hit, stratification result organization is used to select the stratum with the highest level and use it as the stratum for the model data. For example, if the same user hits both a2 and a4, and a4's level is higher than a2's, the user is placed in a4's stratum. Writing to a Hive table persists the stratification results.
[0152] exist Figure 7 In the example, each user's data is analyzed according to the model and layer configuration information to obtain the (user ID, featureMap (a1, a2, ...), date) field information. Because users may match the filtering rules of multiple layers, the featureMap represents a layer set containing at least one layer. Date represents the statistical time, which can be expressed in days.
[0153] S06: Finally, the detailed data is processed again to obtain hierarchical information and flow information. Figure 8 This is a schematic diagram of a user data analysis process provided by an embodiment of the present invention. Figure 8 As shown, the stratified data is first fully joined with the stratified data from the previous day by user ID, and grouped. The purpose of joining with the stratified data from the previous day is to count any new users entering or leaving strata a1, a2, a3, and a4. Finally, the user data for each stratum, including new users entering or leaving, is counted.
[0154] User turnover analysis can also be performed by comparing previously configured data with data from a certain number of days prior (the turnover period). For example, through a full connection operation, the number of users moving from a1 to a2 is 1298, and the number of users moving from a4 to a2 is 2987. Ultimately, this allows analysis of how many users are actively converting to a positive marketing strategy, thereby achieving user operations goals.
[0155] The solution implemented in this embodiment of the present invention can be applied to internal company systems and to connecting to external enterprise data. The underlying big data tables are configurable, allowing the system to fully understand the meaning of the tables and fields. For example, user IDs, click exposure tables, and order tables can be specified. Specific models can be configured based on user needs. These include, but are not limited to, fields for user access frequency, purchase amount, and purchase categories. For user stratification analysis, users can customize the tier distribution based on their business scenarios and user behavior (for example, stratifying users into A, B, C, and D). User flow charts between tiers can also be defined (e.g., None -> A, C -> D). Based on the above configuration, user stratification rules can be configured. For example, to categorize enterprise users into the "awareness" tier, the filtering rule is: at least 10 views in the last 30 days, but no orders. Similarly, users in different tiers can be defined by the time field in the access table being greater than or equal to the current time minus 30 days, the number of views in the access table being greater than or equal to 10, the time field in the order table being greater than or equal to the current time minus 30 days, and the number of orders in the order table being zero. Finally, based on the above configuration information, the system performs big data processing on the user's original table data to obtain the number of users in each tier. At the same time, based on the changes in user IDs within each tier, user flow analysis is generated over time. This allows companies to understand the current status and flow of users at each tier, allowing for subsequent precision marketing.
[0156] The solution of the embodiment of the present invention can be configured to adapt to the user's original data without the need for pre-cleaning and sorting of data. The model and layered information are dynamically configured to apply to user operation analysis data in different business scenarios of different enterprises.
[0157] Figure 9 FIG. 1 is a schematic diagram of a data analysis device provided by an embodiment of the present invention. Figure 9 As shown, the device includes:
[0158] The first acquisition module 901 is used to obtain user original data and model configuration information;
[0159] The model generation module 902 is used to process the user original data using the model configuration information to generate model data of at least one user model, where the model data includes: a model identifier, a user identifier, and a statistical indicator value;
[0160] The second acquisition module 903 is used to obtain the layering rule information, which includes: layer identification, statistical index filtering rules;
[0161] The layer determination module 904 is configured to determine layer data of at least one user layer according to the corresponding relationship and the model data.
[0162] Optionally, the model configuration information includes: model identification, target table, statistical fields and statistical methods; the model generation module 902 is specifically used to:
[0163] Perform statistical processing on the statistical fields in the target table according to the statistical method to generate statistical indicator values corresponding to the model identifier;
[0164] Generate model data corresponding to the model identifier based on the statistical indicator value corresponding to the model identifier.
[0165] Optionally, the user model information includes: a model identifier, multiple target tables, an aggregation method, statistical fields, and a statistical method; the model generation module 902 is specifically used to:
[0166] For multiple target tables, perform aggregation operations according to the aggregation method;
[0167] Perform statistical processing on the statistical fields in the aggregated table according to the statistical method to generate statistical indicator values corresponding to the model identifier;
[0168] Generate model data corresponding to the model identifier based on the statistical indicator value corresponding to the model identifier.
[0169] Optionally, the device further comprises:
[0170] The information receiving module 905 is used to receive model input information input by the user, where the model input information includes: at least one target table, statistical fields, and statistical methods;
[0171] Generate model configuration information based on the target table, statistical fields and statistical methods.
[0172] Optionally, the information receiving module 905 is specifically configured to:
[0173] Get user permissions;
[0174] Display at least one original data table corresponding to user permissions;
[0175] receiving a selection operation for at least one original data table;
[0176] Determine the target table based on the selection operation.
[0177] Optionally, the layering rule information further includes: a model identifier; and the layering determination module 904 is specifically configured to:
[0178] Determine the target model according to the model identifier in the hierarchical rule information;
[0179] According to the filtering rules, the hierarchical data corresponding to the hierarchical identifier is filtered out from the model data corresponding to the target model.
[0180] Optionally, the stratification rule information further includes: multiple model identifiers, association methods, and filtering rules for multiple statistical indicators; the stratification determination module 904 is specifically configured to:
[0181] determining multiple target models according to the model identifier;
[0182] Using filtering rules, determine the data sets of each target model respectively;
[0183] According to the association method, multiple data sets are processed to generate hierarchical data corresponding to the hierarchical identifiers.
[0184] Optionally, the information receiving module 905 is further configured to:
[0185] Receiving hierarchical input information input by a user, the hierarchical input information including: a model identifier and a filtering rule for a statistical indicator, wherein the model identifier corresponds to the statistical indicator;
[0186] Generate hierarchical rule information based on model identification and filtering rules.
[0187] Optionally, the hierarchical data includes: occurrence time; the device further includes:
[0188] The data analysis module 906 is used to filter out the current data set of the current period from the hierarchical data of the target layer;
[0189] Filtering out a previous data set of a previous period from the hierarchical data of the target hierarchical layer;
[0190] Generate target layered change user data based on the current data set and the previous data set.
[0191] Optionally, the hierarchical data includes: occurrence time; the data analysis module 906 is further configured to:
[0192] Obtain the flow information of the target layer, the flow information includes: outgoing layer identifier;
[0193] Filtering a first data set of a first time period from the hierarchical data of the target hierarchical layer;
[0194] Determine a second time period corresponding to the outflow stratification according to the circulation cycle and the first time period;
[0195] Filtering the second data set of the second time period from the hierarchical data corresponding to the outflow hierarchical identifier;
[0196] Outgoing user data of a target layer in a first time period is determined based on the first data set and the second data set.
[0197] Optionally, the hierarchical data includes: occurrence time; the data analysis module 906 is further configured to:
[0198] Obtain the flow information of the target layer, the flow information includes: the flow layer identifier;
[0199] Filtering a third data set of a third period from the hierarchical data of the target hierarchical layer;
[0200] According to the circulation cycle and the third period, the fourth period corresponding to the inflow stratification is determined;
[0201] Filtering a fourth data set of a fourth time period from the hierarchical data corresponding to the inflow hierarchical identifier;
[0202] Inflow user data of the target layer in a third time period is determined based on the third data set and the fourth data set.
[0203] An embodiment of the present invention provides an electronic device, including:
[0204] one or more processors;
[0205] a storage device for storing one or more programs,
[0206] When one or more programs are executed by one or more processors, the one or more processors implement the method of any of the above embodiments.
[0207] Reference below Figure 10 , which shows a schematic structural diagram of a computer system 1000 of a terminal device suitable for implementing an embodiment of the present invention. Figure 10 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0208] like Figure 10 As shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the system 1000 are also stored in the RAM 1003. The CPU 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0209] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.
[0210] In particular, according to the embodiments disclosed herein, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed herein include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the above-mentioned functions defined in the system of the present invention are performed.
[0211] It should be noted that the computer-readable medium described in the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.
[0212] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0213] The modules described in the embodiments of the present invention may be implemented in software or hardware. The modules described may also be provided in a processor. For example, they may be described as: a first acquisition module, a model generation module, a second acquisition module, and a hierarchical determination module. The names of these modules do not, in some cases, limit the modules themselves. For example, the first acquisition module may also be described as a "module for acquiring user raw data and model configuration information."
[0214] As another aspect, the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiments, or may exist independently without being incorporated into the device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the device, the device includes:
[0215] Obtain user original data and model configuration information;
[0216] Using the model configuration information, processing the user original data to generate model data of at least one user model, the model data including: user identification and statistical indicator values;
[0217] Acquire stratification rule information, wherein the stratification rule information includes: stratification identifiers and filtering rules of statistical indicators;
[0218] Determine hierarchical data of at least one user hierarchy based on the model data and the hierarchical rule information.
[0219] According to the technical solutions of the embodiments of the present invention, different users can configure different model configuration information and stratification rule information based on their needs. The system automatically generates stratified data for different user strata using this information. Compared to manually screening target users, this method saves time and labor costs. Furthermore, different marketing strategies can be set for different user strata, facilitating the implementation of personalized marketing strategies for different users.
[0220] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A data analysis method, characterized in that: include: Obtain user original data and model configuration information; Processing the user original data using the model configuration information to generate model data of at least one user model; The model configuration information includes: model identification, multiple target tables, aggregation methods, statistical fields and statistical methods; the model data includes: user identification and statistical indicator values; multiple user models are set up in the system, and user models are set based on statistical indicators, production systems, and user groups; the model data in the user model is between the user's original data and the user's hierarchical data; the model data in the same user model is shared by multiple user hierarchies; The processing of the user original data using the model configuration information to generate model data of at least one user model includes: performing an aggregation operation on the multiple target tables according to the aggregation method; grouping the statistical fields in the aggregated tables by user identifiers and performing statistical processing according to the statistical method to generate statistical indicator values in the model data; and generating model data corresponding to the model identifier based on the statistical indicator values in the model data; Acquire stratification rule information, including stratification identifiers and filtering rules for statistical indicators; the model configuration information and stratification rule information are configured by the user according to his or her needs; Based on the model data and the stratification rule information, the stratification data of at least one user stratification is determined; if the model data hits multiple strata, the stratification results are sorted to take the stratum with the highest level, and the stratum with the highest level is used as the stratum of the model data; the stratification data is used for user operation analysis.
2. The method according to claim 1, characterized in that The model configuration information includes: model identification, target table, statistical fields and statistical methods; The processing of the user original data using the model configuration information to generate model data of at least one user model includes: The statistical fields in the target table are grouped by user identifiers and statistically processed according to the statistical method to generate statistical index values in the model data; Model data corresponding to the model identifier is generated according to the statistical indicator values in the model data.
3. The method according to claim 1, characterized in that Before the user original data is processed by using the model configuration information to generate model data of at least one user model, the method further includes: Receiving model input information input by a user, wherein the model input information includes: at least one target table, a statistical field, and a statistical method; The model configuration information is generated according to the target table, the statistical fields and the statistical method.
4. The method according to claim 3, characterized in that The receiving of model input information input by the user includes: Get user permissions; Display at least one original data table corresponding to the user authority; receiving a selection operation for the at least one original data table; According to the selection operation, the target table is determined.
5. The method according to claim 1, wherein The layering rule information also includes: a model identifier; The determining, based on the model data and the stratification rule information, stratification data of at least one user stratification includes: Determine the target model according to the model identifier in the hierarchical rule information; According to the filtering rule, the hierarchical data corresponding to the hierarchical identifier is filtered out from the model data corresponding to the target model.
6. The method according to claim 1, wherein the hierarchical rule information further comprises: Multiple model identifiers, association methods, and filtering rules for multiple statistical indicators; The determining, based on the model data and the stratification rule information, stratification data of at least one user stratification includes: Determine multiple target models according to the model identifiers in the hierarchical rule information; Using the filtering rules, respectively determining the data sets of the target models; According to the association method, the plurality of data sets are processed to generate hierarchical data corresponding to the hierarchical identifier.
7. The method according to claim 1, characterized in that Before determining the hierarchical data of at least one user hierarchy according to the model data and the hierarchical rule information, the method further includes: Receiving hierarchical input information input by a user, the hierarchical input information including: a model identifier and a filtering rule of a statistical indicator, the model identifier corresponding to the statistical indicator; The hierarchical rule information is generated according to the model identifier and the filtering rule.
8. The method according to claim 1, characterized in that The hierarchical data includes: occurrence time; After determining the hierarchical data of at least one user hierarchy according to the model data and the hierarchical rule information, the method further includes: Filtering the current data set of the current period from the hierarchical data of the target layer; Filtering out a previous data set of a previous period from the hierarchical data of the target hierarchical layer; The target layered changed user data is generated according to the current data set and the previous data set.
9. The method according to claim 1, wherein the hierarchical data comprises: Time of occurrence; After determining the hierarchical data of at least one user hierarchy according to the model data and the hierarchical rule information, the method further includes: Acquire the flow information of the target layer, wherein the flow information includes: an outgoing layer identifier; Filtering a first data set of a first time period from the hierarchical data of the target hierarchical layer; Determining a second time period corresponding to the outflow stratification according to the circulation cycle and the first time period; Filtering the second data set of the second time period from the hierarchical data corresponding to the outflow hierarchical identifier; The outflow user data of the target layer in the first time period is determined based on the first data set and the second data set.
10. The method according to claim 1, wherein the hierarchical data comprises: Time of occurrence; After determining the hierarchical data of at least one user hierarchy according to the model data and the hierarchical rule information, the method further includes: Acquire flow information of a target layer, the flow information including: an inflow layer identifier; Filtering a third data set of a third time period from the hierarchical data of the target hierarchical layer; Determining a fourth time period corresponding to the inflow layer according to the circulation cycle and the third time period; Filtering the fourth data set of the fourth time period from the hierarchical data corresponding to the inflow hierarchical identifier; Inflow user data of the target layer in the third time period is determined based on the third data set and the fourth data set.
11. A data analysis device, characterized in that: include: The first acquisition module is used to obtain user original data and model configuration information; A model generation module, configured to process the user original data using the model configuration information to generate model data of at least one user model; The model configuration information includes: model identification, multiple target tables, aggregation methods, statistical fields and statistical methods; the model data includes: user identification and statistical indicator values; multiple user models are set up in the system, and user models are set based on statistical indicators, production systems, and user groups; the model data in the user model is between the user's original data and the user's hierarchical data; the model data in the same user model is shared by multiple user hierarchies; The model generation module is specifically configured to: perform aggregation operations on the multiple target tables according to the aggregation method; group the statistical fields in the aggregated tables by user identifiers and perform statistical processing according to the statistical method to generate statistical indicator values in the model data; and generate model data corresponding to the model identifier based on the statistical indicator values in the model data; The second acquisition module is used to obtain stratification rule information, which includes stratification identifiers and filtering rules of statistical indicators. The model configuration information and stratification rule information are configured by the user according to his or her own needs. A layer determination module is used to determine the layer data of at least one user layer based on the model data and the corresponding relationship; if the model data hits multiple layers, the layer results are sorted to take the layer with the highest level, and the layer with the highest level is used as the layer of the model data; the layer data is used for user operation analysis.
12. An electronic device, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 10.
13. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Network flow multidimensional operation analysis method and device
CN112598442A