User data multi-dimensional attribution processing method and device, equipment and storage medium
By constructing an index sub-item tree structure, combining smooth proportion weighting and search strategies, the problem of low multi-dimensional attribution analysis of user business data in the existing technology is solved, and the rapid and accurate multi-dimensional attribution analysis results are achieved.
Patent Information
- Application Number
- CN202510531060.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
The root cause mining algorithm of user business data in the prior art is mainly applicable to single-dimensional and absolute value indicators, with low computational efficiency, and more focused on sorting but lacks direct quantitative representation of influence.
By obtaining the observation data set and comparison data set of user business data, combining the smooth proportion weighting strategy and search strategy, an indicator sub-item tree structure is constructed to determine the attribution analysis results corresponding to each indicator ID.
A multi-dimensional attribution analysis of user business data is realized, and the quantifiable attribution analysis results are quickly and accurately screened out, which improves the calculation efficiency and direct quantitative representation of impact.
Smart Images

Figure CN120448729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method, device, equipment, and storage medium for multi-dimensional attribution processing of user data. Background Art
[0002] After acquiring a large amount of user business data, companies often use dimensional attribution to identify key dimensions. Dimensional attribution of indicators refers to the process of breaking down and analyzing changes in an overall indicator (such as sales growth or conversion rate decline) according to different dimensions (such as region, product category, and customer base) during data analysis. Dimensional attribution provides a deeper understanding of the causes of indicator changes, thus supporting companies in building decision trees using user business data.
[0003] Currently, common root cause mining algorithms for user business data include Adtributor, iDice, and HotSpot. The Adtributor algorithm assumes that all root causes are one-dimensional. It ranks dimensions by calculating their s-value (the sum of the s-values of all elements within the dimension), identifying the most unexpected dimension (e.g., province). It then calculates the explanatory power (EP) of each element within the dimension. When the sum of the explanatory power (e.g., province 1 + province 2) exceeds a threshold, these elements are considered root causes. However, this algorithm is primarily suitable for single-dimensional, absolute value indicators. Other algorithms, such as iDice and HotSpot, also have shortcomings, such as being unsuitable for relative value indicators, low computational efficiency, and a focus on ranking rather than directly quantifying impact. Summary of the Invention
[0004] The embodiments of the present invention provide a multi-dimensional attribution processing method, device and equipment for user data, aiming to solve the problem that the common root cause mining algorithms for user business data in the existing technology are applicable to single-dimensional and absolute value indicators, and the root cause mining calculation efficiency is poor, focusing more on sorting but lacking direct quantitative representation of influence.
[0005] In a first aspect, an embodiment of the present invention provides a method for multi-dimensional attribution processing of user data, comprising:
[0006] In response to the data attribution instruction, obtaining a user service data acquisition statement corresponding to the data attribution instruction; wherein the corresponding data collection time in the user service data acquisition statement includes at least a preset observation time and a preset comparison time;
[0007] If it is detected that the current system time is the preset observation time, the user service data acquisition statement is executed to acquire an observation data set and a comparison data set from the data source, and the observation data set and the comparison data set are merged to obtain the current data set to be processed; wherein the data collection time corresponding to the observation data set is the preset observation time, and the data collection time corresponding to the comparison data set is the preset comparison time;
[0008] Obtaining, based on a preset smoothing proportion weighting strategy, an influence index corresponding to each set of comparative user service data in the current data set to be processed; wherein each set of comparative user service data in the current data set to be processed includes observation data and comparative data having the same indicator ID and parameter cross-combination sub-items;
[0009] Based on a preset search strategy and the influence index corresponding to each group of comparison user service data in the current data set to be processed, determine the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed;
[0010] Based on the indicator item tree structure corresponding to each indicator ID in the current data set to be processed and a preset node screening strategy, the attribution analysis result corresponding to each indicator ID is determined.
[0011] In a second aspect, an embodiment of the present invention further provides a user data multi-dimensional attribution processing device, comprising:
[0012] A data acquisition statement acquisition unit is configured to, in response to a data attribution instruction, acquire a user service data acquisition statement corresponding to the data attribution instruction; wherein the corresponding data collection time in the user service data acquisition statement includes at least a preset observation time and a preset comparison time;
[0013] a current data set to be processed acquiring unit, configured to, if it is detected that the current system time is the preset observation time, execute the user service data acquiring statement to acquire an observation data set and a comparison data set from the data source, and merge the observation data set and the comparison data set to obtain the current data set to be processed; wherein the data collection time corresponding to the observation data set is the preset observation time, and the data collection time corresponding to the comparison data set is the preset comparison time;
[0014] an influence index calculation unit, configured to obtain, based on a preset smoothed proportion weighting strategy, an influence index corresponding to each set of comparative user service data in the current data set to be processed; wherein each set of comparative user service data in the current data set to be processed includes observation data and comparative data having the same indicator ID and parameter cross-combination sub-items;
[0015] A tree structure acquisition unit, configured to determine an indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed based on a preset search strategy and an influence index corresponding to each group of compared user service data in the current data set to be processed;
[0016] The attribution analysis result acquisition unit is used to determine the attribution analysis result corresponding to each indicator ID based on the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed and a preset node screening strategy.
[0017] In a third aspect, an embodiment of the present invention further provides a computer device comprising a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in the first aspect is implemented.
[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method described in the first aspect can be implemented.
[0019] The embodiment of the present invention provides a method, apparatus, device and storage medium for multi-dimensional attribution processing of user data, the method comprising: in response to a data attribution instruction, obtaining a user business data acquisition statement corresponding to the data attribution instruction; wherein the corresponding data collection time in the user business data acquisition statement includes at least a preset observation time and a preset comparison time; if it is detected that the current system time is the preset observation time, executing the user business data acquisition statement to obtain an observation data set and a comparison data set from a data source, and merging the observation data set and the comparison data set to obtain a current data set to be processed; wherein the data collection time corresponding to the observation data set is the preset observation time, and the data collection time corresponding to the comparison data set is the preset observation time. The collection time is set as the preset comparison time; based on the preset smoothing proportion weighting strategy, the influence index corresponding to each group of comparison user business data in the current data set to be processed is obtained; wherein, the influence index corresponding to each group of comparison user business data is determined by the conversion influence index and the structural influence index; based on the preset search strategy and the influence index corresponding to each group of comparison user business data in the current data set to be processed, the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is determined; based on the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed and the preset node screening strategy, the attribution analysis result corresponding to each indicator ID is determined. The embodiment of the present invention can quickly determine the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed composed of user business data in combination with the smoothing proportion weighting strategy and the search strategy, thereby combining the indicator sub-item tree structure to more accurately screen and obtain quantifiable display attribution analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of an application scenario of the multi-dimensional attribution processing method for user data provided by an embodiment of the present invention;
[0022] Figure 2 A flowchart of a multi-dimensional attribution processing method for user data provided by an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of a sub-process of a multi-dimensional attribution processing method for user data provided by an embodiment of the present invention;
[0024] Figure 4 A schematic diagram of a sub-process of a multi-dimensional attribution processing method for user data provided by an embodiment of the present invention;
[0025] Figure 5 A schematic diagram of a sub-process of a multi-dimensional attribution processing method for user data provided by an embodiment of the present invention;
[0026] Figure 6 A schematic block diagram of a device for multi-dimensional attribution processing of user data provided by an embodiment of the present invention;
[0027] Figure 7 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0029] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0030] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0031] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0032] Please also refer to Figure 1 and Figure 2 ,in Figure 1 Schematic diagram of a scenario of a multi-dimensional attribution processing method for user data according to an embodiment of the present invention. Figure 2 FIG is a flow chart of a multi-dimensional attribution processing method for user data provided by an embodiment of the present invention. Figure 1 As shown, the multi-dimensional attribution processing method for user data provided by an embodiment of the present invention is applied to the server 10. The server 10 can also be connected to other servers 20 for communication, and the other servers 20 can be regarded as servers corresponding to the storage area of the data source.
[0033] like Figure 2 As shown, the method includes the following steps S110-S150.
[0034] S110. In response to a data attribution instruction, obtain a user business data acquisition statement corresponding to the data attribution instruction.
[0035] The data collection time corresponding to the user service data acquisition statement includes at least a preset observation time and a preset comparison time.
[0036] In this embodiment, the technical solution is described with the server as the execution subject. In the server, multi-dimensional attribution processing can be performed on the data collected from the data source, thereby constructing an attribution tree. Among them, each user data stored at the data source end has an indicator ID (which can be represented by measure_id), data collection time (which can be represented by time), parameter dimension 1-parameter dimension N (N is a preset upper limit of the number of parameter dimensions, such as a positive integer such as 10, 20, etc., and is represented by dim1 to dimN respectively), indicator type (which can be represented by type), indicator numerator value (which can be represented by fz_value), and indicator denominator value (which can be represented by fm_value).
[0037] If it is necessary to collect user service data corresponding to two time points, namely, a preset observation time and a preset comparison time, from a data source, it is necessary to obtain a user service data acquisition statement whose data collection time includes at least the preset observation time and the preset comparison time. For example, the user service data acquisition statement is shown in the following example:
[0038] select fmeasure_id--Indicator ID,
[0039] ftime--time,
[0040] fdim1--dimension 1,
[0041] fdim2--dimension 2,
[0042] fdim3--dimension 3,
[0043] fdim4--dimension 4,
[0044] fdim5 -- dimension 5
[0045] ...,
[0046] fdim20 -- dimension 20 (corresponding to the example when N = 20, N can be increased as needed),
[0047] ftype--Indicator type (Indicator type is absolute value / relative value),
[0048] find_fz_value--numerator value (when the indicator is an absolute value, it is the indicator value),
[0049] find_fm_value--denominator value (when the indicator is an absolute value, it is 1)
[0050] from xxx.xxxxxx--data source table
[0051] where xxx--indicator restriction conditions
[0052] In the user service data acquisition statement shown in the example above, the xxx.xxxxxx following the from directive defines the data address of the data source table in the data source, and the xxx following the where directive limits the data collection time to a preset observation time and a preset comparison time (generally, the preset observation time and the preset comparison time are limited to the same year). The acquired user service data acquisition statement can be used for subsequent automatic execution to obtain the required user service data from the data table of the data source.
[0053] S120: If it is detected that the current system time is the preset observation time, execute the user service data acquisition statement to acquire an observation data set and a comparison data set from the data source, and merge the observation data set and the comparison data set to obtain a current data set to be processed.
[0054] Each set of comparison user service data in the current data set to be processed includes observation data and comparison data with the same indicator ID and parameter cross-combination items;
[0055] Moreover, the data collection time corresponding to the observation data set is the preset observation time, and the data collection time corresponding to the comparison data set is the preset comparison time.
[0056] In this embodiment, if it is detected that the current system time is the preset observation time, it means that the user service data acquisition statement can be automatically executed immediately to acquire the observation data set and the comparison data set from the data source.
[0057] In one embodiment, if Figure 3 As shown, step S120 includes:
[0058] S121. Obtain the data source address, the preset observation time, and the preset comparison time included in the user service data acquisition statement;
[0059] S122, acquiring the observation data set from the corresponding data source according to the data source address and the preset observation time, and acquiring the comparison data set from the corresponding data source according to the data source address and the preset comparison time;
[0060] S123 , merging the observed data set and the comparison data set in a manner of displaying user service data with the same indicator ID and the same parameter cross-combination items in parallel in the upper and lower rows to obtain the current data set to be processed.
[0061] In this embodiment, for example, when the indicator type is limited to an absolute value, the preset observation time is June 30, and the preset comparison time is May 31, if the current system time is June 30 and is the same as the preset observation time, then the user business data with a data collection time of the entire day of June 30 and meeting the corresponding parameter dimensions defined in the user business data acquisition statement is obtained from the data source and forms an observation data set. The user business data with a data collection time of the entire day of May 31 and meeting the corresponding parameter dimensions defined in the user business data acquisition statement can also be obtained from the data source and forms a comparison data set. Moreover, the obtained observation data set and comparison data set can also be merged and displayed together. For example, two user business data with the same indicator ID and parameter cross-combination items are displayed with the user business data corresponding to the current system time in the upper row, and the user business data corresponding to the preset comparison time is displayed in the lower row. The specific merged current data set to be processed is shown in Table 1 below:
[0062] Table 1
[0063]
[0064] In the current dataset to be processed, which is formed by merging the observation dataset and the comparison dataset, it can be intuitively seen that the user business data with the same indicator ID and parameter cross-combination items correspond to two time points (i.e., the preset observation time and the preset comparison time). It should be noted that although the numerator value in each user business data in Table 1 is abbreviated as xxx, it actually corresponds to different numerator values. For example, the numerator value in the first row of Table 1 is specifically the numerator value. 11 The specific value of the numerator value in the second row is the numerator value 12 The specific value of the numerator value in the third row is the numerator value 21 The specific value of the numerator value in the fourth row is the numerator value 22 Etc. Moreover, taking the first two rows of user service data in Table 1 as an example, the indicator ID of both rows of user service data is 10001, and the parameter cross-combination sub-items of both rows of user service data are "Guangdong Class A...Level I", that is, each row of user service data in Table 1 is composed of corresponding multiple parameter dimensions to form a parameter cross-combination sub-item.
[0065] For another example, when the indicator type is limited to relative value, the preset observation time is June 30, and the preset comparison time is May 31, if the current system time is June 30 and is the same as the preset observation time, the observation data set and the comparison data set are obtained from the data source based on the user business data acquisition statement, and then merged to obtain the current data set to be processed. Similar to Table 1, it should be noted that although the numerator value of each user business data in Table 2 is abbreviated as xxx and the denominator value is abbreviated as yyy, they actually correspond to different numerator and denominator values. For example, the specific value of the numerator value in the first row of Table 2 is the numerator value. 11 And the denominator is the denominator value 11 The specific value of the numerator value in the second row is the numerator value 12 And the denominator is the denominator value 12 The specific value of the numerator value in the third row is the numerator value 21 And the denominator is the denominator value 21 The specific value of the numerator value in the fourth row is the numerator value 22 And the denominator is the denominator value 22 The specific data sets currently to be processed are shown in Table 2 below:
[0066] Table 2
[0067]
[0068] However, whether the current dataset to be processed is in the form of Table 1 or the current dataset to be processed is in the form of Table 2, it does not affect the subsequent construction of the attribution number.
[0069] S130 : Obtaining an influence index corresponding to each group of comparison user service data in the current data set to be processed based on a preset smoothing proportion weighting strategy.
[0070] The influence index corresponding to each group of comparison user business data is determined by the conversion influence index and the structural influence index.
[0071] In this embodiment, still referring to the above example, if the current data set to be processed is to merge the observed data set and the comparison data set in a manner of displaying the user service data with the same indicator ID and parameter cross-combination items in parallel in the upper and lower rows, such as the first two rows of user service data in Table 1 or Table 2 can be regarded as a group of comparison user service data. After obtaining the calculation formula corresponding to the sliding proportion weighting strategy, each user service data in the above group of comparison user service data can be brought into the calculation formula to obtain the influence index corresponding to the group of comparison user service data. More specifically, the calculation formulas corresponding to the smooth proportion weighting strategy include a conversion influence index calculation formula and a structure influence index calculation formula. By substituting the group of comparison user business data into the conversion influence index calculation formula, the conversion influence index corresponding to the group of comparison user business data can be obtained, and by substituting the group of comparison user business data into the structure influence index calculation formula, the structure influence index corresponding to the group of comparison user business data can be obtained. Finally, based on the smooth proportion weighting method corresponding to the calculation formula corresponding to the smooth proportion weighting strategy, the conversion influence index and structure influence index of a group of comparison user business data can be comprehensively calculated to finally determine the influence index corresponding to the group of comparison user business data.
[0072] In one embodiment, when it is determined that the indicator type of each user service data in each group of comparison user service data in the current data set to be processed is a relative value, such as Figure 4 As shown, step S130 includes:
[0073] S131. Obtain observation data and comparison data for each set of comparison user service data in the current data set to be processed;
[0074] S132: Obtain a conversion influence acquisition strategy in the smoothed proportion weighted strategy, substitute the observed data and the comparison data into a calculation formula corresponding to the conversion influence acquisition strategy, and obtain a conversion influence index corresponding to the observed data;
[0075] S133: Obtain a structural influence acquisition strategy in the smoothed proportion weighted strategy, substitute the observed data and the comparison data into a calculation formula corresponding to the structural influence acquisition strategy, and obtain a structural influence index corresponding to the observed data;
[0076] S134 : Sum the conversion influence index and the structure influence index corresponding to the observation data to obtain an influence index corresponding to the observation data.
[0077] In this embodiment, when determining the impact index corresponding to each group of comparative user business data in the current data set to be processed, it is essentially also determining the impact index corresponding to the observation data in the group of comparative user business data. The specific process is to first obtain the observation data and comparative data in the group of comparative user business data (such as in Table 2 above, the user business data in the first row is used as the observation data, and the user business data in the second row is used as the comparative data). Then, the observation data and the comparative data are substituted into the calculation formula corresponding to the conversion influence acquisition strategy in the smoothing proportion weighting strategy to calculate the conversion influence index corresponding to the observation data, and the observation data and the comparative data are substituted into the calculation formula corresponding to the structural influence acquisition strategy in the smoothing proportion weighting strategy to calculate the structural influence index corresponding to the observation data. Finally, the conversion influence index and the structural influence index corresponding to the same observation data are summed to obtain the influence index corresponding to the observation data.
[0078] In one embodiment, the calculation formula corresponding to the conversion influence acquisition strategy is conversion influence index = 0.5*(proportion of comparison data + proportion of observation data)*(numerator value of observation data / denominator value of observation data-numerator value of comparison data / denominator value of comparison data); the calculation formula corresponding to the structural influence acquisition strategy is structural influence index = 0.5*(numerator value of observation data / denominator value of observation data + numerator value of comparison data / denominator value of comparison data)*(proportion of comparison data + proportion of observation data).
[0079] In this embodiment, if a weighted proportion strategy is used to determine the corresponding impact index for each group of comparative user service data (including one observation data and one comparative data) in the current data set to be processed, when the impact is calculated for multiple groups of comparative user service data respectively, the corresponding impact index is calculated based on the data proportion of the comparative user service data of the group * the numerator value of the observation data / the denominator value of the observation data. When the impact indexes of multiple groups of comparative user service data with the same indicator ID are weighted and summed (such as by y=sum(percentage) j *Numerator value j / denominator value j ), which accounts for j Indicates the weight ratio of the jth group of user business data in the current dataset to be processed, the numerator value j Indicates the numerator and denominator of the observed data in the jth group of user business data in the current dataset to be processed. jRepresents the denominator value of the observed data in the j-th group of comparison user business data in the current data set to be processed) to calculate the influence index corresponding to the indicator ID in the current data set to be processed. However, after decomposing the above-mentioned weighted summation expression, if there is a denominator value of 0 for the comparison data or when the observed data is 0, the conversion influence or structural influence of the corresponding group of comparison user business data obtained by using the proportion weighting strategy is 0. Moreover, when the proportion weighting strategy is adopted to determine the conversion influence or structural influence of the comparison user business data, the covariance term derived is directly discarded and then normalized or merged into the structural influence term, which will lead to nonlinear decomposition, or excessive amplification of the structural influence term. However, if the smooth proportion weighting strategy in the present application is adopted, the conversion influence or structural influence corresponding to each group of comparison user business data can be effectively smoothed, the covariance term is eliminated, and the influence of the covariance term amplification on the structural influence term is avoided.
[0080] S140 , based on a preset search strategy and the influence index corresponding to each group of compared user service data in the current data set to be processed, determine an index sub-item tree structure corresponding to each index ID in the current data set to be processed.
[0081] In this embodiment, a search strategy is also preset in the server for searching and sorting the impact indicators of each group of comparative user service data corresponding to each indicator ID in the current dataset to be processed, thereby obtaining a sub-item tree structure of indicators corresponding to each indicator ID in the current dataset to be processed. For example, in Table 2, there are five groups of comparative user service data with indicator ID 10001 (the number of five groups is for example only and is not limited to five groups in a specific implementation, and can be any other positive integer number of groups). In this case, the search strategy is combined to search the five groups of comparative user service data with indicator ID 10001 in combination with their corresponding impact indicators, thereby obtaining a sub-item tree structure of indicators corresponding to the five groups of comparative user service data with indicator ID 10001. Among them, when the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is obtained based on the search strategy, the purpose is to achieve that the influence index of the root node in the indicator sub-item tree structure corresponding to each indicator ID has the maximum value, and the parameter cross-combination sub-item corresponding to the root node is displayed in the root node of the indicator sub-item tree structure, and after determining the root node, the parameter cross-combination sub-item corresponding to the comparison user service data with the maximum value of the second-level influence index is selected as the leaf node, and so on until the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is obtained.
[0082] In one embodiment, as a first embodiment of determining the corresponding indicator item tree structure in combination with the search strategy for the comparison user service data corresponding to each indicator ID in the current data set to be processed in step S140, it includes:
[0083] If it is determined that the total number of comparison user business data corresponding to the indicator ID is greater than the preset number of groups, then based on the Monte Carlo tree search model in the search strategy and the influence index corresponding to each group of comparison user business data in the current data set to be processed, the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is determined.
[0084] In this embodiment, if the total number of comparison user service data corresponding to the indicator ID is determined to be greater than the preset number of groups (e.g., 10, which can of course be set to other positive integer values according to actual user needs during implementation), the Monte Carlo Tree Search model (i.e., MCTS model, MCTS stands for Monte Carlo TreeSearch) is preferentially used to iterate multiple groups of comparison user service data corresponding to each indicator ID in the current dataset to be processed to determine the indicator sub-item tree structure corresponding to each indicator ID. Moreover, each round of iteration in the above process is required to go through the four steps of selection, expansion, simulation, and backpropagation, thereby obtaining the indicator sub-item tree structure corresponding to each indicator ID in the current dataset to be processed.
[0085] After the indicator sub-item tree structure corresponding to each indicator ID is determined, the sub-item influence of all parameter cross-combination sub-items corresponding to the indicator ID can be determined based on the indicator sub-item tree structure, so that several parameter cross-combination sub-items with smaller node depth corresponding to the indicator ID can be further selected as the attribution analysis results.
[0086] In one embodiment, as a second embodiment of determining the corresponding indicator item tree structure in step S140 for the comparison user service data corresponding to each indicator ID in the current data set to be processed in combination with the search strategy, it includes:
[0087] If it is determined that the total number of comparison user business data corresponding to the indicator ID is less than or equal to the preset number of groups, then based on the depth-first search model in the search strategy and the influence index corresponding to each group of comparison user business data in the current data set to be processed, the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is determined.
[0088] In this embodiment, if it is determined that the total number of comparison user service data corresponding to an indicator ID is less than or equal to the preset number of groups, a depth-first search model (i.e., DFS model, DFS stands for Depth-First Search) is preferentially used to determine the indicator sub-item tree structure corresponding to each indicator ID for the multiple groups of comparison user service data corresponding to each indicator ID in the current data set to be processed. When using the depth-first search model to determine the depth of the multiple parameter cross-combination sub-items corresponding to each indicator ID in the indicator sub-item tree structure, generally, one path is followed to the end and then returns to the root node, and a second path is searched again until each parameter cross-combination sub-item in the multiple parameter cross-combination sub-items corresponding to each indicator ID is found.
[0089] Similarly, after the indicator sub-item tree structure corresponding to each indicator ID is determined, the sub-item influence of all parameter cross-combination sub-items corresponding to the indicator ID can be determined based on the indicator sub-item tree structure, so that several parameter cross-combination sub-items with smaller node depth corresponding to the indicator ID can be further selected as the attribution analysis results.
[0090] S150 , based on the indicator item tree structure corresponding to each indicator ID in the current data set to be processed and a preset node screening strategy, determine the attribution analysis result corresponding to each indicator ID.
[0091] In this embodiment, after the server determines the indicator item tree structure corresponding to each indicator ID in the current dataset to be processed, the corresponding nodes can be screened out by combining the characteristics of each node in the indicator item tree structure and a preset node screening strategy, thereby forming an attribution analysis result. Taking an indicator item tree structure as an example, when performing a horizontal comparison, among leaf nodes at the same depth, the leaf nodes that are closer to the left have a greater influence index (i.e., a greater dimensional influence), and when performing a vertical comparison, the nodes with a smaller depth have a greater influence index (i.e., a greater dimensional influence).
[0092] In one embodiment, if Figure 5 As shown, step S150 includes:
[0093] S151. Obtain the preset parameter dimension screening number corresponding to each indicator ID of the current data set to be processed in the node screening strategy;
[0094] S152: Determine the attribution analysis result corresponding to each indicator ID based on the indicator item tree structure corresponding to each indicator ID in the current data set to be processed and the number of parameter dimension screenings corresponding to each indicator ID.
[0095] In this embodiment, after the server determines the indicator item tree structure corresponding to each indicator ID in the currently processed data set, it can further obtain a preset number of parameter dimension screenings corresponding to each indicator ID (such as 3, 4, or 5, which can also be set to other positive integers based on the user's actual needs and is not limited to the above example), thereby screening out multiple parameter cross-combination items corresponding to each indicator ID as attribution analysis results. It can be seen that based on the above method, multiple important parameter dimension combinations corresponding to each indicator ID can be quickly determined as attribution analysis results.
[0096] It can be seen that the embodiment of the method can quickly combine the smoothed proportion weighting strategy and the search strategy for the current data set to be processed composed of user business data to determine the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed, thereby combining the indicator sub-item tree structure to more accurately screen and obtain quantifiable attribution analysis results.
[0097] Figure 6 This is a schematic block diagram of a multi-dimensional attribution processing device for user data provided by an embodiment of the present invention. Figure 6 As shown, corresponding to the above user data multi-dimensional attribution processing method, the present invention also provides a user data multi-dimensional attribution processing device 100. The user data multi-dimensional attribution processing device 100 includes a unit for executing the above user data multi-dimensional attribution processing method. Figure 6 The user data multi-dimensional attribution processing device 100 includes: a data acquisition statement acquisition unit 110, a current to-be-processed data set acquisition unit 120, an influence index calculation unit 130, a tree structure acquisition unit 140 and an attribution analysis result acquisition unit 150.
[0098] The data acquisition statement acquisition unit 110 is configured to, in response to a data attribution instruction, acquire a user service data acquisition statement corresponding to the data attribution instruction.
[0099] The data collection time corresponding to the user service data acquisition statement includes at least a preset observation time and a preset comparison time.
[0100] In this embodiment, the technical solution is described with the server as the execution subject. In the server, multi-dimensional attribution processing can be performed on the data collected from the data source, thereby constructing an attribution tree. Among them, each user data stored at the data source end has an indicator ID (which can be represented by measure_id), data collection time (which can be represented by time), parameter dimension 1-parameter dimension N (N is a preset upper limit of the number of parameter dimensions, such as a positive integer such as 10, 20, etc., and is represented by dim1 to dimN respectively), indicator type (which can be represented by type), indicator numerator value (which can be represented by fz_value), and indicator denominator value (which can be represented by fm_value).
[0101] If you need to collect user service data corresponding to two time points, namely, a preset observation time and a preset comparison time, from a data source, you need to obtain a user service data acquisition statement whose data collection time includes at least the preset observation time and the preset comparison time. The obtained user service data acquisition statement can be used for subsequent automatic execution to obtain the required user service data from the data table of the data source.
[0102] The current data set to be processed acquiring unit 120 is configured to execute the user service data acquiring statement to acquire an observed data set and a comparison data set from a data source if it is detected that the current system time is the preset observation time, and merge the observed data set and the comparison data set to obtain the current data set to be processed.
[0103] Each set of comparison user service data in the current data set to be processed includes observation data and comparison data with the same indicator ID and parameter cross-combination items;
[0104] Moreover, the data collection time corresponding to the observation data set is the preset observation time, and the data collection time corresponding to the comparison data set is the preset comparison time.
[0105] In this embodiment, if it is detected that the current system time is the preset observation time, it means that the user service data acquisition statement can be automatically executed immediately to acquire the observation data set and the comparison data set from the data source.
[0106] In one embodiment, the current to-be-processed data set acquisition unit 120 is configured to:
[0107] Obtaining the data source address, the preset observation time, and the preset comparison time included in the user service data acquisition statement;
[0108] Acquire the observation data set from the corresponding data source according to the data source address and the preset observation time, and acquire the comparison data set from the corresponding data source according to the data source address and the preset comparison time;
[0109] The observed data set and the comparison data set are combined in a manner of displaying user service data with the same indicator ID and the same parameter cross-combination items in parallel in the upper and lower rows to obtain the current data set to be processed.
[0110] In this embodiment, for example, when the indicator type is limited to an absolute value, the preset observation time is June 30, and the preset comparison time is May 31, if the current system time is June 30 and is the same as the preset observation time, then the user business data with a data collection time of the entire day of June 30 and meeting the corresponding parameter dimensions defined in the user business data acquisition statement is obtained from the data source and forms an observation data set. The user business data with a data collection time of the entire day of May 31 and meeting the corresponding parameter dimensions defined in the user business data acquisition statement can also be obtained from the data source and forms a comparison data set. Moreover, the obtained observation data set and comparison data set can also be merged and displayed together. For example, two user business data with the same indicator ID and parameter cross-combination items are displayed with the user business data corresponding to the current system time in the upper row, and the user business data corresponding to the preset comparison time is displayed in the lower row. The specific merged current data set to be processed is shown in Table 1 above.
[0111] In the current dataset to be processed, which is formed by merging the observation dataset and the comparison dataset, it can be intuitively seen that the user business data with the same indicator ID and parameter cross-combination items correspond to two time points (i.e., the preset observation time and the preset comparison time). It should be noted that although the numerator value in each user business data in Table 1 is abbreviated as xxx, it actually corresponds to different numerator values. For example, the numerator value in the first row of Table 1 is specifically the numerator value. 11 The specific value of the numerator value in the second row is the numerator value 12 The specific value of the numerator value in the third row is the numerator value 21 The specific value of the numerator value in the fourth row is the numerator value 22 Etc. Moreover, taking the first two rows of user service data in Table 1 as an example, the indicator ID of both rows of user service data is 10001, and the parameter cross-combination sub-items of both rows of user service data are "Guangdong Class A...Level I", that is, each row of user service data in Table 1 is composed of corresponding multiple parameter dimensions to form a parameter cross-combination sub-item.
[0112] For another example, when the indicator type is limited to relative value, the preset observation time is June 30, and the preset comparison time is May 31, if the current system time is June 30 and is the same as the preset observation time, the observation data set and the comparison data set are obtained from the data source based on the user business data acquisition statement, and then merged to obtain the current data set to be processed. Similar to Table 1, it should be noted that although the numerator value of each user business data in Table 2 is abbreviated as xxx and the denominator value is abbreviated as yyy, they actually correspond to different numerator and denominator values. For example, the specific value of the numerator value in the first row of Table 2 is the numerator value. 11 And the denominator is the denominator value 11 The specific value of the numerator value in the second row is the numerator value 12 And the denominator is the denominator value 12 The specific value of the numerator value in the third row is the numerator value 21 And the denominator is the denominator value 21 The specific value of the numerator value in the fourth row is the numerator value 22 And the denominator is the denominator value 22 The specific current dataset to be processed is shown in Table 2 above. However, whether the current dataset to be processed is in the form of Table 1 or in the form of Table 2, it does not affect the subsequent construction of the attribution number.
[0113] The influence index calculation unit 130 is configured to obtain the influence index corresponding to each set of comparison user service data in the current data set to be processed based on a preset smoothing proportion weighting strategy.
[0114] The influence index corresponding to each group of comparison user business data is determined by the conversion influence index and the structural influence index.
[0115] In this embodiment, still referring to the above example, if the current data set to be processed is to merge the observed data set and the comparison data set in a manner of displaying the user service data with the same indicator ID and parameter cross-combination items in parallel in the upper and lower rows, such as the first two rows of user service data in Table 1 or Table 2 can be regarded as a group of comparison user service data. After obtaining the calculation formula corresponding to the sliding proportion weighting strategy, each user service data in the above group of comparison user service data can be brought into the calculation formula to obtain the influence index corresponding to the group of comparison user service data. More specifically, the calculation formulas corresponding to the smoothed proportion weighting strategy include a conversion influence index calculation formula and a structural influence index calculation formula. By substituting the group of comparison user business data into the conversion influence index calculation formula, the conversion influence index corresponding to the group of comparison user business data can be obtained, and by substituting the group of comparison user business data into the structural influence index calculation formula, the structural influence index corresponding to the group of comparison user business data can be obtained. Finally, based on the smoothed proportion weighting method corresponding to the calculation formula corresponding to the smoothed proportion weighting strategy, the conversion influence index and structural influence index of a group of comparison user business data can be comprehensively calculated to finally determine the influence index corresponding to the group of comparison user business data.
[0116] In one embodiment, when it is determined that the indicator type of each user service data in each set of comparison user service data in the current data set to be processed is a relative value, the influence indicator calculation unit 130 is configured to:
[0117] For each set of comparison user service data in the current data set to be processed, observation data and comparison data are obtained;
[0118] Obtaining a conversion influence acquisition strategy in the smoothed proportion weighted strategy, substituting the observed data and the comparison data into a calculation formula corresponding to the conversion influence acquisition strategy, and obtaining a conversion influence index corresponding to the observed data;
[0119] Obtaining a structural influence acquisition strategy in the smoothed proportion weighted strategy, substituting the observed data and the comparison data into a calculation formula corresponding to the structural influence acquisition strategy, and obtaining a structural influence index corresponding to the observed data;
[0120] The conversion influence index and the structure influence index corresponding to the observation data are summed to obtain the influence index corresponding to the observation data.
[0121] In this embodiment, when determining the impact index corresponding to each group of comparative user business data in the current data set to be processed, it is essentially also determining the impact index corresponding to the observation data in the group of comparative user business data. The specific process is to first obtain the observation data and comparative data in the group of comparative user business data (such as in Table 2 above, the user business data in the first row is used as the observation data, and the user business data in the second row is used as the comparative data). Then, the observation data and the comparative data are substituted into the calculation formula corresponding to the conversion influence acquisition strategy in the smoothing proportion weighting strategy to calculate the conversion influence index corresponding to the observation data, and the observation data and the comparative data are substituted into the calculation formula corresponding to the structural influence acquisition strategy in the smoothing proportion weighting strategy to calculate the structural influence index corresponding to the observation data. Finally, the conversion influence index and the structural influence index corresponding to the same observation data are summed to obtain the influence index corresponding to the observation data.
[0122] In one embodiment, the calculation formula corresponding to the conversion influence acquisition strategy is conversion influence index = 0.5*(proportion of comparison data + proportion of observation data)*(numerator value of observation data / denominator value of observation data-numerator value of comparison data / denominator value of comparison data); the calculation formula corresponding to the structural influence acquisition strategy is structural influence index = 0.5*(numerator value of observation data / denominator value of observation data + numerator value of comparison data / denominator value of comparison data)*(proportion of comparison data + proportion of observation data).
[0123] In this embodiment, if a weighted proportion strategy is used to determine the corresponding impact index for each group of comparative user service data (including one observation data and one comparative data) in the current data set to be processed, when the impact is calculated for multiple groups of comparative user service data respectively, the corresponding impact index is calculated based on the data proportion of the comparative user service data of the group * the numerator value of the observation data / the denominator value of the observation data. When the impact indexes of multiple groups of comparative user service data with the same indicator ID are weighted and summed (such as by y=sum(percentage) j *Numerator value j / denominator value j ), which accounts for j Indicates the weight ratio of the jth group of user business data in the current dataset to be processed, the numerator value j Indicates the numerator and denominator of the observed data in the jth group of user business data in the current dataset to be processed. jRepresents the denominator value of the observed data in the j-th group of comparison user business data in the current data set to be processed) to calculate the influence index corresponding to the indicator ID in the current data set to be processed. However, after decomposing the above-mentioned weighted summation expression, if there is a denominator value of 0 for the comparison data or when the observed data is 0, the conversion influence or structural influence of the corresponding group of comparison user business data obtained by using the proportion weighting strategy is 0. Moreover, when the proportion weighting strategy is adopted to determine the conversion influence or structural influence of the comparison user business data, the covariance term derived is directly discarded and then normalized or merged into the structural influence term, which will lead to nonlinear decomposition, or excessive amplification of the structural influence term. However, if the smooth proportion weighting strategy in the present application is adopted, the conversion influence or structural influence corresponding to each group of comparison user business data can be effectively smoothed, the covariance term is eliminated, and the influence of the covariance term amplification on the structural influence term is avoided.
[0124] The tree structure acquisition unit 140 is configured to determine an indicator item tree structure corresponding to each indicator ID in the current data set to be processed based on a preset search strategy and the influence index corresponding to each group of compared user service data in the current data set to be processed.
[0125] In this embodiment, a search strategy is also preset in the server for searching and sorting the impact indicators of each group of comparative user service data corresponding to each indicator ID in the current dataset to be processed, thereby obtaining a sub-item tree structure of indicators corresponding to each indicator ID in the current dataset to be processed. For example, in Table 2, there are five groups of comparative user service data with indicator ID 10001 (the number of five groups is for example only and is not limited to five groups in a specific implementation, and can be any other positive integer number of groups). In this case, the search strategy is combined to search the five groups of comparative user service data with indicator ID 10001 in combination with their corresponding impact indicators, thereby obtaining a sub-item tree structure of indicators corresponding to the five groups of comparative user service data with indicator ID 10001. Among them, when the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is obtained based on the search strategy, the purpose is to achieve that the influence index of the root node in the indicator sub-item tree structure corresponding to each indicator ID has the maximum value, and the parameter cross-combination sub-item corresponding to the root node is displayed in the root node of the indicator sub-item tree structure, and after determining the root node, the parameter cross-combination sub-item corresponding to the comparison user service data with the maximum value of the second-level influence index is selected as the leaf node, and so on until the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is obtained.
[0126] In one embodiment, as a first embodiment of determining the corresponding indicator sub-item tree structure in the tree structure acquisition unit 140 for the comparison user service data corresponding to each indicator ID in the current data set to be processed in combination with the search strategy, it includes:
[0127] If it is determined that the total number of comparison user business data corresponding to the indicator ID is greater than the preset number of groups, then based on the Monte Carlo tree search model in the search strategy and the influence index corresponding to each group of comparison user business data in the current data set to be processed, the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is determined.
[0128] In this embodiment, if the total number of comparison user service data corresponding to the indicator ID is determined to be greater than the preset number of groups (e.g., 10, which can of course be set to other positive integer values according to actual user needs during implementation), the Monte Carlo Tree Search model (i.e., MCTS model, MCTS stands for Monte Carlo TreeSearch) is preferentially used to iterate multiple groups of comparison user service data corresponding to each indicator ID in the current dataset to be processed to determine the indicator sub-item tree structure corresponding to each indicator ID. Moreover, each round of iteration in the above process is required to go through the four steps of selection, expansion, simulation, and backpropagation, thereby obtaining the indicator sub-item tree structure corresponding to each indicator ID in the current dataset to be processed.
[0129] After the indicator sub-item tree structure corresponding to each indicator ID is determined, the sub-item influence of all parameter cross-combination sub-items corresponding to the indicator ID can be determined based on the indicator sub-item tree structure, so that several parameter cross-combination sub-items with smaller node depth corresponding to the indicator ID can be further selected as the attribution analysis results.
[0130] In one embodiment, as a second embodiment of determining the corresponding indicator sub-item tree structure in the tree structure acquisition unit 140 for the comparison user service data corresponding to each indicator ID in the current data set to be processed in combination with the search strategy, it includes:
[0131] If it is determined that the total number of comparison user business data corresponding to the indicator ID is less than or equal to the preset number of groups, then based on the depth-first search model in the search strategy and the influence index corresponding to each group of comparison user business data in the current data set to be processed, the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is determined.
[0132] In this embodiment, if it is determined that the total number of comparison user service data corresponding to an indicator ID is less than or equal to the preset number of groups, a depth-first search model (i.e., DFS model, DFS stands for Depth-First Search) is preferentially used to determine the indicator sub-item tree structure corresponding to each indicator ID for the multiple groups of comparison user service data corresponding to each indicator ID in the current data set to be processed. When using the depth-first search model to determine the depth of the multiple parameter cross-combination sub-items corresponding to each indicator ID in the indicator sub-item tree structure, generally, one path is followed to the end and then returns to the root node, and a second path is searched again until each parameter cross-combination sub-item in the multiple parameter cross-combination sub-items corresponding to each indicator ID is found.
[0133] Similarly, after the indicator sub-item tree structure corresponding to each indicator ID is determined, the sub-item influence of all parameter cross-combination sub-items corresponding to the indicator ID can be determined based on the indicator sub-item tree structure, so that several parameter cross-combination sub-items with smaller node depth corresponding to the indicator ID can be further selected as the attribution analysis results.
[0134] The attribution analysis result acquisition unit 150 is configured to determine the attribution analysis result corresponding to each indicator ID based on the indicator item tree structure corresponding to each indicator ID in the current data set to be processed and a preset node screening strategy.
[0135] In this embodiment, after the server determines the indicator item tree structure corresponding to each indicator ID in the current dataset to be processed, the corresponding nodes can be screened out by combining the characteristics of each node in the indicator item tree structure and a preset node screening strategy, thereby forming an attribution analysis result. Taking an indicator item tree structure as an example, when performing a horizontal comparison, among leaf nodes at the same depth, the leaf nodes that are closer to the left have a greater influence index (i.e., a greater dimensional influence), and when performing a vertical comparison, the nodes with a smaller depth have a greater influence index (i.e., a greater dimensional influence).
[0136] In one embodiment, the attribution analysis result obtaining unit 150 is configured to:
[0137] Obtaining the preset parameter dimension screening number corresponding to each indicator ID of the current data set to be processed in the node screening strategy;
[0138] Based on the indicator item tree structure corresponding to each indicator ID in the current data set to be processed and the number of parameter dimension screening corresponding to each indicator ID, the attribution analysis result corresponding to each indicator ID is determined.
[0139] In this embodiment, after the server determines the indicator item tree structure corresponding to each indicator ID in the currently processed data set, it can further obtain a preset number of parameter dimension screenings corresponding to each indicator ID (such as 3, 4, or 5, which can also be set to other positive integers based on the user's actual needs and is not limited to the above example), thereby screening out multiple parameter cross-combination items corresponding to each indicator ID as attribution analysis results. It can be seen that based on the above method, multiple important parameter dimension combinations corresponding to each indicator ID can be quickly determined as attribution analysis results.
[0140] It can be seen that the implementation of the embodiment of the device can quickly combine the smooth proportion weighting strategy and the search strategy for the current data set to be processed composed of user business data to determine the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed, thereby combining the indicator sub-item tree structure to more accurately screen and obtain quantifiable attribution analysis results.
[0141] The above-mentioned user data multi-dimensional attribution processing device can be implemented in the form of a computer program. The computer program can be used in Figure 7 Runs on the computer equipment shown.
[0142] See also Figure 7 , Figure 7 This is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device integrates any of the user data multi-dimensional attribution processing devices provided by an embodiment of the present invention.
[0143] See Figure 7 The computer device 400 includes a processor 402 , a memory, and a network interface 405 connected via a system bus 401 , wherein the memory may include a storage medium 403 and an internal memory 404 .
[0144] The storage medium 403 may store an operating system 4031 and a computer program 4032. The computer program 4032 includes program instructions, which, when executed, may enable the processor 402 to execute a method for multi-dimensional attribution processing of user data.
[0145] The processor 402 is used to provide computing and control capabilities to support the operation of the entire computer device.
[0146] The internal memory 404 provides an environment for the operation of the computer program 4032 in the storage medium 403. When the computer program 4032 is executed by the processor 402, the processor 402 can execute the above-mentioned user data multi-dimensional attribution processing method.
[0147] The network interface 405 is used to communicate with other devices through the network. Figure 7 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0148] The processor 402 is configured to execute a computer program 4032 stored in the memory to implement the multi-dimensional attribution processing method for user data as described above.
[0149] It should be understood that in the embodiment of the present invention, the processor 402 may be a central processing unit (CPU), and the processor 402 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0150] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0151] Therefore, the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to execute the above-mentioned multi-dimensional attribution processing method for user data.
[0152] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.
[0153] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0154] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0155] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0156] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.
[0157] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A multi-dimensional attribution processing method for user data, characterized in that: include: In response to the data attribution instruction, obtaining a user service data acquisition statement corresponding to the data attribution instruction; wherein the corresponding data collection time in the user service data acquisition statement includes at least a preset observation time and a preset comparison time; If it is detected that the current system time is the preset observation time, the user service data acquisition statement is executed to acquire an observation data set and a comparison data set from the data source, and the observation data set and the comparison data set are merged to obtain the current data set to be processed; wherein the data collection time corresponding to the observation data set is the preset observation time, and the data collection time corresponding to the comparison data set is the preset comparison time; Obtaining, based on a preset smoothing proportion weighting strategy, an influence index corresponding to each set of comparative user service data in the current data set to be processed; wherein each set of comparative user service data in the current data set to be processed includes observation data and comparative data having the same indicator ID and parameter cross-combination sub-items; Based on a preset search strategy and the influence index corresponding to each group of comparison user service data in the current data set to be processed, determine the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed; Based on the indicator item tree structure corresponding to each indicator ID in the current data set to be processed and a preset node screening strategy, the attribution analysis result corresponding to each indicator ID is determined.
2. The method according to claim 1, characterized in that The executing the user service data acquisition statement to acquire an observation data set and a comparison data set from a data source, and merging the observation data set and the comparison data set to obtain a current data set to be processed, includes: Obtaining the data source address, the preset observation time, and the preset comparison time included in the user service data acquisition statement; Acquire the observation data set from the corresponding data source according to the data source address and the preset observation time, and acquire the comparison data set from the corresponding data source according to the data source address and the preset comparison time; The observed data set and the comparison data set are combined in a manner of displaying user service data with the same indicator ID and the same parameter cross-combination items in parallel in the upper and lower rows to obtain the current data set to be processed.
3. The method according to claim 1, characterized in that The influence index corresponding to each group of comparison user service data is determined by the conversion influence index and the structural influence index; when it is determined that the indicator type of each user service data in each group of comparison user service data in the current data set to be processed is a relative value, the influence index corresponding to each group of comparison user service data in the current data set to be processed based on the preset smoothing proportion weighting strategy includes: For each set of comparison user service data in the current data set to be processed, observation data and comparison data are obtained; Obtaining a conversion influence acquisition strategy in the smoothed proportion weighted strategy, substituting the observed data and the comparison data into a calculation formula corresponding to the conversion influence acquisition strategy, and obtaining a conversion influence index corresponding to the observed data; Obtaining a structural influence acquisition strategy in the smoothed proportion weighted strategy, substituting the observed data and the comparison data into a calculation formula corresponding to the structural influence acquisition strategy, and obtaining a structural influence index corresponding to the observed data; The conversion influence index and the structure influence index corresponding to the observation data are summed to obtain the influence index corresponding to the observation data.
4. The method according to claim 3, characterized in that The calculation formula corresponding to the conversion influence acquisition strategy is conversion influence index = 0.5*(proportion of comparison data + proportion of observation data)*(numerator value of observation data / denominator value of observation data-numerator value of comparison data / denominator value of comparison data); the calculation formula corresponding to the structure influence acquisition strategy is structure influence index = 0.5*(numerator value of observation data / denominator value of observation data + numerator value of comparison data / denominator value of comparison data)*(proportion of comparison data + proportion of observation data).
5. The method according to claim 1, wherein The determining of the index sub-item tree structure corresponding to each index ID in the current data set to be processed based on the preset search strategy and the influence index corresponding to each group of compared user service data in the current data set to be processed includes: If it is determined that the total number of comparison user business data corresponding to the indicator ID is greater than the preset number of groups, then based on the Monte Carlo tree search model in the search strategy and the influence index corresponding to each group of comparison user business data in the current data set to be processed, the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is determined.
6. The method according to claim 1, characterized in that The determining of the index sub-item tree structure corresponding to each index ID in the current data set to be processed based on the preset search strategy and the influence index corresponding to each group of compared user service data in the current data set to be processed includes: If it is determined that the total number of comparison user business data corresponding to the indicator ID is less than or equal to the preset number of groups, then based on the depth-first search model in the search strategy and the influence index corresponding to each group of comparison user business data in the current data set to be processed, the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed is determined.
7. The method according to claim 1, characterized in that The determining of the attribution analysis result corresponding to each indicator ID based on the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed and a preset node screening strategy includes: Obtaining the preset parameter dimension screening number corresponding to each indicator ID of the current data set to be processed in the node screening strategy; Based on the indicator item tree structure corresponding to each indicator ID in the current data set to be processed and the number of parameter dimension screening corresponding to each indicator ID, the attribution analysis result corresponding to each indicator ID is determined.
8. A multi-dimensional attribution processing device for user data, characterized in that: include: A data acquisition statement acquisition unit is configured to, in response to a data attribution instruction, acquire a user service data acquisition statement corresponding to the data attribution instruction; wherein the corresponding data collection time in the user service data acquisition statement includes at least a preset observation time and a preset comparison time; a current data set to be processed acquiring unit, configured to, if it is detected that the current system time is the preset observation time, execute the user service data acquiring statement to acquire an observation data set and a comparison data set from the data source, and merge the observation data set and the comparison data set to obtain the current data set to be processed; wherein the data collection time corresponding to the observation data set is the preset observation time, and the data collection time corresponding to the comparison data set is the preset comparison time; an influence index calculation unit, configured to obtain, based on a preset smoothed proportion weighting strategy, an influence index corresponding to each set of comparative user service data in the current data set to be processed; wherein each set of comparative user service data in the current data set to be processed includes observation data and comparative data having the same indicator ID and parameter cross-combination sub-items; A tree structure acquisition unit, configured to determine an indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed based on a preset search strategy and an influence index corresponding to each group of compared user service data in the current data set to be processed; The attribution analysis result acquisition unit is used to determine the attribution analysis result corresponding to each indicator ID based on the indicator sub-item tree structure corresponding to each indicator ID in the current data set to be processed and a preset node screening strategy.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the user data multi-dimensional attribution processing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the method for multi-dimensional attribution processing of user data according to any one of claims 1 to 7 can be implemented.