A user consumption data analysis processing method for a cake mall system
By constructing a unified user consumption dataset in the cake e-commerce system and performing data cleaning and standardization, combined with causal inference models and user group segmentation, the problems of data consistency and causal inference were solved, achieving high accuracy and consistency analysis of user consumption data, dynamically updating user profiles, and identifying consumption driving factors and effect differences.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING LIFE MODEL NETWORK TECHNOLOGY CO LTD
- Filing Date
- 2025-11-03
- Publication Date
- 2026-06-12
AI Technical Summary
Existing technologies struggle to effectively identify the consistency between user behavior response patterns and dynamic consumption paths in cake e-commerce systems, leading to inconsistent data quality, affecting the reliability of analysis results, and lacking a causal inference mechanism for potential influencing factors, thus failing to accurately reflect the heterogeneity of user groups with different characteristics.
Collect and integrate multi-source consumption data to construct a unified user consumption dataset. Through data cleaning and standardization, an initial set of user characteristics is formed. Data quality is corrected by combining user group segmentation and dynamic paths of consumption behavior over time. A causal inference model is constructed using propensity score matching to identify potential influencing factors and calculate individual-level effect biases. Heterogeneous effect analysis is conducted to dynamically update user profiles.
It improves the integrity and accuracy of data, reduces analytical bias, ensures the accuracy and consistency of user feature classification, can identify the differences in consumption drivers and effect strength of different feature subgroups, dynamically reflects changes in user features, and enhances the structural consistency and classification accuracy of user data.
Smart Images

Figure CN121563582B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method for analyzing and processing user consumption data in a cake e-commerce system. Background Technology
[0002] In the analysis of user consumption data in cake e-commerce systems, most methods rely on integrating multi-source data and extracting basic features to build user profiles or segment groups. However, when processing multi-source consumption data, such methods may struggle to effectively identify the consistency between user behavior response patterns and dynamic consumption paths, leading to inconsistent data quality and affecting the reliability of the analysis results. For example, when analyzing core indicators such as user spending amount and frequency, existing methods mostly focus on correlation analysis, lacking causal inference mechanisms for potential influencing factors. Due to the failure to fully consider the effect bias caused by feature differences at the individual level, the analysis results may not accurately reflect the heterogeneity of user groups with different characteristics. Summary of the Invention
[0003] The technical problem to be solved by this invention is to provide a method for analyzing and processing user consumption data in a cake e-commerce system, thereby constructing a unified user consumption dataset and improving data integrity, accuracy, and consistency.
[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0005] Firstly, a method for analyzing and processing user consumption data in a cake e-commerce system, the method comprising:
[0006] Step 1: Collect and integrate multi-source consumption data of users in the cake mall system, including order data, payment data, membership data and behavior log data, to build a unified user consumption dataset;
[0007] Step 2: Clean and standardize the user consumption dataset to form an initial user feature set; based on the initial user feature set, identify different user behavior response patterns by segmenting user groups, and perform data quality correction by combining the dynamic path of consumption behavior in the time dimension to obtain preprocessed user feature data with consistent quality.
[0008] Step 3: Based on the preprocessed user characteristic data, a causal inference model is constructed using the propensity score matching method. The potential influencing factors of the core indicators of user consumption amount and frequency are analyzed, the overall average treatment effect of each factor on consumption performance is estimated, and the effect bias caused by individual differences in characteristics is calculated to obtain the characteristic influence effect estimation results including the overall effect and the individual effect.
[0009] Step 4: Correlate the feature impact effect estimation results with the preprocessed user feature data, and conduct heterogeneity effect analysis by dividing the user feature subgroups to identify the differences in driving factors and effect strength of different feature subgroups in terms of consumption amount and frequency indicators, thus forming the heterogeneity effect analysis results.
[0010] Step 5: Based on the heterogeneity effect analysis results, perform fine-grained segmentation and labeling of user characteristics, and dynamically update user profiles to achieve structural consistency and classification accuracy of user data.
[0011] In a second aspect, a computing device includes:
[0012] One or more processors;
[0013] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0014] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0015] The above-described solution of the present invention has at least the following beneficial effects:
[0016] By collecting multi-source consumption data such as orders, payments, memberships, and behavior logs and constructing a unified dataset, we reduce analytical omissions caused by data dispersion and analytical biases caused by differences in data sources. Through cleaning, standardization, and quality correction, we reduce issues such as inconsistent data formats and inconsistent indicator definitions, making the initial user feature set more accurate. At the same time, we combine user group segmentation and dynamic paths of consumption behavior over time to correct data quality and reduce data biases caused by differences in user behavior patterns or fluctuations in time series data.
[0017] By constructing a causal inference model, we identify potential influencing factors of core indicators such as user spending amount and frequency. By estimating the overall average treatment effect and individual effect bias, we not only determine the overall strength of each influencing factor on consumption performance but also capture the effect differences caused by different user characteristics. This makes the analysis of influencing factors more targeted and avoids analytical errors caused by ignoring individual heterogeneity. By using the estimation results of the influence effect of related features and preprocessed user feature data, we perform user feature subgrouping and heterogeneity effect analysis. This identifies the differences in the driving factors of consumption indicators and the differences in effect strength among different subgroups. Based on the heterogeneity effect analysis results, we perform fine-grained segmentation and labeling of user features, making user feature classification more consistent with the actual behavioral patterns and effect characteristics of each subgroup. This reduces the user data classification bias caused by coarse classification. At the same time, dynamically updating user profiles can reflect changes in user characteristics and behavioral patterns in a timely manner, ensuring that user data always maintains structural consistency and classification accuracy, and reducing data classification lag or bias caused by dynamic changes in user features. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a user consumption data analysis and processing method for a cake e-commerce system provided by an embodiment of the present invention.
[0019] Figure 2 This is a schematic diagram provided by an embodiment of the present invention, which illustrates how user characteristics are finely segmented and labeled based on the results of heterogeneity effect analysis, and user profiles are dynamically updated to achieve structural consistency and classification accuracy of user data. Detailed Implementation
[0020] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0021] like Figure 1 As shown, an embodiment of the present invention proposes a method for analyzing and processing user consumption data in a cake e-commerce system. The method includes the following steps:
[0022] Step 1: Collect and integrate multi-source consumption data of users in the cake mall system, including order data, payment data, membership data and behavior log data, to build a unified user consumption dataset;
[0023] Step 2: Clean and standardize the user consumption dataset to form an initial user feature set; based on the initial user feature set, identify different user behavior response patterns by segmenting user groups, and perform data quality correction by combining the dynamic path of consumption behavior in the time dimension to obtain preprocessed user feature data with consistent quality.
[0024] Step 3: Based on the preprocessed user characteristic data, a causal inference model is constructed using the propensity score matching method. The potential influencing factors of the core indicators of user consumption amount and frequency are analyzed, the overall average treatment effect of each factor on consumption performance is estimated, and the effect bias caused by individual differences in characteristics is calculated to obtain the characteristic influence effect estimation results including the overall effect and the individual effect.
[0025] Step 4: Correlate the feature impact effect estimation results with the preprocessed user feature data, and conduct heterogeneity effect analysis by dividing the user feature subgroups to identify the differences in driving factors and effect strength of different feature subgroups in terms of consumption amount and frequency indicators, thus forming the heterogeneity effect analysis results.
[0026] Step 5: Based on the heterogeneity effect analysis results, perform fine-grained segmentation and labeling of user characteristics, and dynamically update user profiles to achieve structural consistency and classification accuracy of user data.
[0027] In this embodiment of the invention, by collecting multi-source consumption data such as orders, payments, memberships, and behavior logs and constructing a unified dataset, the problem of analytical omissions caused by data dispersion is reduced, and the analytical bias caused by differences in data sources is reduced. Cleaning, standardization, and quality correction reduce problems such as inconsistent data formats and inconsistent indicator definitions, making the initial user feature set more accurate. At the same time, data quality correction is carried out by combining user group segmentation and dynamic paths of consumption behavior over time, reducing data bias caused by differences in user behavior patterns or fluctuations in time series data.
[0028] By constructing a causal inference model, we identify potential influencing factors of core indicators such as user spending amount and frequency. By estimating the overall average treatment effect and individual effect bias, we not only determine the overall strength of each influencing factor on consumption performance but also capture the effect differences caused by different user characteristics. This makes the analysis of influencing factors more targeted and avoids analytical errors caused by ignoring individual heterogeneity. By using the estimation results of the influence effect of related features and preprocessed user feature data, we perform user feature subgrouping and heterogeneity effect analysis. This identifies the differences in the driving factors of consumption indicators and the differences in effect strength among different subgroups. Based on the heterogeneity effect analysis results, we perform fine-grained segmentation and labeling of user features, making user feature classification more consistent with the actual behavioral patterns and effect characteristics of each subgroup. This reduces the user data classification bias caused by coarse classification. At the same time, dynamically updating user profiles can reflect changes in user characteristics and behavioral patterns in a timely manner, ensuring that user data always maintains structural consistency and classification accuracy, and reducing data classification lag or bias caused by dynamic changes in user features.
[0029] In a preferred embodiment of the present invention, step 1 above, which involves collecting and integrating multi-source consumption data of users in the cake e-commerce system, including order data, payment data, membership data, and behavior log data, to construct a unified user consumption dataset, may include:
[0030] In this embodiment of the invention, order data focuses on the basic transaction information generated after a user places an order. This includes collecting the order number, user ID, order time, main product information, paid accessory information, order status, and delivery requirements. It also records the association between the default tableware configuration and paid accessories to ensure the ability to analyze the combined purchase logic of the main product and accessories. Payment data revolves around the order payment process, collecting the payment order number, user ID, payment time, payment amount, payment method, discount information, actual payment amount, and payment status. This comprehensively records the composition of the payment amount and the use of discounts, providing a basis for analyzing the impact of discounts on consumption decisions. Membership data centers on the user's membership identity, collecting user ID, membership level, membership registration time, membership validity period, membership points, and membership benefit records. It also associates this with basic user information to determine the relationship between membership identity and consumption behavior. Behavior log data tracks all user operations within the e-commerce system, collecting user ID, behavior occurrence time, behavior type, behavior association information, and device information to reconstruct the complete behavioral path before the user places an order.
[0031] For different data generation scenarios, corresponding data collection channels are established to ensure real-time and accurate data acquisition. Order data collection relies on the cake mall's order generation model. When a user completes an order placement, the system automatically triggers the data collection mechanism, writing relevant order field information into the order database in real time. Simultaneously, when the order status changes, the system synchronously captures the status change record and corresponding time, updating it to the order database to ensure the timeliness and completeness of order data. Payment data collection connects the data interface between the mall system and the payment system. When a user initiates a payment operation, the payment system synchronizes the payment request information to the mall's data collection layer in real time. After payment is completed, the payment platform immediately provides feedback on the payment result. After receiving the data, the collection layer correlates and matches it with the corresponding order number and writes it into the payment database, ensuring that payment data and order data are consistent. According to the one-to-one correspondence; member data collection is based on the mall's member management system. When a user registers as a member, basic member information is collected and stored. When a member's level changes, points change, or member benefits are used, the member system records the changes, time of change, and related scenarios in real time, and synchronizes the data to the member database to ensure dynamic updates of member data. Behavior log data collection embeds a behavior tracking component in the front-end page of the mall system. When a user performs operations such as browsing, clicking, searching, and adding to cart, the component captures the behavior events in real time, carrying the user ID, behavior time, behavior type, and related information, and sends them to the back-end behavior log collection server through asynchronous requests. The collection server receives and stores the behavior data in chronological order to form the raw behavior log library, while filtering invalid behavior data to ensure the validity of the behavior data.
[0032] By using a unified association identifier, order, payment, membership, and behavior log data scattered across different databases are linked to construct a unified user consumption dataset. The user ID serves as the core association identifier across data types, linking membership data, behavior log data, and user identity. The order number acts as the bridge between order data and payment data, ensuring accurate matching of payment information for each order. Simultaneously, in the behavior log data, corresponding order numbers are added to record order-related behaviors such as adding to cart and placing orders, establishing a link between behavior logs and order data, forming a link between user ID, order number, and multi-source data. The raw data from each data source is cleaned, handling data anomalies. For example, duplicate order records in order data are deleted, abnormal amounts in payment data are corrected, missing fields in membership data are supplemented, and invalid data in behavior logs is filtered out. Furthermore, data formats are standardized, such as adjusting all time fields to year-month-day hour:minute:second format, and uniformly retaining two decimal places in amount fields, ensuring consistency across different data sources.
[0033] By matching order information in the order database with payment information in the payment database using order numbers, an order-payment association dataset is formed, clearly defining the payment details for each order. By matching member levels, points usage records, and other information from the member database to the order-payment dataset using user IDs, new fields such as member level and points deduction amount are added, forming a member-order-payment association dataset to bind user identity with transaction information. By matching user behavior log data, such as browsing, searching, and adding items to cart before placing an order, to the corresponding user's member-order-payment data using user IDs. For behaviors with generated order numbers, precise association is achieved directly using the order number; for behaviors without generated order numbers, matching user IDs and behavior times with the time range of the user's subsequent orders ultimately forms a unified user consumption dataset containing user identity, behavior path, transaction details, and payment discounts.
[0034] In a preferred embodiment of the present invention, step 2 above involves cleaning and standardizing the user consumption dataset to form an initial user feature set; based on the initial user feature set, different user behavior response patterns are identified through user group segmentation, and data quality is corrected by combining the dynamic path of consumption behavior over time to obtain preprocessed user feature data of consistent quality, which may include:
[0035] In this embodiment of the invention, step 220 involves performing missing value imputation, outlier correction, and duplicate record deletion operations on the user consumption dataset to obtain cleaned user consumption data. Specifically, this includes: first, processing missing values. For numeric fields in the dataset, such as single purchase amount, quantity of goods in an order, payment duration, etc., first summing all non-missing values under that field, dividing the sum by the number of non-missing values to obtain the average, and then using this average to fill all missing values in that field. For categorical fields, such as payment method, delivery address area, category of purchased goods, etc., counting the frequency of each category in that field, identifying the category with the highest frequency, and using this category to fill the missing values in that field. Next, outlier correction is performed. For numeric fields, first calculating the average and standard deviation of all non-missing values in that field, and then calculating... The method involves dividing the sum of all non-missing values by the number to obtain the average. Then, the square of the difference between each non-missing value and the average is calculated. These squared values are summed and divided by the number to obtain the variance. The square root of the variance is the standard deviation. Values exceeding the average plus three standard deviations or falling below the average minus three standard deviations are marked as outliers. For outliers exceeding the upper limit, the maximum value of the field is used to replace them; for outliers falling below the lower limit, the minimum value of the field is used to replace them. Finally, duplicate records are deleted. Each record in the dataset is checked row by row. Records with identical user ID, order generation time, order number, SKU code of purchased goods, and total consumption amount are identified as duplicate records. For duplicate records, only the one with the earliest generation time is kept, and the rest are deleted. After the above operations, the cleaned user consumption data is obtained.
[0036] Step 221: Based on the cleaned user consumption data, calculate the consumption frequency, average consumption amount, and most recent consumption time interval. Combine this with membership level information and platform behavior records to construct a user feature set. Specifically, this includes: when calculating the consumption frequency, for each user, count the number of valid orders generated by that user in the past 30 days; this number represents the user's consumption frequency. When calculating the average consumption amount, count the total amount of all valid orders for that user in the past 30 days, divide this total by the number of valid orders, and the result is the user's average consumption amount. If a user has no valid orders in the past 30 days, the average consumption amount is recorded as 0. When calculating the most recent consumption time interval, first obtain the time of the user's last valid consumption order, then determine the current system date, and subtract the last valid consumption time interval from the current system date. The number of days between a valid order and a recent purchase is the most recent purchase interval. If a user has not made any valid purchases in the past 30 days, the most recent purchase interval is recorded as 31 days. Membership level information for each user is directly extracted from the cleaned user consumption data, such as ordinary member, silver member, gold member, diamond member, etc. Platform behavior records for each user are statistically analyzed from the platform's behavior logs, including the number of times the user viewed product detail pages in the past 30 days, the number of times the user added products to the shopping cart but did not complete the checkout, the number of times the user clicked on the homepage promotional banner, and the number of times the user participated in member points redemption activities. The consumption frequency, average consumption amount, most recent purchase interval, extracted membership level information, and statistical platform behavior records obtained above are integrated together to form a feature set for each user, i.e., the user feature set.
[0037] Step 222: Standardize and normalize the numerical features in the user feature set, and perform one-hot encoding on the categorical features to generate the initial user feature set. Specifically, this includes: first, distinguishing between numerical and categorical features in the user feature set. Values such as purchase frequency, average purchase amount, recent purchase time interval, number of times browsing product detail pages, number of times adding items to cart but not yet settled, number of times clicking promotional banners, and number of times participating in points redemption are numerical features, while membership level is a categorical feature. When standardizing the numerical features, for each feature, first calculate the mean and standard deviation of all users' values for that feature. The calculation method is to divide the sum of all users' feature values by the number of users to obtain the mean, then calculate the square of the difference between each user's feature value and the mean, sum these squares, and divide by the number of users to obtain the variance. The square root of the variance is the standard deviation. Then, subtract the mean from each user's feature value, and divide the difference by the standard deviation. The result is the standardized value for that user on that feature. The standardized numerical features are then normalized. For each numerical feature, the maximum and minimum standardized values for all users on that feature are found. The minimum value is subtracted from each user's standardized value, and the difference is divided by the difference between the maximum and minimum values. The result is the normalized value for that user on that feature, ensuring that all values for that feature are between 0 and 1. When performing one-hot encoding on categorical features, if the membership level includes four categories: Ordinary Member, Silver Member, Gold Member, and Diamond Member, a new feature field is created for each category. When a user's membership level is Ordinary Member, the new feature field corresponding to Ordinary Member has a value of 1, while the new feature fields corresponding to Silver Member, Gold Member, and Diamond Member have a value of 0. When a user's membership level is Silver Member, the new feature field corresponding to Silver Member has a value of 1, while the other three new feature fields have a value of 0, and so on. All the standardized and normalized numerical features, as well as all the one-hot encoded categorical features, are integrated together to form the initial user feature set.
[0038] Step 223: Calculate the feature distances for each user in the initial user feature set across multiple dimensions, including purchase frequency, average purchase amount, recent purchase time interval, membership level, and behavioral indicators. Based on a preset distance threshold, group users with similar features into the same group to generate user group segmentation results for data quality verification. Specifically, this includes: determining the dimensions for which feature distances need to be calculated, including purchase frequency, average purchase amount, recent purchase time interval, one-hot encoded membership level field, number of times browsing product detail pages, number of times adding items to cart but not yet settled, number of times clicking promotional banners, and number of times participating in points redemption. For any two users, calculate the feature distance for each dimension. The distance for numerical dimensions is the absolute difference between the normalized values of the two users in that dimension. For example, if user A's purchase frequency normalized value is 0.6 and user B's is 0.4, the distance for that dimension is 0.6 minus the absolute value of 0.4, which is 0.2. The distance for the one-hot encoded membership level dimension is the difference between the values of the two users in that field. For example, if user A's ordinary membership field is 1 and user B's ordinary membership field is 0, the distance for that dimension is 1. Subtracting the difference of 0 by 1, if two users have the same value in this field, the distance is 0. Add the feature distances across all dimensions to obtain the total feature distance between the two users. For example, if the distances of the two users in the above example are 0.2, 0.1, 0.3, 1, 0.2, 0.1, 0.1, and 0.2 respectively across the eight dimensions, the total feature distance is 0.2 + 0.1 + 0.3 + 1 + 0.2 + 0.1 + 0.1 + 0.2 = 2.2. A preset feature distance threshold is set, which is determined based on the distribution of the total feature distance between all users. For example, calculating the total feature distance between all users... The average total feature distance between users is used as the threshold. If the average is 5, then the threshold is set to 2.5. When the total feature distance between two users is less than or equal to 2.5, the two users are considered to have similar features. Following the above method, all users in the dataset are compared pairwise, and users with similar features are grouped into the same group. If user A is similar to user B, and user B is similar to user C, then users A, B, and C are grouped into the same group, ultimately forming multiple user groups, which is the result of user group division for data quality inspection.
[0039] Step 224: For each user group, extract the consumption characteristic change sequence of the corresponding group within a continuous time window to generate dynamic path features reflecting the evolution of consumption behavior of each group. Specifically, this includes: setting continuous time windows, each lasting 7 days, for a total of 12 consecutive time windows covering the past 84 days. The first time window is from day 71 to day 84, the second from day 64 to day 70, the third from day 57 to day 63, and so on, with the twelfth time window being the most recent 7 days. For each user group, within each time window, calculate the average consumption frequency of that group by summing the number of valid orders placed by all users in that group within that time window and dividing by the total number of users in that group. Also, calculate the average consumption amount for that group by summing the average consumption amount placed by all users in that group within that time window. (The total spending amount of each user in this window is divided by the number of valid orders). These average spending amounts are summed and then divided by the total number of users in the group. The average most recent spending interval for this group is calculated. The most recent spending intervals for all users in this group within this time window are counted. These intervals are summed and divided by the total number of users in the group. The average spending frequency of this group within the 12 time windows is arranged sequentially (from the first to the twelfth), forming a sequence of changes in the group's spending frequency. Similarly, the average spending amount within the 12 time windows is arranged sequentially, forming a sequence of changes in the average spending amount. The average most recent spending interval within the 12 time windows is arranged sequentially, forming a sequence of changes in the average most recent spending interval. These three sequences are integrated together as a dynamic path feature reflecting the evolution of the group's consumption behavior over a continuous period. Each user group corresponds to one dynamic path feature.
[0040] Step 225 involves comparing the consumption characteristic change sequences of each user with the dynamic path characteristics point by point to identify abnormal data points that deviate from the expected evolution pattern. Data smoothing and trend calibration are then performed to generate preprocessed user characteristic data of consistent quality. Specifically, for each user, based on their user group, within the 12 consecutive time windows set in step 224, the user's consumption frequency (number of valid orders in each window), average consumption amount (total consumption amount divided by the number of valid orders in each window), and most recent consumption time interval (most recent consumption time interval in each window) are extracted and arranged in chronological order (from the first to the twelfth) to form the user's consumption characteristics. The feature change sequence compares the user's consumption characteristic change sequence with the dynamic path characteristics of their group point by point. Specifically, it compares the user's consumption frequency within the first time window with the average consumption frequency of the group within the first window of the dynamic path, calculates the difference, and subtracts the group's average consumption frequency from the user's consumption frequency. Similarly, it compares the consumption frequency differences within the second to twelfth time windows, as well as the differences in average consumption amount and the most recent consumption time interval across all time windows. A threshold is preset for each feature difference; for example, the consumption frequency difference threshold is 40% of the group's average consumption frequency for the corresponding window. If the absolute value of the difference between a user's consumption frequency in a certain window and the group's average consumption frequency exceeds the group's average consumption frequency for that window... If a user's consumption frequency exceeds 40% of the average consumption frequency in a given window, that user's consumption frequency value in that window is marked as an outlier. The same method is used to determine anomalies in average spending amount and the most recent spending time interval; if the absolute value of the difference exceeds 40% of the average value of the corresponding window for the group, it is marked as an anomaly. Data smoothing is performed on the marked outlier data points. If the outlier data point is in the 5th time window, the user's corresponding feature values from the 4th and 6th windows are added together and divided by 2; the result replaces the outlier value in the 5th window. If the outlier window is the 1st, the feature value from the 2nd window is used; if the outlier window is the 12th, the feature value from the 11th window is used. Trend correction is then applied to the smoothed user consumption characteristic change sequence. First, analyze the overall trend of the dynamic path of the group. For example, calculate the average consumption frequency of the group in the first 6 windows and the average consumption frequency of the last 6 windows. If the average consumption frequency of the last 6 windows is 30% higher than that of the first 6 windows, it is determined that the group's consumption frequency is on the rise. If the average consumption frequency of the last 6 windows in the smoothed consumption frequency sequence of the user is 10% lower than that of the first 6 windows, which is inconsistent with the group trend, then increase the consumption frequency value of the last 6 windows of the user by 15% so that the average consumption frequency of the last 6 windows of the adjusted user sequence is 5% higher than that of the first 6 windows, which is consistent with the group's upward trend. After the above anomaly correction, smoothing and trend calibration processing, the feature data of all users are integrated to obtain preprocessed user feature data of consistent quality.
[0041] By filling in missing information in user consumption data, correcting abnormal values that do not fall within the normal range, and removing duplicate and redundant records, noise and errors in the data are reduced from the source, ensuring the integrity and accuracy of the data. By comparing the dynamic paths of individuals and groups and correcting abnormal data, the preprocessed user characteristic data maintains consistency in time trends and group consistency, improving the stability of data quality and avoiding misjudgments of user consumption behavior due to inconsistent data quality.
[0042] In a preferred embodiment of the present invention, step 3 above, based on preprocessed user characteristic data, uses propensity score matching to construct a causal inference model, analyzes the potential influencing factors of the core indicators of user consumption amount and frequency, estimates the overall average treatment effect of each factor on consumption performance, and calculates the effect bias caused by individual differences in characteristics, obtaining the characteristic influence effect estimation result including the overall effect and the individual effect, which may include:
[0043] In this embodiment of the invention, step 330 involves determining potential influencing factors to be analyzed as processing variables based on preprocessed user feature data, and using the remaining features as matching covariates. Specifically, this includes: filtering potential influencing factors strongly correlated with the cake mall consumption scenario from the preprocessed user feature data, such as whether the user has participated in a birthday cake special discount activity in the past 30 days, whether they have used a 200-50 member coupon, whether they have received a new cake push SMS, etc., and determining these factors as processing variables. Each processing variable is a binary feature, marked as 1 if yes and 0 if no. Other features besides the processing variables are used as matching covariates, including the user's consumption frequency, average consumption amount, recent consumption time interval, total time spent browsing the cake category details page in the past 90 days, the percentage of orders paid for within 24 hours after adding items to the cart, the number of times the homepage promotion zone is clicked, and the membership level after unique hot coding, etc. All user feature data are classified and stored according to two categories of fields: processing variables and matching covariates, to ensure that each piece of data for each user can be clearly matched to its category.
[0044] Step 331: Using matching covariates, estimate the probability of each user being assigned to a processing group through a logistic regression model to obtain the user processing group assignment probability distribution. Specifically, this includes: constructing a logistic regression model for each processing variable. The core of the logistic regression model is to calculate the probability of a user being assigned to a processing group using matching covariate data. The process consists of four steps: First, preprocessing the input and output data of the logistic regression model. Determine the dependent variable of the logistic regression model as whether the user belongs to a processing group, denoted as Y. The processing method is: if the user's processing variable value is 1, they belong to the processing group, and Y is 1; if the user's processing variable value is 0, they belong to the control group, and Y is 0. Iterate through all users in the preprocessed user feature data, determine the corresponding Y value for each user, and generate a dependent variable dataset containing user ID and Y value. Determine the independent variables of the logistic regression model as all matching covariates, denoted as X, specifically including consumption frequency X1, average consumption amount X2, recent consumption time interval X3, browsing history X4, and other parameters. The calculations are as follows: Duration X4, Add-to-Cart Conversion Rate X5, Promotion Clicks X6, Membership Level After Hot-Coded Subscription X7 to X10. Each independent variable is standardized using the following logic: First, calculate the mean of all values for that independent variable (the sum of all user values divided by the total number of users). Then, calculate the standard deviation of all values for that independent variable (the sum of the squares of the differences between each user's value and the mean, divided by the total number of users, and then taking the square root). Finally, subtract the mean from each user's original value for that independent variable, and divide the difference by the standard deviation to obtain the standardized value for that user. For example, when calculating the standardized value of X1, first calculate the mean of all user purchase frequencies, then calculate the standard deviation of all user purchase frequencies, and then subtract the mean from each user's original purchase frequency, dividing the result by the standard deviation to obtain the standardized value for user X1. This process ensures that the influence weights of independent variables of different magnitudes on the model are comparable, generating a dataset containing user IDs and the standardized values of each independent variable.
[0045] The second step is to solve for the parameters of the logistic regression model. The parameters include the intercept term β0 and the regression coefficients β1 to β10 of each independent variable. These are solved iteratively using the maximum likelihood estimation method. Specifically, first, all parameters are initialized to 0. Then, for each user, a composite score is calculated using the current parameter values. This is done by adding the intercept term β0 to β1 multiplied by the user's standardized X1 value, then adding β2 multiplied by the user's standardized X2 value, and so on, until β10 multiplied by the user's standardized X10 value is added. Next, the composite score is substituted into the logistic function to calculate the predicted probability. The logic of the logistic function is as follows: first, calculate the negative power of the composite score (i.e., the natural constant raised to the power of (0 minus the composite score)); then add this result to 1 to obtain a sum; finally, divide this sum by 1 to obtain the predicted probability for that user. Afterward, the likelihood function value is calculated by taking the predicted probability of each user raised to the power of the Y value, then subtracting the predicted probability from the sum of the predicted probability and the Y value. The likelihood value for a user is obtained by multiplying the two results by a power of 1e-6. Then, the likelihood values for all users are multiplied to obtain the likelihood function value for the entire model. To determine if the likelihood function has converged, the likelihood function value of the current iteration is compared with that of the previous iteration. If the difference is greater than 1e-6, it indicates that convergence has not occurred and parameter adjustments are needed. Parameter adjustments use the gradient ascent method. The new value of each parameter is equal to its old value plus the learning rate multiplied by the partial derivative of the likelihood function with respect to that parameter. The learning rate is set to 0.01. The derivative is calculated as follows: for all users, subtract the predicted probability from the user's Y value, multiply by the value of the user's corresponding independent variable (if it is the intercept term β0, multiply by 1), and then add this result for all users. The sum is the partial derivative of the likelihood function with respect to the parameter. Repeat the steps of calculating the comprehensive score, predicted probability, likelihood function value, and parameter adjustment until the difference of the likelihood function values between two iterations is less than or equal to 1e-6. At this point, the likelihood function converges, and the final parameter values β0 to β10 are obtained.
[0046] The third step is to calculate the probability of user being assigned to a processing group. For each user, the standardized values of X are extracted from the independent variable dataset and substituted into the comprehensive score formula calculated using the final parameters. That is, the comprehensive score is equal to β0 plus β1 multiplied by standardized X1 plus β2 multiplied by standardized X2, and so on, up to β10 multiplied by standardized X10, to obtain the user's comprehensive score. Then, the comprehensive score is substituted into the logistic function, and the calculation method is the same as the logistic function in the second step. The resulting predicted probability is the probability that the user is assigned to a processing group.
[0047] The fourth step is to verify the fit of the logistic regression model. Calculate the AIC value of the logistic regression model by multiplying 2 by the total number of model parameters (including the intercept term and all regression coefficients) and subtracting 2 multiplied by the natural logarithm of the likelihood function value. Calculate the BIC value of the logistic regression model by adding the AIC value to the natural logarithm of the total number of parameters multiplied by the sample size and then subtracting the sample size, where the sample size is the total number of users. Compare the AIC and BIC values of the current logistic regression model with those of the simplified model after removing any independent variable. If both values of the current logistic regression model are lower, the current logistic regression model is considered superior. Plot the ROC curve, with the horizontal axis representing the false positive rate (the number of users predicted to be in the treatment group but actually in the control group divided by the total number of users in the actual control group), and the vertical axis representing the true positive rate (the number of users predicted to be in the treatment group and actually in the treatment group divided by the total number of users in the actual control group). (Total number of users in the treatment group); calculate the area under the ROC curve, i.e., the AUC value. If the AUC value is greater than or equal to 0.7, it indicates that the logistic regression model has the ability to distinguish between the treatment group and the control group. Calculate the variance inflation factor (VIF) of each independent variable. The calculation logic is as follows: first, calculate the square of the multiple correlation coefficient between the independent variable and all other independent variables, i.e., the determination coefficient of the regression model when the independent variable is the dependent variable and the other independent variables are independent variables; then subtract this square value from 1 to get a difference; finally, divide 1 by this difference to get the VIF value of the independent variable. If the VIF values of all independent variables are less than 5, it indicates that there is no serious collinearity among the independent variables, and the parameters of the logistic regression model are reliable. After all validations are passed, arrange the treatment group allocation probabilities of all users in ascending order of user ID to form a user treatment group allocation probability distribution table, which contains two columns of data: user ID and corresponding allocation probability.
[0048] Step 332: Based on the user processing group allocation probability distribution, the nearest neighbor matching algorithm is used to find users in the control group whose probabilities are closest to the processing group users for pairing, generating a balanced matching user dataset. Specifically, this includes: first, dividing the processing group and control group. For example, regarding the processing variable of whether or not a 200-minimum-50 coupon was used, the processing group consists of users who have used the coupon in the past 30 days, and the control group consists of users who have not used the coupon during the same period. For each user in the processing group, their processing group allocation probability is extracted. Then, all users in the control group are traversed in ascending order by user ID. The absolute difference between the processing group allocation probability of a control group user and the probability of that user in the processing group is calculated. The control group user with the smallest difference is recorded as a candidate matching object. If the smallest difference is less than 0.05, then... Once a user in the control group is identified as a match, if the minimum difference is greater than or equal to 0.05, the traversal range is expanded and the calculation is repeated until a user with a difference less than 0.05 is found. Each user in the control group is matched only once. If a user is already matched, they will not participate in other pairings. After all users in the processing group have been matched, all feature data of the users in the processing group and the users in the matching control group are integrated to form an initial matching user dataset. The balance of the dataset is checked. For each matching covariate, the absolute difference between the mean of the processing group and the mean of the control group is calculated, and then divided by the average of the two group means to obtain the relative difference. If the relative difference of all covariates is less than 10%, the dataset is considered balanced. Otherwise, pairings with excessive differences are removed and rematched until the balance condition is met, and finally a balanced matching user dataset is generated.
[0049] Step 333: Based on the matched user dataset, calculate the overall average treatment effect of each treatment variable on consumption performance by comparing the differences between the processing group and the control group in terms of consumption amount and frequency indicators. Specifically, this includes: for the consumption amount indicator in the matched user dataset, calculating the total amount of valid orders for all users in the processing group within 30 days after matching, dividing the total amount by the number of users in the processing group to obtain the average consumption amount of the processing group; similarly, calculating the total amount of valid orders for all users in the control group during the same period, dividing by the number of users in the control group to obtain the average consumption amount of the control group; subtracting the average consumption amount of the control group from the average consumption amount of the processing group to obtain the overall average treatment effect of the treatment variable on consumption amount. For the consumption frequency indicator, calculating the total number of valid orders for users in the processing group during the same period, dividing by the number of users to obtain the average consumption frequency of the processing group; calculating the total number of valid orders for users in the control group during the same period, dividing by the number of users to obtain the average consumption frequency of the control group; subtracting the average consumption frequency of the control group from the average consumption frequency of the processing group to obtain the overall average treatment effect of the treatment variable on consumption frequency. Repeat the above calculations for each treatment variable to obtain the corresponding overall average treatment effect.
[0050] Step 334: Based on the overall average treatment effect, analyze the differences in the characteristic distribution of individual users in the matched user dataset, and calculate the deviation between the individual treatment effect and the overall average treatment effect. Specifically, in the matched user dataset, each user in the treatment group is paired with a user in the control group. For each pairing, calculate the individual treatment effect: the total effective consumption amount of the treatment group users within 30 days after matching is subtracted from the total effective consumption amount of the control group users during the same period to obtain the individual consumption amount treatment effect; the number of effective orders of the treatment group users during the same period is subtracted from the number of effective orders of the control group users during the same period to obtain the individual consumption frequency treatment effect. The individual consumption amount effect bias is obtained by comparing the individual consumption amount effect with the corresponding overall average treatment effect. Similarly, the individual consumption frequency effect bias is obtained by subtracting the overall average consumption frequency effect from the individual consumption amount effect. For example, if the treatment group's consumption amount is 60 yuan more than the control group in a pair (individual effect), and the overall average effect of the treatment variable is 40 yuan, then the individual consumption amount effect bias is 20 yuan. If the individual effect is 30 yuan, then the bias is -10 yuan. This calculation is performed for all pairs to obtain the individual effect bias for each user.
[0051] Step 335 integrates the overall average treatment effect and individual effect bias to generate a characteristic impact effect estimation result that includes both overall and individual effects. Specifically, this involves creating an effect analysis table for each treatment variable, containing three basic columns: treatment variable name, overall average spending amount treatment effect, and overall average spending frequency treatment effect. Then, an individual effect bias column is added to the table, associating the ID of each paired treatment group user with the corresponding individual spending amount effect bias and individual spending frequency effect bias, ensuring that each user's individual bias corresponds to the overall effect of their respective treatment variable. For example, in the table for the treatment variable "spend 200 yuan and get 50 yuan off," it records both the overall average increase in spending amount by 40 yuan and spending frequency by 0.3 times, and also records user A's individual bias as +20 yuan and +0.1 times, user B's individual bias as -10 yuan and -0.2 times, etc. Integrating the effect analysis tables for all treatment variables forms a complete characteristic impact effect estimation result.
[0052] By using nearest neighbor matching to generate a balanced dataset, it is possible to ensure that the distribution of matching covariates in the treatment group and the control group tends to be consistent, providing a reliable data foundation for accurately inferring causal relationships. The overall average treatment effect is calculated to clarify the actual causal impact of each influencing factor on the amount and frequency of user consumption, thus avoiding misjudging correlations as causal relationships.
[0053] In a preferred embodiment of the present invention, step 4 above, which involves correlating the feature impact effect estimation results with the preprocessed user feature data, performing heterogeneity effect analysis through user feature subgroup division, identifying differences in driving factors and effect strength among different feature subgroups in terms of consumption amount and frequency indicators, and forming heterogeneity effect analysis results, may include:
[0054] In this embodiment of the invention, step 440 involves associating and integrating the feature impact effect estimation results with the preprocessed user feature data according to the user identifier, forming an augmented analysis dataset that includes user feature data and corresponding impact effects. Specifically, this includes: determining the user identifier as a unique user ID, which serves as the reference field for association; extracting the overall average treatment effect and individual effect bias corresponding to each user ID from the feature impact effect estimation results; extracting the consumption characteristics, behavioral characteristics, and membership level characteristics corresponding to each user ID from the preprocessed user feature data; matching the two parts of data through the user ID, that is, for each user ID, integrating all effect data in the feature impact effect estimation results with all feature data in the preprocessed user feature data to form a complete record containing the user ID, all feature data, and all effect data; summarizing the complete records of all users to form an augmented analysis dataset, ensuring that the features and effects of each user in the dataset correspond one-to-one, without omissions or mismatches.
[0055] Step 441: Based on the user demographic and consumption behavior characteristics in the augmented analysis dataset, users are divided into multiple feature subgroups according to preset rules. Specifically, this includes: extracting user demographic and consumption behavior characteristics from the augmented analysis dataset; setting preset division rules; dividing users into high-consumption, medium-consumption, and low-consumption subgroups based on average consumption amount; dividing them into high-frequency, medium-frequency, and low-frequency subgroups based on consumption frequency; and dividing them into diamond, gold, silver, and ordinary member subgroups based on membership level. Simultaneously, cross-subgroup divisions are performed, such as high-consumption-high-frequency subgroups and low-consumption-low-frequency subgroups. For each user in the augmented analysis dataset, the corresponding subgroup division rules are matched according to their feature data, and the user is assigned to the corresponding feature subgroup, ensuring that each user belongs to only one basic subgroup and one cross-subgroup.
[0056] Step 442: For each identified feature subgroup, extract the feature impact effect data for that subgroup and analyze the core factors driving consumption amount and frequency to identify the core driving factors affecting consumption amount and frequency within each subgroup. Specifically, this includes: for each feature subgroup, selecting feature impact effect data for all users in that subgroup from the augmented analysis dataset, including the overall average treatment effect and individual effect bias of each treatment variable on consumption amount and frequency; calculating the average effect value of each treatment variable within that subgroup; and for consumption amount, summing the individual consumption amount treatment effects of all users in the subgroup under that treatment variable. Divide the average effect of the treatment variable on the amount of consumption within the subgroup by the number of users in the subgroup. For the frequency of consumption, calculate the average effect of the treatment variable on the frequency of consumption within the subgroup in the same way. Sort the average effect values of all treatment variables within the subgroup from largest to smallest absolute value, and select the top two treatment variables as the core driving factors of the subgroup. For example, in the high-consumption subgroup, the average amount of consumption effect of participating in the new product tasting activity is 80 yuan (the largest absolute value), and the average amount of consumption effect of using the diamond member exclusive coupon is 60 yuan (the second largest absolute value). These two treatment variables are the core driving factors of the amount of consumption in the high-consumption subgroup.
[0057] Step 443 involves cross-analysis of the core driving factors of different characteristic subgroups to quantify the differences in the effect strength of the same factor in different subgroups and identify the distribution pattern of factor effects as user characteristics change. Specifically, this includes: collecting the core driving factors of all characteristic subgroups and their average effect values within the subgroups, establishing a comparison table, with table fields including the name of the treatment variable, the name of the subgroup, the average consumption amount effect within the subgroup, and the average consumption frequency effect within the subgroup. For the same treatment variable, the average effect value is extracted in different subgroups, the difference in effect strength is calculated, and the difference in consumption amount dimension is obtained by subtracting the average effect value of the low-consumption subgroup from the average effect value of the high-consumption subgroup. The difference in consumption frequency is obtained by subtracting the average effect value of the low-frequency subgroup from the average effect value of the high-frequency subgroup. For example, the average consumption amount effect of using membership coupons is 50 yuan in the high-consumption subgroup and 30 yuan in the low-consumption subgroup, with a difference of 20 yuan; the average consumption frequency effect is 0.4 times in the high-frequency subgroup and 0.8 times in the low-frequency subgroup, with a difference of -0.4 times. The relationship between these differences and subgroup characteristics is analyzed. For example, as the average consumption amount of users increases, the consumption amount effect of using membership coupons shows an upward trend; as the consumption frequency increases, the consumption frequency effect of this treatment variable shows a downward trend, thereby identifying the distribution law of factor effects as user characteristics change.
[0058] Step 444: Based on the analysis of differences in driving factors and the analysis of differences in effect strength, generate the results of the heterogeneity effect analysis describing each characteristic subgroup. Specifically, this includes: for each characteristic subgroup, summarizing its core driving factors, such as participation in new product tasting and exclusive coupons for diamond members for the high-consumption subgroup, the average effect value of each core driving factor, and combining the comparison results from step 443 to describe the differences between this subgroup and other subgroups in terms of driving factors, such as the core driving factor for the high-consumption subgroup being new product-related activities, while for the low-consumption subgroup it is basic discount activities, as well as the differences in the effect strength of the same factors, such as the effect of the same discount activity being 1.5 times that of the high-consumption subgroup. This information is then categorized and organized by subgroup to form a structured heterogeneity effect analysis result, including the subgroup name, a list of core driving factors and their effect values, differences in driving factors with other subgroups, differences in effect strength with other subgroups, and distribution patterns, ensuring that the results clearly reflect the heterogeneity of different characteristic subgroups in terms of consumption drivers.
[0059] By associating feature data with effect data, the integrity of the analysis data can be ensured, and subgrouping and effect analysis can be based on a unified user dimension, avoiding analytical bias caused by data fragmentation. Subgrouping features according to preset rules can accurately capture user groups with similar features, providing a basis for exploring behavioral differences among different groups and avoiding the analysis of all users as homogeneous groups.
[0060] In a preferred embodiment of the present invention, step 5 above, which involves fine-grained segmentation and labeling of user features based on the heterogeneity effect analysis results, and dynamic updating of user profiles to achieve structural consistency and classification accuracy of user data, may include:
[0061] In this embodiment of the invention, step 550 involves generating fine-grained user classification results by combining and analyzing multiple feature dimensions based on the differences in driving factors and effect strength among different feature subgroups in the heterogeneity effect analysis results. Specifically, this includes extracting core feature dimensions from the heterogeneity effect analysis results, including spending power dimension, consumption frequency dimension, membership level dimension, and driving factor sensitivity dimension. Each feature dimension is further subdivided and defined. For example, in the spending power dimension, high spending is defined as an average spending amount in the past 90 days ≥ 1.5 times the average of all users; medium spending is defined as 0.5 to 1.5 times the average; and low spending is defined as ≤ 0.5 times the average. In the driving factor sensitivity dimension, sensitivity to new product activities refers to an average spending amount effect ≥ 6 times the average of participating in new product tasting activities. The initial cost was 0 yuan. Then, a multi-dimensional combination analysis was conducted. First, spending power and spending frequency were combined to generate nine basic combinations, such as high-spending-high-frequency and high-spending-medium-frequency. Next, each basic combination was combined with membership level to generate 36 secondary combinations, such as high-spending-high-frequency-diamond and high-spending-high-frequency-gold. Finally, each secondary combination was combined with a sensitive dimension of driving factors to generate 144 candidate categories, such as high-spending-high-frequency-diamond-sensitive to new product promotions and high-spending-high-frequency-gold-sensitive to discount coupons. Categories with actual users were selected from the candidate categories, and the effect patterns of each subgroup in the heterogeneity effect analysis results were combined to merge categories with similar characteristics, ultimately generating 60-80 fine-grained user classification results. Each category contains a unique combination feature identifier.
[0062] Step 551: Using the fine-grained user classification results as the basis for annotation, add feature labels to each user category to form a complete annotation system. Specifically, this includes: defining feature labels for each category based on the combined features of the fine-grained user classification. Label types include consumption potential labels, price sensitivity labels, and category preference labels. Consumption potential labels are determined based on consumption capacity and effect strength. For example, in the high-consumption-high-frequency-diamond category, if its average consumption amount effect ranks in the top 20% of all categories, it is labeled as high potential; the medium-consumption-medium-frequency-gold category is labeled as medium potential; and the low-consumption-low-frequency-ordinary category is labeled as low potential. Price sensitivity labels are determined based on the effect strength on discount-type driving factors. For example, for a certain category... If the average spending amount using discount coupons is ≥50 yuan, it is labeled as highly sensitive; if the effect is between 30-49 yuan, it is labeled as moderately sensitive; and if it is <30 yuan, it is labeled as low sensitive. Category preference labels are determined based on the sensitivity to activities related to a specific cake category. For example, if the average consumption frequency effect of participating in birthday cake special discounts is the highest in a certain category, it is labeled as birthday cake preference; if the effect of participating in new mousse cake tastings is the highest, it is labeled as mousse cake preference. Each fine-grained category must include at least one consumption potential label, one price sensitivity label, and one category preference label. All labels for all categories are summarized to form a labeling system table that includes label name, label definition, and corresponding category range, ensuring that the labels correspond one-to-one with the category characteristics.
[0063] Step 552 involves integrating the completed user classification results into the existing user profiles according to a unified standard to form structured profile data. Specifically, this includes: determining the basic fields of the existing user profiles, and adding four new fields: fine-grained classification, consumption potential tags, price-sensitive tags, and category preference tags. The fine-grained classification field stores the unique identifier of the category; the consumption potential tag, price-sensitive tag, and category preference tag fields store the text values of the corresponding tags, with a unified field format: user ID is an 18-digit numeric code, category identifier is a letter combination string, and tags are Chinese text; all fields use UTF-8 encoding, numeric fields retain two decimal places (if any), and text fields remove special characters. The fine-grained classification results and corresponding tags for each user are filled into the newly added fields according to the above format, and integrated with the existing profile fields to form complete user profile data containing basic information, consumption characteristics, classification results, and feature tags. After integration, the data is format-validated to ensure that the newly added fields in each user's profile record have no empty values and that the format conforms to the standard, ultimately forming a structured profile dataset that can be directly used for querying and analysis in the e-commerce system.
[0064] Step 553: Based on the newly generated user consumption data, regularly update and maintain the user profile data, adjust user classification and feature annotations, and verify the accuracy of user classification and the consistency of data structure through preset consistency verification rules. Specifically, this includes: setting the update cycle to once a week, automatically collecting newly generated user consumption data from the previous week every Monday morning, cleaning and standardizing the new data, calculating each user's latest spending power, spending frequency, membership level, and effect value on each driving factor, comparing the newly calculated feature values with the feature standards of the user's current fine-grained classification, keeping the classification and labels unchanged if the user features still meet the current classification standards, and re-matching the fine-grained classification and updating the corresponding feature labels after the update is completed, performing consistency verification to check the classification accuracy, checking whether the latest user features are consistent with the new classification standards, marking inconsistencies as abnormal and readjusting them; verifying data structure consistency, checking whether the number, format, and encoding of fields in all user profiles are uniform, correcting formats for inconsistencies, recording abnormal data and structural problems, generating a weekly update report, ensuring that the adjusted user profile data is accurately classified and structurally consistent, and maintaining the long-term reliability of user data.
[0065] Adding feature labels to each category and forming a labeling system can transform abstract classification results into intuitive feature descriptions, making the differences in user characteristics between different categories identifiable and understandable. Integrating these features into user profiles according to unified standards can ensure that newly added category and label data are compatible with the existing profile structure, avoid application obstacles caused by data format chaos, and achieve structural consistency of user data.
[0066] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0067] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0068] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for analyzing and processing user consumption data in a cake e-commerce system, characterized in that, The method includes: Step 1: Collect and integrate multi-source consumption data of users in the cake mall system, including order data, payment data, membership data and behavior log data, to build a unified user consumption dataset; Step 2 involves cleaning and standardizing the user consumption dataset to form an initial user feature set. Based on this initial user feature set, different user behavior response patterns are identified through user group segmentation, and data quality is corrected by combining the dynamic path of consumption behavior over time, resulting in preprocessed user feature data of consistent quality, including: Calculate the feature distance of each user in the initial user feature set across multiple dimensions, including consumption frequency, average consumption amount, recent consumption time interval, membership level, and behavioral indicators. Based on a preset distance threshold, group users with similar features into the same group to generate user group segmentation results for data quality verification. For each user group, extract the consumption characteristic change sequence of the corresponding group within a continuous time window to generate dynamic path features that reflect the evolution of consumption behavior of each group; The consumption characteristic change sequence of each user is compared with the dynamic path characteristics point by point to identify abnormal data points that deviate from the expected evolution pattern. Through data smoothing and trend calibration, preprocessed user characteristic data of consistent quality is generated. Step 3: Based on the preprocessed user characteristic data, a causal inference model is constructed using the propensity score matching method. The potential influencing factors of the core indicators of user consumption amount and frequency are analyzed, the overall average treatment effect of each factor on consumption performance is estimated, and the effect bias caused by individual differences in characteristics is calculated to obtain the characteristic influence effect estimation results including the overall effect and the individual effect. Step 4: Correlate the feature impact effect estimation results with the preprocessed user feature data, and conduct heterogeneity effect analysis by dividing the user feature subgroups. Identify the differences in driving factors and effect strengths among different feature subgroups in terms of consumption amount and frequency indicators, and form the heterogeneity effect analysis results, including: The feature impact effect estimation results are linked and integrated with the preprocessed user feature data according to user identifiers to form an enhanced analysis dataset that includes user feature data and corresponding impact effects; Based on the user demographic and consumption behavior characteristics in the augmented analytics dataset, users are divided into multiple feature subgroups according to preset rules. These multiple feature subgroups include high-spending and low-spending groups divided by consumption amount, high-frequency and low-frequency groups divided by consumption frequency, and different user groups divided by membership level. For each of the identified feature subgroups, feature impact data for the corresponding subgroups are extracted, and the core factors driving consumption amount and frequency are analyzed to identify the core driving factors affecting consumption amount and frequency within each subgroup. Cross-analysis of the core driving factors of different feature subgroups was conducted to quantify the differences in the effect intensity of the same factor in different subgroups and to identify the distribution pattern of factor effect as user characteristics change. Based on the analysis of differences in driving factors and the analysis of differences in effect intensity, the results of the analysis of heterogeneity effects describing each characteristic subgroup are generated. Step 5: Based on the heterogeneity effect analysis results, perform fine-grained segmentation and labeling of user characteristics, and dynamically update user profiles to achieve structural consistency and classification accuracy of user data.
2. The method for analyzing and processing user consumption data in a cake e-commerce system according to claim 1, characterized in that, The user consumption dataset is cleaned and standardized to form an initial set of user features, including: The user consumption dataset is cleaned by performing missing value imputation, outlier correction, and duplicate record deletion operations. Based on the cleaned user consumption data, the consumption frequency, average consumption amount, and recent consumption time interval are calculated, and a user feature set is constructed by combining membership level information and platform behavior records. The numerical features in the user feature set are standardized and normalized respectively, and the categorical features are one-hot encoded to generate the initial user feature set.
3. The method for analyzing and processing user consumption data in a cake e-commerce system according to claim 2, characterized in that, Based on preprocessed user characteristic data, a causal inference model is constructed using propensity score matching. This model analyzes the potential influencing factors of core indicators such as user spending amount and frequency, estimates the overall average treatment effect of each factor on spending performance, and calculates the effect bias at the individual level due to characteristic differences. The resulting estimation results of the characteristic impact effect, including both overall and individual effects, are as follows: Based on preprocessed user feature data, potential influencing factors to be analyzed are identified as processing variables, and the remaining features are used as matching covariates. By using matched covariates, the probability of each user being assigned to a processing group is estimated through a logistic regression model, thus obtaining the probability distribution of user processing group assignment. Based on the probability distribution of user processing groups, the nearest neighbor matching algorithm is used to find the user in the control group whose probability is closest to the user in the processing group and pair them up to generate a balanced matching user dataset. Based on a matched user dataset, the overall average treatment effect of each treatment variable on consumption performance is calculated by comparing the differences between the treatment group and the control group in terms of consumption amount and frequency indicators. Based on the overall average treatment effect, the differences in the feature distribution of individual users in the matched user dataset are analyzed, and the deviation between the individual treatment effect and the overall average treatment effect is calculated. By integrating the overall average treatment effect with the individual effect bias, a characteristic impact effect estimation result that includes both overall and individual level effects is generated.
4. The user consumption data analysis and processing method for a cake e-commerce system according to claim 3, characterized in that, Based on the heterogeneity effect analysis results, user characteristics are finely segmented and labeled, and user profiles are dynamically updated to achieve structural consistency and classification accuracy of user data, including: Based on the differences in driving factors and effect intensity among different feature subgroups in the heterogeneity effect analysis results, a fine-grained user classification result is generated by combining multiple feature dimensions. Using fine-grained user classification results as the basis for annotation, feature labels are added to each user category to form a complete annotation system; The user classification results that have been annotated will be integrated into the existing user profiles according to a unified standard to form structured profile data; Based on newly generated user consumption data, the user profile data is updated and maintained regularly, user classification and feature labeling are adjusted, and the accuracy of user classification and consistency of data structure are verified through preset consistency verification rules.
5. The user consumption data analysis and processing method for a cake e-commerce system according to claim 4, characterized in that, The feature tags include consumption potential level, price sensitivity, and category preference.
6. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Shopping mall user portrait construction method
CN117575648A
Commercial truck portrait construction and label distribution method under multi-source heterogeneous data fusion
CN118821027A