User Total Active Days Prediction Method and System Based on Distributed Supervised Learning
Through a step-by-step supervised learning method, using recent data for iterative prediction, the problem of traditional TAD prediction methods relying on old data is solved, and more accurate prediction of user active days is achieved.
Patent Information
- Application Number
- CN202510199177.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The traditional supervised total active days (TAD) prediction method relies on historical data over a long period of time and cannot accurately reflect the current user behavior and product status, resulting in bias in prediction results.
Using a step-by-step supervised learning method, through distributed prediction, we maximize the use of recent data as samples for prediction. Specific steps include: grouping channel data by month, extracting multidimensional feature sets; using supervised machine learning models for initial training and optimization; gradually expanding training samples through iterative prediction mechanisms, predicting TAD values for shortening cycles in stages until they cover the entire stage; and finally building a prediction model based on the complete data set.
It significantly improves the accuracy of forecasting of users' total active days, can more accurately reflect the current status of the product and user active behavior patterns, and adapt to product iteration and market changes.
Smart Images

Figure CN119693049B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular, to a method and system for predicting the total active days of users based on distributed supervised learning. Background Art
[0002] In the current Internet and mobile Internet fields, user behavior analysis is crucial for product optimization, market promotion, and user experience improvement. The total active days of users (TAD, Total Active Days per User) is one of the important indicators to measure the user stickiness and activity of a product. It refers to the total active days of each user, that is, within a statistical period, the average number of days each user uses a certain application. If the statistical period is more than one year, the total active days of each user can basically reflect the number of days the user uses the application before churn. For the operation team and product developers, accurately predicting the total active days of users helps to formulate more effective operation strategies, optimize product functions, and improve user experience.
[0003] However, traditional supervised TAD prediction methods face many challenges in practical applications. Traditional methods usually need to use historical data of a relatively long time period, such as data 360 days ago, for model training and prediction. However, this method has several significant problems. First, over time, the product itself may have gone through multiple rounds of iteration, and there have been significant changes in functions and interfaces, which has led to changes in user behavior patterns and activity levels. Therefore, using historical data from an earlier time as a sample may not accurately reflect the true situation of current users, resulting in deviations in prediction results.
[0004] Secondly, with the changes in the market environment, the channel types and structures of products may also change significantly. These changes will also affect user acquisition and activity levels, and thus affect the prediction results of TAD.
[0005] Furthermore, the TAD situation of the overall user group may also be significantly different from that one year ago. This may be the result of the combined effects of various factors such as changes in the user group, evolution of market trends, and strategic adjustments of competitors. Therefore, using data from a relatively long time ago as a sample for prediction may not be able to adapt to these changes, resulting in inaccurate prediction results.
[0006] In view of the above problems, there is an urgent need for a more accurate and efficient TAD prediction method to fully utilize recent data that can better reflect the current situation of the product as a sample for prediction. This application proposes a method and system for predicting the total active days of users based on distributed supervised learning, aiming to solve the problems existing in traditional methods and improve the accuracy of prediction. Summary of the Invention
[0007] The object of the present invention is to provide a method and system for predicting the total active days of users based on distributed supervised learning, aiming to maximize the use of recent data that can better reflect the current situation of the product as samples for prediction through distributed prediction, so as to effectively improve the accuracy of prediction.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] In the first aspect, the present application provides a method for predicting the total active days of users based on distributed supervised learning, including the following steps:
[0010] S1. Group the channel data with the new addition time within a set time window before the target prediction date (for example, several months back from the target prediction date) by month, and extract a multi-dimensional feature set, including channel attributes, time attributes, TAD values at each stage, and TAD ratios at each stage;
[0011] S2. Use the longest cycle TAD as the target variable, and the remaining features as input features, and train through a supervised machine learning model to obtain an initial prediction model, and optimize the model parameters based on the prediction error rate index;
[0012] S3. Adopt an iterative prediction mechanism, gradually expand the training samples to a data set including historical prediction results, and predict the TAD values of shortened cycles in stages until the full-stage TAD of the target prediction cycle is covered; the specific steps include:
[0013] S3.1. At the first prediction, use the initial sample data within the set time window to predict the longest cycle TAD, and generate the first set of prediction results;
[0014] S3.2. In each subsequent iteration, incorporate the prediction results of the previous step into the training samples to expand the training data set;
[0015] S3.3. Based on the expanded data set, gradually shorten the target TAD cycle, and increase the multi-stage TAD prediction. The target cycle of each iteration is reduced by a set number of days, and at least one TAD prediction of the shortened cycle is added;
[0016] S3.4. Repeat steps S3.2 - S3.3 until the training data covers all shortened-stage TADs of the target prediction cycle;
[0017] S4. Perform final modeling based on the complete cycle data set, extract the full-stage TAD values and TAD ratios, and train to generate a final prediction model;
[0018] S5. Input the new channel data into the final prediction model, and output the TAD prediction results of the target cycle.
[0019] In a preferred embodiment, in step S1, the channel attributes include a unique ID for identifying different channels, the type to which the channel belongs, and the channel increment;
[0020] The time attribute represents the date of data recording and identifies whether the date is a weekend or a holiday;
[0021] The TAD values at each stage refer to the total active days within different time periods starting from the date of user addition, including but not limited to 30-day, 60-day, 90-day to 360-day TAD;
[0022] The TAD ratios at each stage refer to the ratios between the TAD values at different stages, including the ratios of adjacent cycle TAD and the combination of TAD ratios of the initial cycle and subsequent cycles (such as 30-day / 60-day, 30-day / 90-day to 30-day / 360-day, etc.), which are used to reflect the change trend of user activity.
[0023] In a preferred embodiment, in step S2, the supervised machine learning model is a LightGBM model or an XGBoost model.
[0024] In a preferred embodiment, in step S2, the prediction error rate index is the mean absolute percentage error MAPE.
[0025] In a preferred embodiment, in step S3.3, the target cycle is reduced by 30 days each time, and the number of predicted TAD for the shortened cycle newly added in each iteration increases by one successively.
[0026] In a preferred embodiment, in step S4, the full-stage TAD value includes the 360-day full-cycle active days starting from the date of user addition, and the TAD ratio includes the TAD growth ratio for consecutive dates.
[0027] In a preferred embodiment, in step S5, the newly added channel data needs to meet the data accumulation condition of at least 3 days before it can be input into the final prediction model for target cycle TAD prediction.
[0028] In a second aspect, the present application provides a user total active days prediction system based on distributed supervised learning for implementing the user total active days prediction system based on distributed supervised learning according to any one of the first aspects. The system includes the following modules:
[0029] A data preprocessing module for grouping the channel data with the new addition time within a set time window before the target prediction date by month and extracting a multi-dimensional feature set. The multi-dimensional feature set includes channel attributes, time attributes, TAD values at each stage, and TAD ratios at each stage;
[0030] An initial modeling module, which takes the longest cycle TAD as the target variable and the remaining features as input features, trains through a supervised machine learning model to generate an initial prediction model, and optimizes the model parameters based on the prediction error rate index;
[0031] An iterative prediction module, which is used to execute an iterative prediction mechanism, gradually expand the training samples to a data set including historical prediction results, and predict the TAD values of the shortened cycle in stages until the full-stage TAD of the target prediction cycle is covered; the iterative prediction module includes:
[0032] An initial prediction unit, which is used to predict the longest cycle TAD using the initial sample data within a set time window during the first prediction and generate the first set of prediction results;
[0033] A data expansion unit, which is used to incorporate the prediction results of the previous step into the training samples and expand the training data set during each subsequent iteration;
[0034] A cycle adjustment unit, which is used to gradually shorten the target TAD cycle based on the expanded data set and increase the multi-stage TAD prediction. The target cycle of each iteration is reduced by a set number of days, and at least one TAD prediction of the shortened cycle is added;
[0035] A loop control unit, which is used to repeatedly call the data expansion unit and the cycle adjustment unit until the training data covers all the shortened-stage TADs of the target prediction cycle;
[0036] A final modeling module, which is used to perform final modeling based on the complete cycle data set, extract the full-stage TAD values and TAD ratios, and train to generate a final prediction model;
[0037] A prediction output module, which inputs the new channel data into the final prediction model and outputs the target cycle TAD prediction result.
[0038] In a preferred embodiment, in the data preprocessing module:
[0039] The channel attributes include a unique ID for identifying different channels, the type to which the channel belongs, and the channel increment;
[0040] The time attributes include the date of data recording and the identification of whether it is a weekend or a holiday;
[0041] The TAD values of each stage include the total active days within different time periods starting from the date of user addition, including but not limited to 30-day, 60-day, 90-day to 360-day TAD;
[0042] The TAD ratios of each stage include the ratio of adjacent cycle TADs and the combination of the TAD ratios of the initial cycle and subsequent cycles, which are used to reflect the change trend of user activity.
[0043] In a preferred embodiment, in the initial modeling module, the supervised machine learning model is a LightGBM model or an XGBoost model.
[0044] In a preferred embodiment, in the initial modeling module, the prediction error rate index is the mean absolute percentage error MAPE.
[0045] In a preferred embodiment, in the cycle adjustment unit, the target cycle is reduced by 30 days each time, and the number of predicted shortened cycle TADs newly added in each iteration increases by one successively.
[0046] In a preferred embodiment, in the final modeling module, the full-stage TAD value includes the active days of the 360-day full cycle starting from the date when the user is newly added, and the TAD ratio includes the TAD growth ratio of consecutive dates.
[0047] In a preferred embodiment, in the prediction output module, the newly added channel data needs to meet the data accumulation condition of at least 3 days before it can be input into the final prediction model for target cycle TAD prediction.
[0048] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0049] The present invention provides a method and system for predicting the total active days of users based on distributed supervised learning. The method includes: processing channel data within a set time window by monthly grouping to extract a multi-dimensional feature set; using a supervised machine learning model for initial training and optimization; gradually expanding training samples through an iterative prediction mechanism, and predicting the TAD values of the shortened cycle in stages until the full stage is covered; finally, constructing a prediction model based on the complete data set to achieve accurate prediction of the newly added channel data. Through the distributed prediction method, the present invention maximally utilizes recent data as samples, more accurately reflects the current state of the product and the user active behavior pattern, thereby significantly improving the prediction accuracy. In addition, the distributed prediction strategy enables the model to gradually adapt to the changing trend of the data, and still maintain high prediction performance when the product iterates frequently and the channel structure changes significantly. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1It is a schematic flow chart of the method for predicting the total active days of users based on distributed supervised learning of the present invention;
[0052] Figure 2 It is a data table for stage prediction of user active days (TAD) shown in the preferred embodiment of the present invention. Detailed implementation manners
[0053] In order to make the above and other features and advantages of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art, and are only exemplary, not restrictive.
[0054] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0055] Embodiment 1:
[0056] Referring to FIG. 1, this embodiment provides a method for predicting the total active days of users based on distributed supervised learning, which specifically includes the following steps:
[0057] Step S1: Group the channel data with the new addition time within a set time window before the target prediction date (for example, several months back from the target prediction date) by month, and extract a multi-dimensional feature set, including channel attributes, time attributes, TAD values at each stage, and TAD ratios at each stage.
[0058] Among them, the channel attributes include a unique ID for identifying different channels, the type to which the channel belongs, and the channel new addition quantity;
[0059] The time attribute represents the date of data recording and identifies whether the date is a weekend or a holiday;
[0060] The TAD values at each stage refer to the total active days within different time periods starting from the date of user new addition, including but not limited to 30-day, 60-day, 90-day to 360-day TAD;
[0061] The TAD ratios at each stage refer to the ratios between TAD values at different stages, including the ratio of adjacent cycle TAD and the combination of TAD ratios of the initial cycle and subsequent cycles (such as 30-day / 60-day, 30-day / 90-day to 30-day / 360-day, etc.), which are used to reflect the change trend of user activity.
[0062] Step S2: Use the longest cycle TAD (such as 360-day TAD) as the target variable, and the remaining features as input features. Train through a supervised machine learning model (such as LightGBM model or XGBoost model) to obtain an initial prediction model, and optimize the model parameters based on a prediction error rate metric (such as Mean Absolute Percentage Error MAPE).
[0063] Step S3: Adopt an iterative prediction mechanism to gradually expand the training samples to a dataset including historical prediction results, and predict the TAD values of shortened cycles in stages until covering the full-stage TAD of the target prediction cycle. The specific steps include:
[0064] Step S3.1: At the first prediction, use the initial sample data within the set time window to predict the longest cycle TAD and generate the first set of prediction results;
[0065] Step S3.2: In each subsequent iteration, incorporate the prediction results of the previous step into the training samples to expand the training dataset;
[0066] Step S3.3: Based on the expanded dataset, gradually shorten the target TAD cycle (such as 330 days, 300 days), and increase multi-stage TAD predictions. The target cycle of each iteration decreases by a set number of days (such as 30 days), and at least one TAD prediction of a shortened cycle is added.
[0067] Step S3.4: Repeat steps S3.2 - S3.3 until the training data covers all shortened-stage TADs of the target prediction cycle.
[0068] Step S4: Conduct final modeling based on the complete cycle dataset, extract the full-stage TAD values and TAD ratios, and train to generate the final prediction model. The full-stage TAD values include the active days of the 360-day full cycle from the date when the user is newly added, and the TAD ratios include the TAD growth ratios of consecutive dates.
[0069] Step S5: Input the new channel data into the final prediction model to output the TAD prediction results of the target cycle. Among them, the new channel data needs to meet the data accumulation condition of at least 3 days before it can be input into the final prediction model for TAD prediction of the target cycle.
[0070] To more clearly illustrate the technical solution of this application, a specific example is described below.
[0071] Assume that the current date is January 1, 2025, and it is necessary to predict the total active days (TAD) of users in each stage of the whole year of 2024. In this embodiment, a distributed supervised learning method is used to iteratively model using recent data to improve the prediction accuracy of 360-day TAD.
[0072] Step 1: Sample data preparation.
[0073] Time window setting: Select the newly added data of the channels from October 2023 to the present as samples and group them monthly (a total of 15 months). For specific reference, Figure 2 the data table shown.
[0074] Step 2: Predict the 360-day TAD for the data grouping in January 2024.
[0075] 2.1) Use the data from October 2023 to December 2023 as samples to process and calculate the following indicators: channel ID, channel type, channel new addition, whether the current date is a weekend, whether the current date is a holiday, TAD values at each stage (30-day, 60-day, 90-day... 360-day), and TAD ratios at each stage (such as 30-day / 60-day, 30-day / 90-day... 30-day / 330-day).
[0076] 2.2) After obtaining the above indicators, use the 360-day TAD (the longest cycle) as the target variable and the remaining indicators as feature variables, and call supervised machine learning models such as LightGBM (abbreviation: lgb) and XGBoost (abbreviation: xgb) for training. By comparing the prediction error rates (MAPE) of different models, select the model and parameters with the smallest error rate as the optimal model.
[0077] 2.3) Apply the optimal model for prediction: Apply the optimal model and parameters obtained in step 2.2), input the features of the model under the samples in January 2024, and predict and output the 360-day TAD in January 2024.
[0078] Step 3: Predict the 330-day TAD and 360-day TAD for the data grouping in February 2024.
[0079] 3.1) Use the results predicted in step 2, and use the data from October 2023 to January 2024 as samples to process and calculate the following indicators: channel ID, channel type, channel new addition, whether the current date is a weekend, whether the current date is a holiday, TAD values at each stage (30-day, 60-day, 90-day... 360-day), and TAD ratios at each stage (such as 30-day / 60-day, 30-day / 90-day... 30-day / 330-day).
[0080] 3.2) After obtaining the above indicators, use the 330-day TAD and 360-day TAD as the target variables and the remaining indicators as feature variables, call supervised machine learning models such as lgb and xgb, and select the model and parameters with the smallest prediction error rate (MAPE).
[0081] 3.3) Apply the model and parameters obtained in step 3.2), input the features of the model under the February 2024 sample, and predict and output the 330-day TAD and 360-day TAD for February 2024.
[0082] Step 4: Predict the 300-day TAD, 330-day TAD, and 360-day TAD for the data grouping in March 2024.
[0083] 4.1) Using the results predicted in Step 3, with the data from October 2023 to February 2024 as the sample, process and calculate the following indicators: channel ID, channel type, channel increment, whether the current date is a weekend, whether the current date is a holiday, TAD values for each stage (30-day, 60-day, 90-day... 360-day), and TAD ratios for each stage (such as 30-day / 60-day, 30-day / 90-day... 30-day / 270-day).
[0084] 4.2) After obtaining the above indicators, using the 300-day TAD, 330-day TAD, and 360-day TAD as the target variables and the other indicators as the feature variables, call supervised machine learning models such as lgb and xgb, and select the model and parameters with the smallest prediction error rate (MAPE).
[0085] 4.3) Apply the model and parameters obtained in step 4.2), input the features of the model under the March 2024 sample, and predict and output the 300-day TAD, 330-day TAD, and 360-day TAD for March 2024.
[0086] Step 5: And so on, repeatedly build models and predictions to supplement the TAD for each stage under the full-year data of 2024 (refer to Figure 2 the yellow part of the data table in
[0087] Through the method of rolling prediction, gradually supplement the TAD prediction values for each month in 2024 at different stages (such as 30-day, 60-day... 360-day).
[0088] Step 6: After supplementation, perform the final model building.
[0089] 6.1) Using the data from October 2023 to December 2024 as the sample (including the predicted TAD values), process and calculate the following more comprehensive indicators: channel ID, channel type, channel increment, whether the current date is a weekend, whether the current date is a holiday, TAD values for each stage (1-day, 2-day, 3-day... 360-day), and TAD ratios for each stage (such as 2-day / 1-day, 3-day / 1-day, etc.).
[0090] 6.2) After obtaining the above indicators, using the 360-day TAD as the target variable and the remaining indicators as feature variables, call supervised machine learning models such as lgb and xgb, select the model and parameters with the smallest prediction error rate (MAPE), and train to generate the final prediction model.
[0091] Step Seven: Prediction of new channels.
[0092] Apply the final prediction model and parameters obtained in Step Six. When the subsequent new channel reaches 3 days, input the model features (including channel ID, channel type, channel new volume, whether the current date is a weekend, whether the current date is a holiday, and the calculated TAD value, etc.) to predict the 360-day TAD of this channel (refer to Figure 2 the blue part of the data table in
[0093] In summary, through the iterative mechanism of this application, recent data and prediction results are effectively integrated. Each iteration incorporates the latest prediction results into the training samples, solving the problem of historical data lag, and enabling the model to be optimized based on the current product release status at all times (for example, the prediction results in January 2024 are used in the model construction in February). By gradually shortening the target prediction period (reducing 30 days each time) and increasing the prediction stages (such as predicting 360 days for the first time, and predicting 330 days and 360 days for the second time), the model can synchronously capture short-term and long-term user activity trends.
[0094] Compared with traditional methods, the technical solution of this embodiment significantly reduces the dependence on obsolete data. Traditional methods usually rely on historical data 360 days ago, while the technical solution of this embodiment only requires 3 months of initial data (from October to December 2023) to start prediction, and subsequent samples are expanded through iteration, reducing the dependence on obsolete data. In the embodiment, the prediction in January 2024 only uses 3 months of data, and gradually expands to the whole-year data later, with more reasonable resource occupancy. For new channels, only 3 days of data need to be accumulated to predict the user active days (TAD) for 360 days through the finally constructed model, thus meeting the real-time requirements of the business.
[0095] Embodiment 2:
[0096] Based on the same design concept, this embodiment also provides a user total active days prediction system based on distributed supervised learning for performing the user total active days prediction method based on distributed supervised learning described in Embodiment 1.
[0097] Specifically, the system includes: a data preprocessing module, an initial modeling module, an iterative prediction module, a final modeling module, and a prediction output module.
[0098] The data preprocessing module is used to group the channel data with the new addition time within a set time window before the target prediction date by month, and extract a multi-dimensional feature set. The multi-dimensional feature set includes channel attributes, time attributes, TAD values at each stage, and TAD ratios at each stage.
[0099] Among them, the channel attributes include a unique ID for identifying different channels, the type to which the channel belongs, and the channel new addition quantity; the time attributes include the date of data record and an identifier indicating whether it is a weekend or a holiday; the TAD values at each stage include the total active days within different time periods starting from the day of user new addition, including but not limited to 30-day, 60-day, 90-day to 360-day TAD; the TAD ratios at each stage include the ratio of adjacent cycle TAD and the combination of the TAD ratio of the initial cycle and subsequent cycles, which is used to reflect the change trend of user activity.
[0100] The initial modeling module is used to take the longest cycle TAD as the target variable, and the remaining features as input features, and train through a supervised machine learning model to generate an initial prediction model, and optimize the model parameters based on the prediction error rate index. In a preferred embodiment, the supervised machine learning model is a LightGBM model or an XGBoost model; the prediction error rate index is the mean absolute percentage error MAPE.
[0101] The iterative prediction module is used to execute an iterative prediction mechanism, gradually expand the training samples to a data set including historical prediction results, and predict the TAD values of shortened cycles in stages until the full-stage TAD of the target prediction cycle is covered.
[0102] Specifically, the iterative prediction module includes:
[0103] An initial prediction unit, which is used for the first prediction, to predict the longest cycle TAD using the initial sample data within the set time window and generate the first set of prediction results;
[0104] A data expansion unit, which is used in each subsequent iteration to incorporate the prediction results of the previous step into the training samples and expand the training data set;
[0105] A cycle adjustment unit, which is used based on the expanded data set to gradually shorten the target TAD cycle and increase the multi-stage TAD prediction. The target cycle of each iteration is reduced by a set number of days (such as 30 days), and at least one TAD prediction of a shortened cycle is added;
[0106] A loop control unit, which is used to repeatedly call the data expansion unit and the cycle adjustment unit until the training data covers all the shortened-stage TADs of the target prediction cycle.
[0107] The final modeling module is used to perform final modeling based on the complete cycle dataset, extract the full-stage TAD value and TAD ratio, and train and generate a final prediction model. Among them, the full-stage TAD value includes the total active days in the 360-day full cycle starting from the date when the user is newly added, and the TAD ratio includes the TAD growth ratio for consecutive dates.
[0108] The prediction output module is used to input the newly added channel data into the final prediction model and output the target cycle TAD prediction result. Among them, the newly added channel data can be input into the final prediction model for target cycle TAD prediction only after meeting the data accumulation condition of at least 3 days.
[0109] It can be understood that the various modules described in the user total active days prediction system based on distributed supervised learning correspond to Figure 1 the respective steps in the user total active days prediction method based on distributed supervised learning described. Therefore, the operations, features, and beneficial effects described above for the user total active days prediction method based on distributed supervised learning also apply to the user total active days prediction system based on distributed supervised learning and the modules included therein, and will not be elaborated here.
[0110] The specific embodiments of the present invention have been described in detail above, but they are only examples, and the present invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, equivalent transformations and modifications made without departing from the spirit and scope of the present invention should all be covered within the scope of the present invention.
Claims
1. A method for predicting the total number of active days of users based on step-by-step supervised learning, characterized in that: The steps include: S1. Group the channel data added within the set time window before the target forecast date by month, and extract the multidimensional feature set, including channel attributes, time attributes, TAD values of each stage, and TAD ratios of each stage; S2, taking the longest cycle TAD as the target variable and the other features as input features, the supervised machine learning model is trained to obtain the initial prediction model, and the model parameters are optimized based on the prediction error rate indicator; S3, adopting an iterative prediction mechanism, gradually expanding the training samples to the data set containing historical prediction results, and predicting the TAD value of the shortened period in stages until the full-stage TAD of the target prediction period is covered; The specific steps include: S3.
1. When making the first prediction, use the initial sample data within the set time window to predict the longest period TAD and generate the first set of prediction results; S3.2, in each subsequent iteration, the prediction results of the previous step are incorporated into the training samples to expand the training data set; S3.
3. Based on the expanded data set, gradually shorten the target TAD period and add multi-stage TAD predictions. The target period of each iteration is reduced by a set number of days, and at least one TAD prediction with a shortened period is added. S3.4, repeat steps S3.2-S3.3 until the training data covers all shortened stages TAD of the target prediction cycle; S4. Perform final modeling based on the complete cycle data set, extract the TAD value and TAD ratio of all stages, and train to generate the final prediction model; S5. Input the new channel data into the final prediction model and output the target period TAD prediction result.
2. The method for predicting total active days of users based on step-by-step supervised learning according to claim 1, characterized in that: In step S1, the channel attributes include a unique ID for identifying different channels, the channel type, and the amount of new channels added; The time attribute indicates the date of the data record and identifies whether the date is a weekend or a holiday; The TAD value of each stage refers to the total number of active days in different time periods starting from the date the user was added; The TAD ratio of each stage refers to the ratio between TAD values in different stages, including the ratio of TADs of adjacent periods and the combination of TAD ratios of an initial period and subsequent periods, and is used to reflect the changing trend of user activity.
3. The method for predicting total active days of users based on step-by-step supervised learning according to claim 1, characterized in that: In step S2, the supervised machine learning model is a LightGBM model or an XGBoost model.
4. The method for predicting total active days of users based on step-by-step supervised learning according to claim 1, characterized in that: In step S2, the prediction error rate indicator is the mean absolute percentage error MAPE.
5. The method for predicting total active days of users based on step-by-step supervised learning according to claim 1, characterized in that: In step S3.3, the target period is reduced by 30 days each time, and the number of TAD predictions for the shortened period added in each iteration increases by one each time.
6. The method for predicting total active days of users based on step-by-step supervised learning according to claim 1, characterized in that: In step S4, the full-stage TAD value includes the number of active days in the 360-day full cycle from the date the user is added, and the TAD ratio includes the TAD growth ratio of consecutive dates.
7. The method for predicting total active days of users based on step-by-step supervised learning according to claim 1, characterized in that: In step S5, the newly added channel data must meet the data accumulation condition of at least 3 days before it can be input into the final prediction model for target period TAD prediction.
8. A user total active days prediction system based on step-by-step supervised learning, characterized in that: Includes the following modules: A data preprocessing module is used to group the channel data newly added within a set time window before the target prediction date by month, and extract a multidimensional feature set, wherein the multidimensional feature set includes channel attributes, time attributes, TAD values of each stage, and TAD ratios of each stage; The initial modeling module is used to train the longest cycle TAD as the target variable and the remaining features as input features through a supervised machine learning model to generate an initial prediction model and optimize the model parameters based on the prediction error rate indicator; Iterative prediction module, used to execute the iterative prediction mechanism, gradually expand the training samples to the data set containing historical prediction results, and predict the TAD value of the shortened period in stages until the full-stage TAD of the target prediction period is covered; The iterative prediction module comprises: An initial prediction unit, used to predict the longest period TAD using the initial sample data within a set time window during the first prediction, and generate a first set of prediction results; The data expansion unit is used to incorporate the prediction results of the previous step into the training samples in each subsequent iteration to expand the training data set; A cycle adjustment unit, which is used to gradually shorten the target TAD cycle based on the expanded data set and add multi-stage TAD predictions, wherein the target cycle of each iteration is reduced by a set number of days and at least one TAD prediction with a shortened cycle is added; A loop control unit, used for repeatedly calling the data expansion unit and the cycle adjustment unit until the training data covers all shortened stages TAD of the target prediction cycle; The final modeling module is used to perform final modeling based on the complete cycle data set, extract the TAD value and TAD ratio of the entire stage, and train and generate the final prediction model; The prediction output module is used to input the final prediction model into the newly added channel data and output the target period TAD prediction result.
9. The user total active days prediction system based on step-by-step supervised learning according to claim 8, characterized in that: In the data preprocessing module: The channel attributes include a unique ID for identifying different channels, the channel type, and the amount of new channels added; The time attributes include the date of the data record and whether it is a weekend or holiday; The TAD value of each stage includes the total number of active days in different time periods starting from the date when the user was added; The TAD ratios of each stage include the ratios of TADs of adjacent periods and the combination of TAD ratios of an initial period and subsequent periods, which are used to reflect the changing trend of user activity.
10. The user total active days prediction system based on step-by-step supervised learning according to claim 8, characterized in that: In the final modeling module, the full-stage TAD value includes the 360-day full-cycle active days from the date the user is added, and the TAD ratio includes the TAD growth ratio of consecutive dates.
Citation Information
Patent Citations
A daily active user number prediction method and device
CN109711897A
Data expansion method for monthly power generation prediction of new energy
CN110110908A