User mobile traffic dynamic prediction and evaluation method based on big data
Through dynamic prediction and evaluation methods of user mobile traffic based on big data, the traditional mobile traffic prediction methods in terms of data permissions and quality are solved, and accurate prediction and evaluation of mobile traffic usage is achieved to adapt to diversified business needs.
Patent Information
- Application Number
- CN202510118087.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional mobile traffic prediction methods are difficult to adapt to diversified business needs, have limited data permissions, and face the inability to fully guarantee data quality, resulting in low mobile traffic prediction accuracy.
The dynamic prediction and evaluation method of user mobile traffic based on big data is adopted, and the prediction value is calculated by collecting user historical traffic data, classifying and processing it, and using the mean threshold and exponential smoothing method, and finally evaluating the prediction accuracy through the error coefficient.
It realizes accurate prediction of users' usage of mobile traffic in the remaining time of this month, ensuring that high prediction accuracy and reliability are maintained while data permissions are restricted and data quality is uneven, and adapting to diversified business needs.
Smart Images

Figure CN120034893A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of big data, and specifically relates to a method for dynamically predicting and evaluating user mobile traffic based on big data. Background Art
[0002] In today's digital marketing environment, personalized services and precision marketing for users have become one of the key factors in improving customer satisfaction and market competitiveness. Among them, the prediction of mobile traffic usage is crucial for formulating effective marketing strategies. However, in actual applications, the demand direction in different scenarios is often not clear enough, data permissions are limited, and there is a problem that data quality cannot be fully guaranteed. These problems pose challenges to the accurate prediction of mobile traffic usage.
[0003] Traditional mobile traffic prediction methods usually rely on a large amount of historical data and complex feature dimensions, which requires high data quality and extensive access rights. However, in real application scenarios, data may be stored in a decentralized manner, access is limited, or even incomplete or inaccurate. In addition, in the face of dynamically changing market demands and user behavior patterns, fixed prediction models are difficult to adapt to diverse business needs. Therefore, the present invention provides a method for dynamic prediction and evaluation of user mobile traffic based on big data. Summary of the invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a dynamic prediction and evaluation method for user mobile traffic based on big data, which is used to solve the technical problems that traditional mobile traffic prediction methods are difficult to adapt to diversified business needs, have limited data permissions, and face the inability to fully guarantee data quality, resulting in low mobile traffic prediction accuracy.
[0005] To achieve the above object, the first aspect of the present invention provides a method for dynamic prediction and evaluation of user mobile traffic based on big data, comprising the following steps:
[0006] Step 1: Collect historical traffic data used by users and classify the historical traffic data;
[0007] Step 2: Divide the historical traffic data into two data sets, marked as the first data set and the second data set, and calculate the mean threshold of the two data sets respectively;
[0008] Step 3: Calculate the first prediction value of the traffic data used in the user prediction period according to the mean threshold of the two types of data sets;
[0009] Step 4: Calculate the second predicted value of the traffic data used in the user prediction period according to the exponential smoothing method;
[0010] Step 5: Calculate the final predicted value of the traffic data used in the user prediction period according to the first predicted value and the second predicted value;
[0011] Step 6: Calculate the error coefficient of the final prediction value; based on the error coefficient, evaluate the prediction accuracy.
[0012] Preferably, the collecting of historical traffic data used by users and classifying the historical traffic data include:
[0013] Extract some traffic data of users from historical data and classify the traffic data by time; among them, the traffic data is classified into type one according to the beginning, middle and end of the month; it is classified into type two according to weekdays and weekends; holidays are classified as weekends, and working days caused by adjusted holidays are classified as weekdays;
[0014] Combining type 1 and type 2, we get 6 combination types: beginning of the month - within the week, beginning of the month - weekend, middle of the month - within the week, middle of the month - weekend, end of the month - within the week, and end of the month - weekend;
[0015] The extracted traffic data are classified into corresponding combination types according to time.
[0016] Preferably, dividing the historical traffic data into two types of data sets and calculating the mean thresholds of the two types of data sets respectively include:
[0017] Based on the combination types of historical traffic data, the means of each combination type in the first data set and the second data set are calculated respectively, and marked as the mean thresholds of the first data set and the second data set; wherein the means of each combination type include the mean of the beginning of the month and within a week, the mean of the beginning of the month and the weekend, the mean of the middle of the month and within a week, the mean of the middle of the month and the weekend, the mean of the end of the month and within a week, and the mean of the end of the month and the weekend.
[0018] Preferably, the step of calculating the first predicted value of the traffic data used in the user prediction period according to the mean threshold of the two types of data sets includes:
[0019] By formula y 1 =k 1 ×S 1 +k 2 ×S 2 Calculate the first predicted value y1 of the user's traffic data; where k 1 , k 2 is the weight coefficient; S1 is the traffic prediction value of the first data set; S2 is the traffic prediction value of the second data set.
[0020] Preferably, the traffic prediction value of the first data set is the product of the mean threshold of each type combination in the first data set and the number of days in each type combination;
[0021] The traffic prediction value of the second data set is the product of the mean threshold of each type combination in the second data set and the number of days in each type combination.
[0022] Preferably, the step of calculating the second predicted value of the traffic data used within the user prediction period according to the exponential smoothing method includes:
[0023] Step P1: Calculate the first-order smoothing value of the flow data within the forecast period:
[0024]
[0025] Where T is the forecast period; is the first-order smoothing value of the flow data in T; Y T is the actual value of the flow data within T; is the first-order smoothing value of the flow data in T-1; α is the smoothing coefficient; the initial value is the mean of the first three items
[0026] Step P2: Calculate the second-order smoothing value of the flow data within the forecast period according to the first-order smoothing value of the flow data within the forecast period:
[0027]
[0028] in, is the second-order smoothed value of the flow data within T; is the second-order smoothed value of the flow data in T-1;
[0029] Step P3: Calculate the third-order smoothing value of the flow data within the forecast period according to the second-order smoothing value of the flow data within the forecast period:
[0030]
[0031] in, is the third-order smoothed value of the flow data within T; is the third-order smoothed value of the flow data in time T-1;
[0032] Step P4: According to the first-order smoothing value, second-order smoothing value and third-order smoothing value of the flow data within the prediction period, the prediction function is fitted as follows:
[0033]
[0034] Among them, t is the advance prediction period; at, bt, ct are the parameters of the prediction function;
[0035]
[0036] Step P5: Calculate the optimal smoothing coefficient based on the prediction function, return to step P1, and obtain the second prediction value
[0037] Preferably, the optimal smoothing coefficient is obtained in the following manner, including:
[0038]
[0039] Where i is the number of days, i∈[1,7]; SSE i is the residual sum of squares; SL MSE The minimum corresponding smoothing coefficient is the optimal smoothing coefficient.
[0040] Preferably, the calculating the final predicted value of the traffic data used in the user prediction period according to the first predicted value and the second predicted value includes:
[0041] By the formula Q = S 1 / S 2 Calculate the volatility coefficient Q;
[0042] when When , the final prediction value is the first prediction value;
[0043] when When the final prediction value = m 1 ×y 1 +(1-m 1 )×y 2 ;in,
[0044] When Q∈[0.4,0.75)], the final prediction value = m 2 ×y 1 +(1-m 2 )×y 2 ;in,
[0045] When Q∈[0,0.4)∪(2.5,+∞), the final prediction value is the second prediction value.
[0046] Preferably, the calculation of the error coefficient of the final predicted value includes:
[0047] By formula Calculate the error coefficient R of the final prediction value 2 ; Among them, SSE is the residual sum of squares; SST is the deviation sum of squares.
[0048] Preferably, the evaluating the accuracy of the prediction result based on the error coefficient includes:
[0049] When the error coefficient is within the preset error threshold range, the prediction result is accurate; when the error coefficient is not within the preset error threshold range, the prediction result is inaccurate.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] The present invention realizes accurate prediction of the user's mobile traffic usage for the rest of the month, and effectively evaluates the prediction results and system performance. Taking into account high interpretability and low feature dimension dependency, it ensures that high prediction accuracy and reliability can be maintained even when data permissions are limited, data quality is uneven, and feature field data quality is poor. Combining the two prediction values to determine the final prediction result, the prediction takes into account both the trend analysis based on historical data and the feature information of the current data set, which improves the comprehensiveness and accuracy of the prediction results. Finally, the prediction accuracy is evaluated by calculating the error coefficient to ensure that the model can maintain good adaptability and robustness in different application contexts. In addition, it also reduces excessive reliance on any single data source or prediction method, and improves the overall adaptability and robustness through diversified data processing methods and prediction models. This makes the present invention not only suitable for traditional mobile traffic prediction scenarios, but also can flexibly respond to various complex application requirements and practical marketing problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0053] Figure 1 It is a schematic diagram of the process of the present invention;
[0054] Figure 2 It is a schematic flow chart of the second prediction value analysis method of mobile traffic data of the present invention. DETAILED DESCRIPTION
[0055] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] See also Figure 1 The first aspect of the present invention provides a method for dynamically predicting and evaluating user mobile traffic based on big data, comprising the following steps:
[0057] Step 1: Collect historical traffic data used by users and classify the historical traffic data;
[0058] Extract some traffic data of users from historical data and classify the traffic data by time; among them, the traffic data is classified into type one according to the beginning, middle and end of the month; it is classified into type two according to weekdays and weekends; holidays are classified as weekends, and working days caused by adjusted holidays are classified as weekdays;
[0059] Combining type 1 and type 2, we get 6 combination types: beginning of the month - within the week, beginning of the month - weekend, middle of the month - within the week, middle of the month - weekend, end of the month - within the week, and end of the month - weekend;
[0060] The extracted traffic data are classified into corresponding combination types according to time.
[0061] For example, extract the data of a user in the previous 63 days from the current time from the historical data. The data in the previous 63 days is the daily mobile traffic data usage and is divided by natural day.
[0062] Every day in each natural month is divided into Type 1 according to the beginning / middle / end of the month: beginning of the month (1st to 10th), middle of the month (11th to 20th), and end of the month (21st and later); according to the weekday / end of the month, it is divided into Type 2: weekdays (Monday, Tuesday, Wednesday, Thursday, Friday), weekends (Saturday, Sunday);
[0063] Step 2: Divide the historical traffic data into two data sets, marked as the first data set and the second data set, and calculate the mean threshold of the two data sets respectively;
[0064] Specifically, based on the combination types of historical traffic data, the means of each combination type in the first data set and the second data set are calculated respectively, and marked as the mean thresholds of the first data set and the second data set; wherein the means of each combination type include the mean of the beginning of the month and the week, the mean of the beginning of the month and the weekend, the mean of the middle of the month and the week, the mean of the middle of the month and the weekend, the mean of the end of the month and the week, and the mean of the end of the month and the weekend.
[0065] For example, suppose that the traffic data of a user in the last 63 days is collected, and the collected traffic data is divided into the first data set: data from the previous 1-28 days, and the second data set: data from the previous 29-63 days;
[0066] Calculate the mean threshold for each combination: the mean threshold for the first data set (first 1-28 days): the mean a11 of the beginning of the month and the week, the mean a12 of the beginning of the month and the weekend, the mean a13 of the middle of the month and the week, the mean a14 of the middle of the month and the weekend, the mean a15 of the end of the month and the week, and the mean a16 of the end of the month and the weekend; the mean threshold for the second data set (first 29-63 days): the mean a21 of the beginning of the month and the week, the mean a22 of the beginning of the month and the weekend, the mean a23 of the middle of the month and the week, the mean a24 of the middle of the month and the weekend, the mean a25 of the end of the month and the week, and the mean a26 of the end of the month and the weekend.
[0067] Step 3: Calculate the first prediction value of the traffic data used in the user prediction period according to the mean threshold of the two types of data sets;
[0068] Specifically, the product of the mean threshold of each type combination in the first data set and the number of days in each type combination is calculated to obtain the traffic prediction value of the first data set, which is marked as S1; the product of the mean threshold of each type combination in the second data set and the number of days in each type combination is calculated to obtain the traffic prediction value of the second data set, which is marked as S2;
[0069] For example, to calculate the predicted traffic usage value for the remaining 4 days of a month, the 4 days are of the "end of the month - weekend" combination type, and the predicted value for this part is: S 15 =4×a 15 , S 25 =4×a 25 .
[0070] By formula y 1 =k 1 ×S 1 +k 2 ×S 2 Calculate the first predicted value y1 of the user's traffic data; where k 1 , k 2 is the weight coefficient.
[0071] According to previous exploration and data rules, the present invention believes that for most users, user data in the more recent time series dimension has a greater impact on the user's subsequent traffic usage, and takes k1=0.6 and k2=0.4 for calculation.
[0072] Here, you can also set the values of k1 and k2 according to different business needs. For example, for marketing security considerations, if you hope that the predicted traffic usage is less than the user's actual traffic usage, it is safer. You can set k_1=0.55 and k_2=0.35 for exploration.
[0073] Step 4: Calculate the second predicted value of the traffic data used in the user prediction period according to the exponential smoothing method;
[0074] See also Figure 2, the process of obtaining the second prediction value is as follows:
[0075] Step P1: Calculate the first-order smoothing value of the flow data within the forecast period:
[0076]
[0077] Where T is the forecast period; is the first-order smoothing value of the flow data in T; Y T is the actual value of the flow data within T; is the first-order smoothing value of the flow data in T-1; α is the smoothing coefficient; the initial value is the mean of the first three items
[0078] Step P2: Calculate the second-order smoothing value of the flow data within the forecast period according to the first-order smoothing value of the flow data within the forecast period:
[0079]
[0080] in, is the second-order smoothed value of the flow data within T; is the second-order smoothed value of the flow data in T-1;
[0081] Step P3: Calculate the third-order smoothing value of the flow data within the forecast period according to the second-order smoothing value of the flow data within the forecast period:
[0082]
[0083] in, is the third-order smoothed value of the flow data within T; is the third-order smoothed value of the flow data in time T-1;
[0084] Step P4: According to the first-order smoothing value, second-order smoothing value and third-order smoothing value of the flow data within the prediction period, the prediction function is fitted as follows:
[0085]
[0086] Among them, t is the advance prediction period; at, bt, ct are the parameters of the prediction function;
[0087]
[0088] Step P5: Calculate the optimal smoothing coefficient based on the prediction function, return to step P1, and obtain the second prediction value
[0089]
[0090] Where i is the number of days, i∈[1,7]; SSE i is the residual sum of squares; SL MSE The minimum corresponding smoothing coefficient is the optimal smoothing coefficient; the smoothing coefficient α is between 0 and 1, and the smaller α is, the stronger the smoothing effect is.
[0091] For example: T = 28, that is, prediction starts from the 28th cycle of data. For example, when predicting the 29th cycle of data, t is 1, and when predicting the 30th cycle of data, t is 2, and so on, to get the second prediction value.
[0092] Step 5: Calculate the final predicted value of the traffic data used in the user prediction period according to the first predicted value and the second predicted value;
[0093] By the formula Q = S 1 / S 2 Calculate the volatility coefficient Q;
[0094] It should be noted that the fluctuation coefficient Q reflects the user's recent traffic usage habits and changes in usage habits at longer time intervals. When the Q value is too large or too small, the system makes adjustments based on the second predicted value.
[0095] when When , the user's traffic usage habits are relatively stable, and the final prediction value is the first prediction value;
[0096] when When the final prediction value = m 1 ×y 1 +(1-m 1 )×y 2 ;in,
[0097] When Q∈[0.4,0.75)], the final prediction value = m 2 ×y 1 +(1-m 2 )×y 2 ;in,
[0098] When Q∈[0,0.4)∪(2.5,+∞), the traffic usage of some users before and after is too different, and the final prediction value is the second prediction value.
[0099] Step 6: Calculate the error coefficient of the final prediction value; based on the error coefficient, evaluate the prediction accuracy;
[0100] By formula Calculate the error coefficient R of the final prediction value 2; Among them, SSE is the residual sum of squares, that is, the sum of squares of predicted value minus actual value; SST is the deviation sum of squares, that is, the sum of squares of actual value minus the average of actual value.
[0101] When the error coefficient is within the preset error threshold range, the prediction result is accurate; when the error coefficient is not within the preset error threshold range, the prediction result is inaccurate.
[0102] Part of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is a formula closest to the actual situation obtained by software simulation of a large amount of collected data; the preset parameters and preset thresholds in the formula are set by technical personnel in this field according to actual conditions or obtained through simulation of a large amount of data.
[0103] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A method for dynamic prediction and evaluation of user mobile traffic based on big data, characterized in that: include: Collect historical traffic data used by users and classify the historical traffic data; The historical traffic data is divided into two types of data sets, which are marked as the first data set and the second data set, and the mean thresholds of the two types of data sets are calculated respectively; Calculate the first prediction value of the traffic data used in the user prediction period according to the mean threshold of the two types of data sets; Calculate the second prediction value of the traffic data used in the user prediction period according to the exponential smoothing method; Calculate the final predicted value of the traffic data used in the user prediction period according to the first predicted value and the second predicted value; Calculate the error coefficient of the final predicted value; Based on the error coefficient, the prediction accuracy is evaluated.
2. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 1 is characterized in that: The collecting of historical traffic data used by users and classifying the historical traffic data include: Extract some traffic data of users from historical data and classify the traffic data by time; among them, the traffic data is classified into type one by month and type two by week; Combine type 1 and type 2 to obtain a combined type; The extracted traffic data are classified into corresponding combination types according to time.
3. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 2 is characterized in that: The method of dividing the historical traffic data into two types of data sets and calculating the mean thresholds of the two types of data sets respectively includes: Based on the combination types of the historical traffic data, the means of each combination type in the first data set and the second data set are calculated respectively, and are marked as the mean thresholds of the first data set and the second data set.
4. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 3 is characterized in that: The step of calculating the first predicted value of the traffic data used in the user prediction period according to the mean threshold of the two types of data sets includes: The first predicted value y1 of the user's traffic data is calculated by the formula y1=k1×S1+k2×S2; wherein k1 and k2 are weight coefficients; S1 is the traffic prediction value of the first data set; and S2 is the traffic prediction value of the second data set.
5. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 4 is characterized in that: The traffic prediction value of the first data set is the product of the mean threshold of each type combination in the first data set and the number of days in each type combination; The traffic prediction value of the second data set is the product of the mean threshold of each type combination in the second data set and the number of days in each type combination.
6. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 4 is characterized in that: The step of calculating the second predicted value of the traffic data used within the user prediction period according to the exponential smoothing method includes: Step P1: Calculate the first-order smoothing value of the flow data within the forecast period: Where T is the forecast period; is the first-order smoothing value of the flow data in T; Y T is the actual value of the flow data within T; is the first-order smoothing value of the flow data in T-1; α is the smoothing coefficient; the initial value is the mean of the first three items Step P2: Calculate the second-order smoothing value of the flow data within the forecast period according to the first-order smoothing value of the flow data within the forecast period: in, is the second-order smoothed value of the flow data within T; is the second-order smoothed value of the flow data in T-1; Step P3: Calculate the third-order smoothing value of the flow data within the forecast period according to the second-order smoothing value of the flow data within the forecast period: in, is the third-order smoothed value of the flow data within T; is the third-order smoothed value of the flow data in time T-1; Step P4: According to the first-order smoothing value, second-order smoothing value and third-order smoothing value of the flow data within the prediction period, the prediction function is fitted as follows: Among them, t is the advance prediction period; at, bt, ct are the parameters of the prediction function; Step P5: Calculate the optimal smoothing coefficient based on the prediction function, return to step P1, and obtain the second prediction value 7. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 6 is characterized in that: The optimal smoothing coefficient is obtained in the following manner, including: Where i is the number of days, i∈[1,7]; SSE i is the residual sum of squares; SL MSE The minimum corresponding smoothing coefficient is the optimal smoothing coefficient.
8. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 6 is characterized in that: The calculating, according to the first prediction value and the second prediction value, a final prediction value of the traffic data used within the user prediction period includes: The volatility coefficient Q is calculated by the formula Q = S1 / S2; when When , the first predicted value is marked as the final predicted value; when When , the final prediction value is calculated by the formula final prediction value = m1×y1+(1-m1)×y2; where, When Q∈[0.4,0.75)], the final prediction value is calculated by the formula final prediction value = m2×y1+(1-m2)×y2; where, When Q∈[0,0.4)∪(2.5,+∞), the second predicted value is marked as the final predicted value.
9. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 1 is characterized in that: The calculation of the error coefficient of the final prediction value includes: By formula Calculate the error coefficient R of the final prediction value 2 ; Among them, SSE is the residual sum of squares; SST is the deviation sum of squares.
10. The method for dynamic prediction and evaluation of user mobile traffic based on big data according to claim 9 is characterized in that: The accuracy of the prediction results is evaluated based on the error coefficient, including: When the error coefficient is within the preset error threshold range, the prediction result is marked as accurate; when the error coefficient is not within the preset error threshold range, the prediction result is marked as inaccurate.