New store sales prediction method, device, equipment and storage medium
By acquiring and processing sales data from the target region and using the Prophet algorithm to train a time series forecasting model, the problem of insufficient data in new store sales forecasting was solved, achieving automated and accurate sales forecasting and optimizing new store product strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 创优数字科技(广东)有限公司
- Filing Date
- 2024-12-30
- Publication Date
- 2026-05-01
AI Technical Summary
In the chain retail industry, new stores lack their own sales history data, which increases the difficulty of predicting the demand for seasonal products and holiday best-selling products. Existing technologies rely on manual judgment and have limited front-end data volume and interpretability.
By acquiring sales data of the target area within a preset time period, cleaning and stitching the data, and using time series prediction models such as the Prophet algorithm to train the optimal parameter combination, the sales volume of the new store can be predicted.
It has achieved automated sales forecasting for new stores, increased the amount and validity of data, improved the accuracy and practicality of forecasts, and enabled the optimization of product strategies in target stores.
Smart Images

Figure CN119740710B_ABST
Abstract
Description
New store sales forecasting methods, devices, equipment, and storage media Technical Field
[0001] This application relates to the field of data analysis, and in particular to a method, apparatus, device, and storage medium for predicting new store sales. Background Technology
[0002] In the chain retail industry, accurate forecasting of future sales for new stores is crucial to providing a data foundation for rational inventory management, marketing strategy development, and supply chain management. However, the lack of historical sales data for new stores significantly increases the difficulty of forecasting demand for seasonal and holiday-driven best-selling items.
[0003] Currently, for new store sales forecasting, especially time-related forecasting, the common approach is to substitute new store sales demand with sales demand from similar historical stores, and the forecasting task is completed manually. This method has limitations in terms of front-end data volume and interpretability, and its forecasting effectiveness relies on human experience. Therefore, an automated new store sales forecasting method is needed to address these issues. Summary of the Invention
[0004] The purpose of this application is to at least address one of the aforementioned technical deficiencies, particularly the limitations in the amount and interpretability of front-end data in the new store sales forecasting process of existing technologies, and the fact that the forecasting effect depends on human experience.
[0005] Firstly, this application provides a method for predicting sales of new stores, the method including:
[0006] Obtain the first sales data for the target area within a preset time period;
[0007] The first sales data includes sales data of multiple existing stores in the target area and sales data of the target store;
[0008] The first sales data is preprocessed to obtain the second sales data;
[0009] The data preprocessing includes data cleaning and data splicing based on the time-series attributes of the corresponding sales data;
[0010] Set the target parameters and corresponding parameter search range for the time series forecasting model, and train the time series forecasting model based on the second sales data;
[0011] The third sales data of the target store is predicted using the time series prediction model.
[0012] As an optional implementation, the step of preprocessing the first sales data to obtain the second sales data includes:
[0013] Based on the time-series attributes of the first sales data, outlier data is identified and processed to obtain cleaned data;
[0014] Based on the temporal attributes of the cleaning data, data splicing is performed to obtain the second sales data.
[0015] As an optional implementation, the step of performing data concatenation based on the temporal attributes of the cleaned data to obtain second sales data includes:
[0016] Determine the first pieced data in the cleaning data that belongs to the target store;
[0017] Based on the temporal attributes of the first spliced data, determine the second spliced data belonging to the existing store in the cleaning data within the preset time period;
[0018] Wherein, the temporal attributes of the second spliced data precede the temporal attributes of the first spliced data;
[0019] The second and first spliced data are spliced together according to the time sequence attribute to obtain the second sales data.
[0020] As an optional implementation, the method further includes:
[0021] Determine the target calendar based on the temporal activity characteristics of the target area;
[0022] Based on the target calendar, determine the holiday marking parameters corresponding to the second sales data, and use the holiday marking parameters as the time classification attribute of the second sales data, merging them into the second sales data.
[0023] As an optional implementation, the target parameters include:
[0024] Seasonal model parameters, growth model parameters, holiday prior scale parameters, weekly seasonality parameters, annual seasonality parameters, change point prior scale parameters, change point range parameters, and holiday marking parameters;
[0025] After setting the target parameters and corresponding parameter search range of the time series prediction model, the method further includes:
[0026] Determine the parameter search process;
[0027] The parameter search process includes one or more of the following: grid search, prior search, and random search.
[0028] Based on the search results for the target parameters according to the parameter search process, the parameter combination of the time series prediction model is determined.
[0029] As an optional implementation, the method further includes:
[0030] Obtain the building distribution information for the target area;
[0031] Based on the building distribution and the first sales data, the user distribution is determined;
[0032] Based on the user distribution, the fourth sales data of the target store is predicted using the time series prediction model.
[0033] The fourth sales data is used to indicate product attributes and corresponding user profiles.
[0034] As an optional implementation, after predicting the fourth sales data of the target store based on the user distribution using the time series prediction model, the method further includes:
[0035] Based on the fourth sales data of the target store, determine the optimization plan;
[0036] The optimization scheme is used to indicate the parameter adjustment method for the brand, category, model, and placement of each product in the target store.
[0037] Secondly, this application provides a new store sales forecasting device, the device comprising:
[0038] The acquisition module is used to acquire the first sales data of the target area within a preset time period;
[0039] The first sales data includes sales data of multiple existing stores in the target area and sales data of the target store;
[0040] The processing module is used to preprocess the first sales data to obtain the second sales data;
[0041] The data preprocessing includes data cleaning and data splicing based on the time-series attributes of the corresponding sales data;
[0042] The processing module is also used to set the target parameters and corresponding parameter search range of the time series prediction model, and to train the time series prediction model based on the second sales data.
[0043] The processing module is also used to predict and obtain the third sales data of the target store through the time series prediction model.
[0044] Thirdly, this application provides a computer device including one or more processors and a memory storing computer-readable instructions that, when executed by the one or more processors, perform the steps of the method described in the first aspect.
[0045] Fourthly, this application provides a storage medium storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the method described in the first aspect.
[0046] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0047] Based on any of the above embodiments, this application obtains the first sales data of the target area within a preset time period, thereby obtaining the total sales situation since the opening of the new store and the sales situation of the existing store within a certain period of time, avoiding the problem of insufficient new store data. While focusing on the new store data itself, it expands the data volume. The sales data of the existing store and the target store are used together as the basic data for sales forecasting, namely the first sales data. Further, based on the first sales data, data cleaning and data splicing preprocessing tasks are performed to obtain the second sales data. This allows the time series prediction model with preset search parameters to be trained, obtaining the time series prediction model under the optimal parameter combination, and then predicting the third sales data of the target store to indicate the expected sales situation of each category. Thus, by adding the sales data of the new store itself on the basis of historical store sales data, the front-end data volume and data effectiveness are improved, and automated new store sales forecasting is realized. Furthermore, the new store sales forecast can be used to optimize the products in the target store, which also improves the practicality of new store sales forecasting. Attached Figure Description
[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0049] Figure 1 is a flowchart illustrating a new store sales forecasting method provided in one embodiment of this application;
[0050] Figure 2 is a schematic diagram of an application scenario for a new store sales forecasting method provided in an embodiment of this application;
[0051] Figure 3 is an internal structure diagram of the computer device provided in an embodiment of this application. Detailed Implementation
[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] Demand forecasting for new chain retail stores refers to predicting future sales demand for a new store by analyzing various data and market factors before its opening. This process is crucial for the success of a new store because it helps businesses rationally manage inventory, develop marketing strategies, and optimize supply chain management. Because new stores lack their own historical sales data, forecasting becomes more complex, especially for seasonal items and items with high demand during holidays. Currently, a common solution is to use the sales demand of similar historical stores as a benchmark for the new store's sales demand, based on assessments of front-end operations and sales.
[0054] Specifically, in some implementations, existing technologies mainly rely on judgments from the front end of the supply chain, using the product demand of similar stores as the demand for new stores, which may have the following limitations:
[0055] Insufficient front-end data: Sales forecasting often relies on historical data and market trends. However, the lack of long-term historical data at the front end means that forecasts are often based only on recent data or some incidental events, which affects the accuracy of forecasts and thus the effectiveness of decision-making.
[0056] Lack of data analysis skills: Salespeople are often adept at dealing with customers but may lack professional data analysis skills. Data analysis requires specialized skills and tools, while salespeople may rely more on intuition and experience, which can lead to biased predictions.
[0057] Judgment granularity is too small: Sales forecasts are often granular, and may only focus on certain specific products or stores. From a statistical point of view, the smaller the granularity, the less accurate the forecast.
[0058] Ignoring sales data of new stores: Short-term sales history of new stores is also a reference, such as recent sales trends. If the existing sales data of new stores is not considered and only the data of replacement stores is used, the accuracy of prediction will be further affected.
[0059] Therefore, the technical concept of this application lies in the fact that the method provided by this application obtains the first sales data of the target area within a preset time period, thereby obtaining the total sales situation since the opening of the new store and the sales situation of the existing store within a certain period of time, avoiding the problem of insufficient new store data. While focusing on the new store data itself, it expands the data volume. The sales data of the existing store and the target store are used together as the basic data for sales prediction, namely the first sales data. Furthermore, based on the first sales data, data cleaning and data splicing preprocessing tasks are performed to obtain the second sales data. This allows the time series prediction model with preset search parameters to be trained, obtaining the time series prediction model under the optimal parameter combination, and then predicting the third sales data of the target store to indicate the expected sales situation of each category. Thus, by adding the sales data of the new store itself on the basis of historical store sales data, the front-end data volume and data effectiveness are improved, and automated new store sales prediction is realized. Furthermore, the new store sales prediction can be used to optimize the products in the target store, which also improves the practicality of new store sales prediction.
[0060] To address the aforementioned technical issues, this application provides an automated new store sales forecasting method that utilizes sufficient data, including both existing and new store data, and can be used simultaneously for judging overall macro-level sales performance or optimizing micro-level performance. Based on a time series forecasting model, such as the Prophet algorithm, a complete forecasting process is implemented. This method can make reasonable baseline forecasts for products with sales patterns based on a large amount of historical data, without relying on front-end input, thus solving the problems of insufficient front-end data and lack of data analysis capabilities. The algorithm integrates existing data from new stores and sales data from the region where the new stores are located, leveraging regional economies of scale to improve the forecasting accuracy for new stores, thereby solving the problems of excessively fine-grained judgments and the neglect of new store sales data. Specific solutions can be found in the corresponding descriptions of each implementation.
[0061] Please refer to Figure 1, which is a flowchart illustrating a new store sales forecasting method provided in an embodiment of this application. As shown in Figure 1, the method includes:
[0062] S101. Obtain the first sales data of the target area within a preset time period;
[0063] The first sales data includes sales data of multiple existing stores in the target area and sales data of the target store;
[0064] Specifically, relevant data can be collected from the data platform according to the required forecast time range. If the target store is a completely unopened new store, the regional sales data for the new store's products will be the median sales volume of the products in the same region. However, in some feasible application scenarios in this application, historical sales data from multiple existing stores and all sales data from the target store can be obtained, and then filtered, cleaned, and stitched together in the subsequent data processing.
[0065] S102. Perform data preprocessing on the first sales data to obtain the second sales data;
[0066] The data preprocessing includes data cleaning and data splicing based on the time-series attributes of the corresponding sales data;
[0067] Specifically, sales data from other stores and new stores within the region can be preprocessed, and existing store sales data can be combined with that of the target store to form a continuous time series as input for Prophet. In this way, Prophet can learn from the region's historical sales and apply it to new store forecasting.
[0068] S103. Set the target parameters and corresponding parameter search range of the time series prediction model, and train the time series prediction model based on the second sales data.
[0069] S104. Using the time series prediction model, predict and obtain the third sales data of the target store.
[0070] In this application, the time series forecasting model may include the aforementioned Prophet algorithm, which is an open-source time series forecasting algorithm that performs well when processing time series data with complex patterns such as seasonality and holiday effects.
[0071] The Prophet algorithm will be introduced below to illustrate its compatibility with the application scenarios provided in this application.
[0072] Algorithm characteristics:
[0073] Automatic handling of seasonality and trends: The Prophet algorithm can automatically identify and process seasonal and trend variations in time series data. It eliminates the need for users to perform complex seasonal decomposition or trend fitting operations on the data beforehand; instead, it automatically captures these characteristics through built-in model structures and parameter settings. For example, for sales data, it can automatically identify annual peak and off-peak sales seasons (seasonality) and long-term sales growth or decline trends (trends).
[0074] Considering the holiday effect: This algorithm pays special attention to the impact of holidays on time series data. It allows users to easily define holiday calendars and adjust forecasts based on the characteristics of different holidays. For example, when forecasting sales data in the retail industry, it can accurately account for the significant fluctuations in sales during major holidays such as Spring Festival and National Day, thus making the forecasts more closely reflect reality.
[0075] Flexibility and Interpretability: The Prophet algorithm offers high flexibility, allowing users to adapt to different data characteristics and prediction needs by adjusting a range of parameters. Furthermore, its model structure is relatively simple and interpretable, unlike some deep learning algorithms which are often described as black boxes. Users can generally understand how the model makes predictions based on factors such as trends, seasonality, and holidays in the data.
[0076] Algorithm principle:
[0077] Model Composition: The basic model of the Prophet algorithm can be represented in the following form:
[0078] y(t) = g(t) + s(t) + h(t) + e(t)
[0079] Where y(t) is the time series value to be predicted (such as the sales volume of a certain product at time t).
[0080] g(t) is a trend function used to describe the long-term growth or decline trend of a time series. Common trend functions include linear trend functions and saturated growth trend functions (such as the logistic growth function), and the specific form can be selected according to the characteristics of the data and user settings.
[0081] s(t) is a seasonality function used to capture seasonal fluctuations in time series. It can be an additive or multiplicative seasonality pattern. When using a seasonality function, the number of seasonal cycles needs to be determined (e.g., if there are 12 months in a year, the number of seasonal cycles may be 12).
[0082] In fact, the seasons in the seasonal function can be understood as a periodic pattern, rather than necessarily the seasons in the natural calendar. They can be set and understood according to the specific implementation method.
[0083] h(t) is the holiday function, used to handle the impact of holidays on the time series. Based on the user-defined holiday calendar, i.e., the target calendar mentioned in this application, a corresponding impact function is set for each holiday. These impact functions are usually functions derived from experience or data fitting, used to describe the fluctuations of the time series during the holiday period.
[0084] e(t) is the error term, which is assumed to follow a normal distribution with a mean of zero, and is used to represent random fluctuations that the model cannot completely capture.
[0085] Trend Fitting: When determining the trend function, the Prophet algorithm employs a piecewise linear fitting method called "change point detection." It identifies points in the data where sales volumes change significantly (change points), and then uses linear fitting between these change points to describe the trend. For example, suppose a product underwent a promotional campaign for a period of time, leading to a sudden and substantial increase in sales. The algorithm will detect this change point and apply different linear trends before and after the change point, thus more accurately capturing changes in long-term trends.
[0086] Seasonal Fitting: For seasonal fitting, as mentioned earlier, the algorithm estimates the seasonal function by analyzing historical data based on the seasonal pattern (addition or multiplication) selected by the user. It considers seasonal fluctuations at different periods, such as monthly, weekly, and yearly seasonality, depending on the data and user settings.
[0087] Holiday Fitting: When dealing with the holiday effect, users first need to define an accurate holiday calendar. Then, for each holiday, the algorithm fits a corresponding holiday function based on the fluctuations in the time series of holiday periods in historical data. For example, for the Spring Festival, by analyzing sales data during the Spring Festival in previous years, the function form of the Spring Festival's impact on sales is determined so that the impact of the Spring Festival can be accurately considered in forecasting.
[0088] Algorithm application process:
[0089] Data preparation:
[0090] Collect the time series data to be predicted, such as the sales data corresponding to this application. Ensure the integrity and accuracy of the data, including the correct setting of timestamps and data cleaning (removal of outliers, handling of missing values, etc.).
[0091] Parameter settings:
[0092] Based on the data characteristics and prediction needs, set various parameters for the Prophet algorithm, such as Seasonality_mode, Growth, Holidays_prior_scale, Weekly_seasonality, Yearly_seasonality, Changepoint_prior_scale, and Changepoint_range. Different values for these parameters will affect how the algorithm captures and processes factors such as trends, seasonality, and holidays in the data; their meanings and impacts have been explained in detail previously.
[0093] Model training:
[0094] The prepared data and set parameters are fed into the Prophet algorithm model for training. During training, the algorithm estimates the various parameters in the model by minimizing prediction errors (such as mean squared error, mean absolute error, etc.) based on information such as trends, seasonality, and holidays in the data.
[0095] Model evaluation:
[0096] The trained model is evaluated using metrics such as mean squared error (MSE), mean absolute error (MAE), and root mean square error (RMSE). By comparing the evaluation metric values of the model under different parameter settings, the model's performance can be determined, allowing for further parameter optimization or model improvement.
[0097] Predictive applications:
[0098] Once the model is evaluated and approved, the trained model can be used to predict future time series values, namely, the new store sales forecast described in this application.
[0099] The method provided in this embodiment obtains the first sales data of the target area within a preset time period, thereby obtaining the total sales situation since the opening of the new store and the sales situation of the existing stores within a certain period of time. This avoids the problem of insufficient new store data. While focusing on the new store data itself, it expands the data volume. The sales data of the existing stores and the target stores are used together as the basic data for sales forecasting, namely the first sales data. Further, based on the first sales data, data cleaning and data splicing preprocessing tasks are performed to obtain the second sales data. This allows the time series prediction model with preset search parameters to be trained, obtaining the time series prediction model under the optimal parameter combination, and then predicting the third sales data of the target store to indicate the expected sales situation of each category. Thus, by adding the sales data of the new store itself on the basis of historical store sales data, the front-end data volume and data effectiveness are improved, and automated new store sales forecasting is realized. Furthermore, the new store sales forecast can be used to optimize the products in the target store, which also improves the practicality of new store sales forecasting.
[0100] As an optional implementation, the step of preprocessing the first sales data to obtain the second sales data includes:
[0101] Based on the time-series attributes of the first sales data, outlier data is identified and processed to obtain cleaned data;
[0102] The process of obtaining cleaned data may include: separating sales data for holidays from the time series, using the rolling standard deviation method (3 steps in the rolling window) to identify holidays where commodity sales fluctuate abnormally, and then setting the sales corresponding to these holidays to null values.
[0103] Additionally, outlier detection is performed on weekdays and weekends. Sales data for holidays are removed from the time series, and the Hodrick-Prescott filter is used to identify dates where product sales fluctuate abnormally. The sales figures for these dates are then set to null values.
[0104] In addition to setting it to null, it can also be filled using interpolation methods in numerical analysis to ensure data continuity.
[0105] Based on the temporal attributes of the cleaning data, data splicing is performed to obtain the second sales data.
[0106] This implementation method identifies outliers in the first sales data by utilizing its temporal attributes, such as holidays, weekdays (excluding holidays), and weekends (excluding holidays). It then sets sales-related attributes to null values or uses interpolation to correct the values, obtaining cleaned data. Based on the temporal attributes of the cleaned data, it concatenates the data in chronological order to obtain processed second sales data. This improves the effectiveness and interpretability of the second sales data, ensuring sufficient front-end data volume and validity, and enabling automated new store sales forecasting.
[0107] As an optional implementation, the step of performing data concatenation based on the temporal attributes of the cleaned data to obtain second sales data includes:
[0108] Determine the first pieced data in the cleaning data that belongs to the target store;
[0109] Based on the temporal attributes of the first spliced data, determine the second spliced data belonging to the existing store in the cleaning data within the preset time period;
[0110] Wherein, the temporal attributes of the second spliced data precede the temporal attributes of the first spliced data;
[0111] The second and first spliced data are spliced together according to the time sequence attribute to obtain the second sales data.
[0112] In this embodiment, since the original first sales data includes sales data of multiple existing stores in the target area and sales data of the target store, the cleaned data can also include two parts: sales data of existing stores and sales data of the target store. First, all data belonging to the target store is determined. Then, according to a preset time period, the second concatenated data belonging to the existing store and whose time sequence is mutually exclusive with the time sequence of the first concatenated data corresponding to the target store is determined. Thus, data concatenation can be performed according to the time sequence. During the data processing, the sales data of the new store itself is added on the basis of the historical store sales data, which improves the amount of front-end data and data effectiveness, and realizes automated new store sales prediction.
[0113] As an optional implementation, the method further includes:
[0114] Determine the target calendar based on the temporal activity characteristics of the target area;
[0115] Based on the target calendar, determine the holiday marking parameters corresponding to the second sales data, and use the holiday marking parameters as the time classification attribute of the second sales data, merging them into the second sales data.
[0116] Specifically, based on the patterns of social activities, the data processing method related to the target calendar is described as follows: The holiday calendar built into the Prophet algorithm may deviate from the actual holidays, therefore a custom holiday calendar is needed. Holiday information is obtained using the open-source Python package chinese-calender and then processed into the format required by Prophet. Furthermore, due to work schedule adjustments, sales before and after important holidays may differ from those on regular weekdays and weekends; therefore, these dates can also be treated as holidays, allowing the Prophet algorithm to determine their impact on sales.
[0117] Since the original calendar built into the time series forecasting model may deviate from actual needs, this implementation method determines the target calendar based on the time activity characteristics of user groups or the region as a whole in the target area, and determines the holiday marking parameters of the second sales data as the time classification attribute of the second sales data. These parameters are then merged and added to the second sales data to complete the setting of the target calendar and further processing of the second sales data. This allows the second sales data to be associated with the time activity characteristics of the target area, increasing the effectiveness and interpretability of the second sales data, thereby improving the effect of new store sales forecasting.
[0118] As an optional implementation, the target parameters include:
[0119] Seasonal model parameters, growth model parameters, holiday prior scale parameters, weekly seasonality parameters, annual seasonality parameters, change point prior scale parameters, change point range parameters, and holiday marking parameters;
[0120] After setting the target parameters and corresponding parameter search range of the time series prediction model, the method further includes:
[0121] Determine the parameter search process;
[0122] The parameter search process includes one or more of the following: grid search, prior search, and random search.
[0123] Based on the search results for the target parameters according to the parameter search process, the parameter combination of the time series prediction model is determined.
[0124] Specifically, a feasible interpretation of each target parameter and a feasible search range based on a specific application scenario include:
[0125] Seasonality_mode (seasonality mode parameter):
[0126] Values: Multiplicative (multiplicative seasonality pattern), additive (additive seasonality pattern).
[0127] Meaning and Impact: This parameter determines how the model handles the impact of seasonality on sales. In the additive seasonality model, sales are assumed to be the simple sum of the trend term, seasonality term, and error term: Sales = Trend + Seasonality + Error. For example, for a product, seasonal fluctuations might manifest as a fixed increase or decrease in sales each summer. In the multiplicative seasonality model, sales are calculated by multiplying the trend term by the seasonality term and then adding the error term: Sales = Trend × Seasonality + Error. For instance, sales of a product during holidays might be several times higher than usual; this multiple relationship is better described using the multiplicative seasonality model. The appropriate seasonality model should be chosen based on the actual performance of seasonal fluctuations in the data. If seasonal fluctuations show relatively stable increases or decreases, the additive model may be more suitable; if the multiple relationship changes significantly, the multiplicative model may be better.
[0128] Growth (growth pattern parameter):
[0129] Values: Flat (stable growth, i.e., no growth trend) and linear (linear growth).
[0130] Meaning and Impact: This defines the model's assumptions about the sales growth trend over time. Choosing "Flat" means the model assumes sales will not show a significant increase or decrease in the future, suitable for goods in relatively saturated markets with stable sales. For example, some daily necessities like salt have relatively stable market demand, making the Flat model more appropriate. The linear model, on the other hand, assumes sales grow or decrease linearly over time. For goods in their development stage, such as emerging technological products, sales may show a linear upward trend over time; choosing the linear model in this case better captures this growth trend.
[0131] Holidays_prior_scale (Holidays prior scale parameter):
[0132] Value: 5 (example value).
[0133] Meaning and Impact: This parameter controls the weighting of holiday influence on sales. A larger value means the model will place greater emphasis on the impact of holiday factors on sales, allocating more adjustment margin to holiday-related fluctuations during forecasting. For example, when the value is 5, if a product experiences a significant sales peak during holidays, the model will be more inclined to adjust the forecast based on holiday factors, making the predicted value more reflective of sales changes brought about by holidays. Conversely, a smaller value will cause the model to relatively weaken the impact of holidays.
[0134] Weekly_seasonality (weekly seasonality parameter, a periodicity parameter in weekly units):
[0135] Value: False (example value).
[0136] Meaning and Impact: This parameter determines whether to consider the impact of weekly seasonality on sales. If set to True, the model will attempt to capture recurring weekly sales patterns, such as the difference in sales between weekends and weekdays. For example, for some leisure and entertainment products, weekend sales may be significantly higher than weekday sales; setting it to True allows the model to better reflect this weekly seasonal variation. Setting it to False indicates that the model does not consider this weekly seasonality, suitable for products with insignificant weekly sales fluctuations, such as some office supplies, whose sales may not be affected by the distinction between weekends and weekdays. In this application, sales trend prediction is mainly achieved through a custom target calendar, requiring consideration of actual holidays, summer and winter vacations, and other social activities as corresponding parameters; therefore, this parameter can be set to disabled.
[0137] Specifically, considering that the weekday effect during winter and summer vacations differs from that during the school term, Prophet's built-in `weekly_seasonality` parameter is not used. Instead, a custom feature is used to allow Prophet to identify the weekday effect during and outside of winter and summer vacations. Two features can be used to identify whether a specific date falls within a winter / summer vacation or a school term, or two values of a single feature can be used for labeling. Prophet then adds corresponding seasonal parameters based on these two features, allowing it to identify the different weekday effects during winter / summer vacations and school terms. Dates falling on public holidays are not considered for the weekday effect, and their feature value is 0.
[0138] In a practical application scenario, such as time series forecasting, especially in education-related fields or businesses heavily influenced by school schedules, summer and winter vacations exhibit drastically different consumption or business activity patterns compared to weekdays during the school term. While the Prophet algorithm can handle time series features such as seasonality, its built-in `weekly_seasonality` parameter cannot effectively distinguish the differences in the effects of summer / winter vacations and weekdays during the school term. Therefore, a custom feature approach is needed to address this issue.
[0139] Feature construction:
[0140] First, two custom features need to be constructed to identify the state of a date. One feature indicates whether the date is during winter or summer vacation. For example, a feature named `is_summer_winter_vacation` can be defined, with a value of 1 if the date is during winter or summer vacation and 0 otherwise. The other feature indicates whether the date is in the middle of a semester, named `is_semester`. Correspondingly, the value is 1 if the date is in the middle of a semester and 0 if it is not. The construction of these two features requires data on the school's holiday and semester schedules. This can be determined by using a pre-set holiday schedule or by obtaining relevant information from relevant channels.
[0141] Data integration and Prophet applications:
[0142] The two custom features are integrated with existing time series data (such as sales data and business activity data) to ensure that each point in time corresponds to one of these two feature values. The integrated data is then input into the Prophet algorithm. When processing the data, the Prophet algorithm uses these two features to identify different seasonal patterns. For example, when the "is_summer_winter_vacation" feature value is 1, it learns the special effects of weekdays under these conditions (such as low consumption periods); when the "is_semester" feature value is 1, it learns the normal business patterns of weekdays during the semester (such as peak consumption periods). Meanwhile, for dates that fall within holidays, since their feature values have been set to 0 (whether it's "is_summer_winter_vacation" or "is_semester"), Prophet ignores the impact of these dates on the weekday effect and considers the impact of holidays on the data separately. This allows for a more accurate capture of changes in the weekday effect across different time periods, thereby improving the accuracy of time series forecasting.
[0143] By using this custom feature approach, the Prophet algorithm can better adapt to business scenarios with differences in summer and winter vacations and school semesters, improving the predictive model's ability to capture complex seasonality and weekday effects in the data, thereby generating more accurate prediction results.
[0144] Yearly_seasonality (annual seasonality parameter):
[0145] Values: True, False (example values).
[0146] Meaning and Impact: This setting determines whether the model considers the impact of annual seasonal factors on sales. If set to True, the model analyzes the impact of different seasons (spring, summer, autumn, winter) or specific holidays on sales within the annual cycle. For example, in the clothing industry, summer and winter clothing sales typically differ significantly; setting it to True allows the model to capture this annual seasonal variation, thus predicting sales more accurately. Setting it to False indicates that the model does not consider annual seasonality and is suitable for goods with relatively stable annual sales, such as basic necessities, whose sales may remain relatively stable throughout the year.
[0147] Changepoint_prior_scale (changepoint prior scale parameter):
[0148] Value: 1 (example value).
[0149] Meaning and Impact: This parameter controls the model's sensitivity to change points in the data (i.e., points where sales change significantly). A larger value makes the model more sensitive to change points and gives them more attention during prediction, thus adjusting the forecast to adapt to the changes following the change. For example, if a product experiences a promotional campaign that leads to a sudden and significant increase in sales, a larger `Changepoint_prior_scale` value allows the model to detect this change more quickly and consider its impact in subsequent predictions. Conversely, a smaller value makes the model less sensitive.
[0150] Changepoint_range (changepoint range parameter):
[0151] Values: 0.85, 0.9 (example values).
[0152] Meaning and Impact: This parameter limits the range within which the model searches for change points in the data. For example, values of 0.85 and 0.9 mean the model will only search for change points in the last 85% to 90% of the data. This setting aims to prevent the model from over-searching for change points in the earlier parts of the data, as the early data may be relatively unstable or influenced by initial conditions. By limiting the range of change points, the model can more effectively target the more stable parts of the data to find change points that influence predictions.
[0153] Holiday (custom parameter):
[0154] Values: True, False (example values).
[0155] Meaning and Impact: This parameter indicates whether product sales are affected by holidays; it is not a built-in parameter of Prophet. When this parameter is True, it passes a holiday calendar to Prophet's built-in holiday parameter; when it is False, the holiday parameter is not used.
[0156] Specifically, this parameter determines whether to pass holiday calendar information to Prophet's built-in parameters. When the value is True, the model uses the passed-in holiday calendar to consider the impact of holidays on sales, which is suitable for products whose sales are significantly affected by holidays, such as travel products and gifts. When the value is False, the model does not consider holiday factors, which is suitable for products whose sales are not significantly affected by holidays, such as some industrial raw materials, whose sales depend mainly on production demand rather than holidays.
[0157] For this parameter, this application uses hypothesis testing methods, such as t-statistic testing, to determine whether the sales volume of the same product in stores in the same region has changed significantly before and after the holiday based on the average sales volume before and after the holiday (assuming no significant change, if the p-value is less than the preset value, it means that the original hypothesis is true and no significant change has occurred).
[0158] Parameter tuning strategies and methods:
[0159] Grid search: This method allows for a comprehensive search across all possible combinations of the parameters. For example, there are two possible values for `Seasonality_mode`, two for `Growth`, and so on. All possible combinations of parameters are considered, and the Prophet algorithm is run on each one. The optimal parameter combination is then determined based on the evaluation metrics of the prediction results (such as mean squared error, mean absolute error, etc.). While this method is computationally intensive, it significantly increases the likelihood of finding the global optimum.
[0160] Random search: Unlike grid search, random search randomly selects a certain number of parameter combinations in the parameter space to run the Prophet algorithm. This method has relatively lower computational cost, but it may not guarantee finding the global optimum. However, in practical applications, since grid search is computationally expensive when the parameter space is large, random search is still a feasible method. It can quickly find a relatively optimal parameter region through random search, and then perform a more detailed grid search within that region.
[0161] Prior search: Based on past experience and understanding of the business scenario, initially determine the range or value of some parameters. For example, for products with a known relatively stable market, initially set Growth to Flat; for products significantly affected by holidays, set Holiday to True, etc. Then, combine other parameter tuning methods to further optimize the parameter combination.
[0162] This implementation method determines the target parameters and the corresponding search parameter range, thereby further determining the matching parameter search process to execute the parameter search. Based on the prediction effect of the second sales data, the final search results are determined as the parameter combination of the time series prediction model, thus obtaining the time series prediction model for performing the sales prediction task, ensuring the effectiveness of new store sales prediction.
[0163] As an optional implementation, the method further includes:
[0164] Obtain the building distribution information for the target area;
[0165] Based on the building distribution and the first sales data, the user distribution is determined;
[0166] Based on the user distribution, the fourth sales data of the target store is predicted using the time series prediction model.
[0167] The fourth sales data is used to indicate product attributes and corresponding user profiles.
[0168] This implementation method can determine the user distribution in the target area based on the building distribution and some attributes in the first sales data, such as product category and price data. Then, the user distribution is input into a time series prediction model to predict the product attribute distribution and user profile corresponding to the target area, providing a basis for product optimization and improving the practicality of sales forecasting.
[0169] Please refer to Figure 2. Figure 2 is a schematic diagram of an application scenario for a new store sales forecasting method provided in one embodiment of this application. As an optional implementation, after predicting the fourth sales data of the target store based on the user distribution and using the time series prediction model, the method further includes:
[0170] Based on the fourth sales data of the target store, determine the optimization plan;
[0171] The optimization scheme is used to indicate the parameter adjustment method for the brand, category, model, and placement of each product in the target store.
[0172] Figure 2 illustrates a specific implementation of an optimization scheme for a target store. The target store's fourth sales data is stored in the backend, and the displayed data can be an optimization scheme. Specifically, the optimization scheme includes product category classification, display category classification, specific sub-category classification, display location, current distance based on a reference point, optimized distance calculated by an optimization model based on the target store's fourth sales data and the reference point, the change from the current distance to the optimized distance, and a preset global planning distance. In the retail industry, the preset global planning distance can be a standardized distance for the entire chain provided by the target store's superior data analysis agency.
[0173] In the optimization plan, the optimized display distance and display location are the main variables. During the optimization process, to improve the optimization effect, one or more of the following conditions can be restricted: the total display distance under the major category remains unchanged; the total display distance on the side walls remains unchanged; and the total display distance on the central island remains unchanged. For specific sub-categories of products with preset global planned display distances, the global planned display distances can also be used as a reference optimization condition. In each iteration, the display location and optimized display distance for specific sub-categories can be determined based on the fourth sales data from the backend.
[0174] In fact, the meters in the optimization plan should be understood as a coordinate value. In actual stores, the scale can be planned or calculated in meters, so meters are used as the corresponding value.
[0175] Based on the fourth sales data obtained from the forecast, this implementation method can determine optimization plans related to the product display and inventory of the target store to increase sales, thereby improving the practicality of sales forecasting.
[0176] This application embodiment also provides a new store sales forecasting device, the device comprising:
[0177] The acquisition module is used to acquire the first sales data of the target area within a preset time period;
[0178] The first sales data includes sales data of multiple existing stores in the target area and sales data of the target store;
[0179] The processing module is used to preprocess the first sales data to obtain the second sales data;
[0180] The data preprocessing includes data cleaning and data splicing based on the time-series attributes of the corresponding sales data;
[0181] The processing module is also used to set the target parameters and corresponding parameter search range of the time series prediction model, and to train the time series prediction model based on the second sales data.
[0182] The processing module is also used to predict and obtain the third sales data of the target store through the time series prediction model.
[0183] The method provided in this embodiment obtains the first sales data of the target area within a preset time period, thereby obtaining the total sales situation since the opening of the new store and the sales situation of the existing stores within a certain period of time. This avoids the problem of insufficient new store data. While focusing on the new store data itself, it expands the data volume. The sales data of the existing stores and the target stores are used together as the basic data for sales forecasting, namely the first sales data. Further, based on the first sales data, data cleaning and data splicing preprocessing tasks are performed to obtain the second sales data. This allows the time series prediction model with preset search parameters to be trained, obtaining the time series prediction model under the optimal parameter combination, and then predicting the third sales data of the target store to indicate the expected sales situation of each category. Thus, by adding the sales data of the new store itself on the basis of historical store sales data, the front-end data volume and data effectiveness are improved, and automated new store sales forecasting is realized. Furthermore, the new store sales forecast can be used to optimize the products in the target store, which also improves the practicality of new store sales forecasting.
[0184] As an optional implementation, the specific method by which the processing module preprocesses the first sales data to obtain the second sales data includes:
[0185] Based on the time-series attributes of the first sales data, outlier data is identified and processed to obtain cleaned data;
[0186] Based on the temporal attributes of the cleaning data, data splicing is performed to obtain the second sales data.
[0187] This implementation method identifies outliers in the first sales data by utilizing its temporal attributes, such as holidays, weekdays (excluding holidays), and weekends (excluding holidays). It then sets sales-related attributes to null values or uses interpolation to correct the values, obtaining cleaned data. Based on the temporal attributes of the cleaned data, it concatenates the data in chronological order to obtain processed second sales data. This improves the effectiveness and interpretability of the second sales data, ensuring sufficient front-end data volume and validity, and enabling automated new store sales forecasting.
[0188] As an optional implementation, the specific method by which the processing module performs data splicing based on the temporal attributes of the cleaned data to obtain the second sales data includes:
[0189] Determine the first pieced data in the cleaning data that belongs to the target store;
[0190] Based on the temporal attributes of the first spliced data, determine the second spliced data belonging to the existing store in the cleaning data within the preset time period;
[0191] Wherein, the temporal attributes of the second spliced data precede the temporal attributes of the first spliced data;
[0192] The second and first spliced data are spliced together according to the time sequence attribute to obtain the second sales data.
[0193] In this embodiment, since the original first sales data includes sales data of multiple existing stores in the target area and sales data of the target store, the cleaned data can also include two parts: sales data of existing stores and sales data of the target store. First, all data belonging to the target store is determined. Then, according to a preset time period, the second concatenated data belonging to the existing store and whose time sequence is mutually exclusive with the time sequence of the first concatenated data corresponding to the target store is determined. Thus, data concatenation can be performed according to the time sequence. During the data processing, the sales data of the new store itself is added on the basis of the historical store sales data, which improves the amount of front-end data and data effectiveness, and realizes automated new store sales prediction.
[0194] As an optional implementation, the processing module is further configured to:
[0195] Determine the target calendar based on the temporal activity characteristics of the target area;
[0196] Based on the target calendar, determine the holiday marking parameters corresponding to the second sales data, and use the holiday marking parameters as the time classification attribute of the second sales data, merging them into the second sales data.
[0197] Since the original calendar built into the time series forecasting model may deviate from actual needs, this implementation method determines the target calendar based on the time activity characteristics of user groups or the region as a whole in the target area, and determines the holiday marking parameters of the second sales data as the time classification attribute of the second sales data. These parameters are then merged and added to the second sales data to complete the setting of the target calendar and further processing of the second sales data. This allows the second sales data to be associated with the time activity characteristics of the target area, increasing the effectiveness and interpretability of the second sales data, thereby improving the effect of new store sales forecasting.
[0198] As an optional implementation, the target parameters include:
[0199] Seasonal model parameters, growth model parameters, holiday prior scale parameters, weekly seasonality parameters, annual seasonality parameters, change point prior scale parameters, change point range parameters, and holiday marking parameters;
[0200] The processing module is also used to, after setting the target parameters and corresponding parameter search range of the time series prediction model...
[0201] Determine the parameter search process;
[0202] The parameter search process includes one or more of the following: grid search, prior search, and random search.
[0203] Based on the search results for the target parameters according to the parameter search process, the parameter combination of the time series prediction model is determined.
[0204] This implementation method determines the target parameters and the corresponding search parameter range, thereby further determining the matching parameter search process to execute the parameter search. Based on the prediction effect of the second sales data, the final search results are determined as the parameter combination of the time series prediction model, thus obtaining the time series prediction model for performing the sales prediction task, ensuring the effectiveness of new store sales prediction.
[0205] As an optional implementation, the processing module is further configured to:
[0206] Obtain the building distribution information for the target area;
[0207] Based on the building distribution and the first sales data, the user distribution is determined;
[0208] Based on the user distribution, the fourth sales data of the target store is predicted using the time series prediction model.
[0209] The fourth sales data is used to indicate product attributes and corresponding user profiles.
[0210] This implementation method can determine the user distribution in the target area based on the building distribution and some attributes in the first sales data, such as product category and price data. Then, the user distribution is input into a time series prediction model to predict the product attribute distribution and user profile corresponding to the target area, providing a basis for product optimization and improving the practicality of sales forecasting.
[0211] As an optional implementation, the processing module is further configured to, after predicting and obtaining the fourth sales data of the target store based on the user distribution and using the time series prediction model,
[0212] Based on the fourth sales data of the target store, determine the optimization plan;
[0213] The optimization scheme is used to indicate the parameter adjustment method for the brand, category, model, and placement of each product in the target store.
[0214] Based on the fourth sales data obtained from the forecast, this implementation method can determine optimization plans related to the product display and inventory of the target store to increase sales, thereby improving the practicality of sales forecasting.
[0215] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, a processing module can be a separate processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as program code in the device's memory, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0216] Schematic, as shown in FIG3, FIG3 is a schematic diagram of the internal structure of a computer device 300 provided in an embodiment of this application. The computer device 300 can be provided as a server. Referring to FIG3, the computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by memory 301 for storing instructions executable by the processing component 302, such as application programs. The application programs stored in memory 301 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 302 is configured to execute instructions to perform the text recognition method of any of the above embodiments.
[0217] The computer device 300 may also include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate on an operating system stored in memory 301, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or similar.
[0218] Those skilled in the art will understand that the structure shown in Figure 3 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.
[0219] This application provides a storage medium storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the method provided in any embodiment.
[0220] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0221] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
[0222] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting sales of new stores, characterized in that, The method includes: acquiring first sales data of a target area within a preset time period; wherein the first sales data includes sales data of multiple existing stores in the target area and sales data of the target store; performing data preprocessing on the first sales data to obtain second sales data; wherein the data preprocessing includes data cleaning and data splicing according to the time-series attributes of the corresponding sales data; setting target parameters and corresponding parameter search ranges for a time-series prediction model, and training the time-series prediction model based on the second sales data; predicting and obtaining third sales data of the target store through the time-series prediction model; wherein the step of performing data preprocessing on the first sales data to obtain second sales data includes: identifying outliers based on the time-series attributes of the first sales data and processing the outliers to obtain cleaned data; and performing data processing based on the time-series attributes of the cleaned data. The method involves splicing data to obtain second sales data. Specifically, the step of splicing data based on the temporal attributes of the cleaned data to obtain second sales data includes: determining first spliced data belonging to the target store within the cleaned data; determining second spliced data belonging to the existing store within a preset time period based on the temporal attributes of the first spliced data; wherein the temporal attributes of the second spliced data precede those of the first spliced data; performing data splicing on the second spliced data and the first spliced data according to the order of their temporal attributes to obtain second sales data; and further, the method includes: determining a target calendar based on the temporal activity characteristics of the target area; determining holiday marker parameters corresponding to the second sales data based on the target calendar, and merging the holiday marker parameters as the time classification attribute of the second sales data into the second sales data.
2. The method according to claim 1, characterized in that, The target parameters include: seasonality pattern parameters, growth pattern parameters, holiday prior scale parameters, weekly seasonality parameters, annual seasonality parameters, change point prior scale parameters, change point range parameters, and holiday marker parameters. After setting the target parameters of the time series forecasting model and the corresponding parameter search range, the method further includes: determining the parameter search process; wherein, the parameter search process includes one or more of grid search, prior search, and random search; and determining the parameter combination of the time series forecasting model based on the search results of the target parameters according to the parameter search process.
3. The method according to any one of claims 1-2, characterized in that, The method further includes: obtaining the building distribution of the target area; determining the user distribution based on the building distribution and the first sales data; and predicting the fourth sales data of the target store based on the user distribution using the time series prediction model; wherein the fourth sales data is used to indicate product attributes and corresponding user profiles.
4. The method according to claim 3, characterized in that, After predicting the fourth sales data of the target store based on the user distribution using the time series prediction model, the method further includes: determining an optimization scheme based on the fourth sales data of the target store; wherein the optimization scheme is used to indicate the parameter adjustment method for the brand, category, model, and placement of each product in the target store.
5. A new store sales forecasting device, used to implement the method as described in claim 1, characterized in that, The device includes: an acquisition module for acquiring first sales data of a target area within a preset time period; wherein the first sales data includes sales data of multiple existing stores in the target area and sales data of the target store; a processing module for preprocessing the first sales data to obtain second sales data; wherein the data preprocessing includes data cleaning and data splicing based on the time-series attributes of the corresponding sales data; the processing module is further configured to set target parameters and corresponding parameter search ranges for a time-series prediction model, and train the time-series prediction model based on the second sales data; the processing module is further configured to predict and obtain third sales data of the target store using the time-series prediction model; and the specific method by which the processing module preprocesses the first sales data to obtain the second sales data includes: identifying outlier data based on the time-series attributes of the first sales data and processing the outlier data to obtain cleaned data; According to the temporal attributes of the cleaned data, data splicing is performed to obtain second sales data; and the specific method by which the processing module performs data splicing to obtain second sales data according to the temporal attributes of the cleaned data includes: determining first spliced data belonging to the target store in the cleaned data; determining second spliced data belonging to the existing store in the cleaned data within the preset time period according to the temporal attributes of the first spliced data; wherein the temporal attributes of the second spliced data precede the temporal attributes of the first spliced data; performing data splicing on the second spliced data and the first spliced data in the order of the temporal attributes to obtain second sales data; and the processing module is further configured to: determine a target calendar according to the time activity characteristics of the target area; determine the holiday marking parameters corresponding to the second sales data according to the target calendar, and merge the holiday marking parameters as the time classification attributes of the second sales data into the second sales data.
6. A computer device, characterized in that, The method includes one or more processors and a memory storing computer-readable instructions that, when executed by the one or more processors, perform the steps of the method as described in any one of claims 1-4.
7. A storage medium, characterized in that, The storage medium stores computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Commodity sales prediction method and device, computer equipment and storage medium
CN111274531A
Sales prediction method and device, computing equipment and storage medium
CN118967214A