Enterprise business income prediction method based on multiple time sequence models
Through the enterprise operating income prediction method that integrates multiple time series models, the problem that traditional methods are difficult to fully consider multiple factors is solved, and higher prediction accuracy and robustness are achieved, making the prediction results more reliable.
Patent Information
- Application Number
- CN202510002102.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-03
AI Technical Summary
Traditional corporate operating income forecasting methods are difficult to fully consider various factors such as economic cycles, market changes and seasonal fluctuations, and a single model cannot fully utilize all the characteristics of the data, resulting in insufficient prediction accuracy and robustness.
The integrated method based on multiple time series models (such as XGBoost, SARIMAX, LSTM, etc.) is adopted to improve the accuracy and robustness of prediction through data preprocessing, feature engineering, model selection and training, hyperparameter optimization and model fusion prediction.
By integrating multiple time series models, the complementarity of the models is achieved, the accuracy and robustness of predictions are improved, the risks brought by errors of a single model are reduced, and the prediction results are more reliable.
Smart Images

Figure CN120088000A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data, and particularly relates to a method for predicting enterprise operating income based on multiple time series models. Background Art
[0002] With the increasing dependence of decision-makers on data-driven decision-making, accurately predicting the next stage of economic trends has become crucial for decision-makers' strategic planning. In the current big data era, decision-makers increasingly rely on data analysis to guide decision-making, and traditional empiricist methods can no longer meet the requirements of high accuracy. Therefore, modern machine learning and time series analysis technologies need to be adopted. Enterprise operating income belongs to time series data, so a time series prediction model needs to be constructed.
[0003] Enterprise operating income is affected by various internal and external factors, such as economic cycles, market changes, seasonal fluctuations, etc. Traditional model prediction methods often cannot comprehensively consider these influencing factors. Additionally, in practical applications, there will also be situations where the data is incomplete, and there may be missing values or outliers. If these data problems are not addressed, they may have a negative impact on the accuracy of the model. And different prediction models perform differently when dealing with specific types of data. For example, the XGBoost model is effective in dealing with structured data, while the LSTM model can effectively capture long-term dependencies in time series. Therefore, a single model may not be able to fully utilize all the characteristics of the data. Summary of the Invention
[0004] In view of the above problems, the purpose of the present invention is to provide a method for predicting enterprise operating income based on multiple time series models, which improves the accuracy and robustness of prediction by integrating the advantages of multiple time series models.
[0005] The present invention adopts the following technical solutions:
[0006] A method for predicting enterprise operating income based on multiple time series models includes the following steps:
[0007] Step S1, data collection: Collect historical data for predicting enterprise operating income, including historical input data, external economic indicators, and external influencing factors;
[0008] Step S2, data preprocessing: Preprocess the historical data and convert the data into a format that can be used by the model;
[0009] Step S3, feature engineering: Extract the time features and external factor features of the data;
[0010] Step S4, model selection and training: Select multiple time series models and train each model, and each model is processed in parallel;
[0011] Step S5, Model Evaluation and Optimization: Adjust the hyperparameters through grid search and cross-validation to optimize the performance of each model, select the best-performing parameter combination, and obtain the best model for each model;
[0012] Step S6, Model Fusion Prediction: Obtain the performance metrics of each best model, assign corresponding weights to each best model according to the performance metric results, and calculate the weighted average of the prediction values of the best models. The result obtained is used as the prediction result of the enterprise's operating income.
[0013] Further, the method for predicting the enterprise's operating income further includes the following steps:
[0014] Step S7, Result Display
[0015] Visually display the prediction results and generate a report for decision-making use.
[0016] Further, in step S2, data preprocessing includes missing value handling, outlier detection, and data cleaning.
[0017] Further, in step S4, the selected time series models are XGBoost model, SARIMAX model, Prophet model, LSTM model, and exponential smoothing model.
[0018] Further, the specific process in step S5 is as follows: For each model, set a parameter grid, traverse all possible parameter combinations, evaluate the performance of each combination through cross-validation, and select the best-performing parameter combination; Cross-validation is to divide the dataset into multiple subsets, evaluate the model performance on different training set and validation set partitions, and through K-fold cross-validation, ensure the consistency and robustness of the model under different data partitions, and finally obtain the best model.
[0019] Further, the specific process of step S6 is as follows:
[0020] S61, Obtain the performance metrics of each best model, and the performance metric is the mean squared error MSE:
[0021]
[0022] where n is the total number of samples, Y i is the true value of the i-th sample, is the predicted value of the i-th sample, is the square of the prediction error of the i-th sample;
[0023] S62, Calculate the corresponding weight w k assigned to the best model:
[0024]
[0025] Among them, MSE k is the mean squared error of the k-th best model, and m is the total number of best models, which is 5;
[0026] S63. For each time point, calculate the weighted average of the predicted values of each best model as the final predicted value of the enterprise's operating income:
[0027]
[0028] Among them, is the final predicted value after fusion, w k is the weight of the k-th best model, is the predicted value of the k-th best model.
[0029] The beneficial effects of the present invention are:
[0030] First of all, by integrating multiple time series models (such as XGBoost, SARIMAX, LSTM, etc.), the present invention realizes the complementarity of models. The superiority of each model under specific conditions can be compensated for by the deficiencies of other models, thereby improving the overall prediction accuracy and robustness. This integration method effectively reduces the risk brought by the error of a single model and makes the prediction result more reliable;
[0031] Secondly, in data collection, considering the influence of various economic indicators, external factors, etc., it has good flexibility, enabling the model to adapt to various business scenarios. Whether in the peak sales season or in the off-season, it can provide reasonable predictions;
[0032] Thirdly, through automated data preprocessing, feature engineering, model selection, and model evaluation optimization, etc., not only the efficiency is improved, but also it is ensured that the model can quickly respond in a rapidly changing data environment.
[0033] Through the prediction results provided by the present invention, a reliable basis is provided for the strategic decision-making of decision-makers. By analyzing and interpreting the prediction data, decision-makers can optimize resource allocation, adjust market strategies, improve operational efficiency and profit margins. This data-driven decision support system enables stronger adaptability and market competitiveness in the increasingly competitive market. Brief Description of the Drawings
[0034] Figure 1 is a flowchart of the method for predicting the enterprise's operating income based on multiple time series models provided by the embodiments of the present invention. Detailed Embodiments
[0035] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0036] Accurate operating revenue prediction can help decision-makers identify market trends and potential risks in advance, so as to formulate corresponding strategies. Through model selection, model fusion, etc., the present invention realizes the complementarity of models. The superiority of each model under specific conditions can be compensated by the deficiencies of other models, thereby overall improving the accuracy and robustness of prediction. The provided prediction results provide a reliable basis for the strategic decisions of decision-makers. By analyzing and interpreting the prediction data, decision-makers can optimize resource allocation, adjust market strategies, and improve operational efficiency and profit margins. Such a data-driven decision support system enables stronger adaptability and market competitiveness in an increasingly competitive market. In order to illustrate the technical solutions described in the present invention, specific embodiments will be used for illustration below.
[0037] As Figure 1 shown, the enterprise operating revenue prediction method based on multiple time series models provided in this embodiment includes the following steps:
[0038] Step S1, data collection: Collect historical data for enterprise operating revenue prediction, including historical input data, external economic indicators, and external influencing factors.
[0039] Historical data can be collected from the enterprise's financial system, sales system, market research, etc. Specific historical data generally includes historical revenue data, external economic indicators, and external influencing factors. The storage format of historical data is not limited, and different types of indicators can be stored and data analysis, processing, and modeling are supported. Usually, a structured table format is used, such as a CSV file, Excel spreadsheet, or relational database.
[0040] In this embodiment, historical data is saved monthly, including historical revenue data, external economic indicators, and external influencing factors, etc. As a specific example, in this embodiment, the historical data is in the following table format:
[0041]
[0042] The descriptions of each field are as follows:
[0043] Date: Recorded in months, in the format of YYYY-MM, which is convenient for time series modeling.
[0044] Enterprise Revenue: Records the monthly revenue data of the enterprise, which is the main target variable for prediction.
[0045] Cost of Sales: Reflects production costs and facilitates the analysis of the relationship between revenue and expenses.
[0046] Marketing Spend: Records the company's marketing expenditures and is used to analyze the impact of marketing activities on revenue.
[0047] CPI (Consumer Price Index): Represents the changes in the consumer price level and is an economic indicator used to evaluate the impact of inflation on a company's revenue.
[0048] GDP Growth Rate: Reflects the macroeconomic growth situation and is one of the external economic indicators.
[0049] Seasonal Factor: Used to record seasonal influences such as peak or off-peak seasons so that the model can capture seasonal variations.
[0050] External Events: Contains information on major events or holidays (such as the Spring Festival, promotional activities) and is used to analyze the impact of specific events on revenue.
[0051] Here, the company's revenue, cost of sales, and marketing spend are historical revenue data, the CPI and GDP growth rate are external economic indicators, and the seasonal factor and external events are external influencing factors. Among them, the date data is in string or timestamp format. The numerical data (company revenue, cost of sales, marketing spend, CPI, GDP growth rate, etc.) is in floating-point or integer format. The categorical data (seasonal factor, external events) is in string or categorical encoding format for the model to process non-numerical information. The data storage format is CSV or Excel, which can be conveniently displayed, saved, and viewed in a table and is suitable for medium and small datasets. In addition, a relational database (PostgreSQL) is used to store the company's data, and a NoSQL database (MongoDB) is used to save the externally flexible economic indicators and external influencing factors.
[0052] Step S2, Data Preprocessing: Preprocess the historical data and convert the data into a format that can be used by the model.
[0053] Data preprocessing generally includes handling missing values, detecting outliers, and data cleaning to remove outliers and duplicate values, fill in missing values, ensure the quality of the data, and convert the original data into a format that can be used by the model. The data preprocessing step is crucial, and good data quality is the basis for ensuring the accuracy of the model.
[0054] Step S3, Feature Engineering: Extract the time features and external factor features of the data.
[0055] Feature engineering is the process of transforming raw data into features that can better represent the actual problems processed by the prediction model and improve the accuracy of predicting unknown data. This step implements time feature extraction (such as months, quarters, etc.) and external factor features (such as economic indicators, market trends, etc.). By extracting appropriate features, the model performance can be significantly affected.
[0056] Step S4, Model Selection and Training: Select multiple time series models and train each model, with each model processed in parallel.
[0057] There are many types of time series models. In this embodiment, according to the actual application scenario and data characteristics, the following five time series models are reasonably selected, namely the XGBoost model, the SARIMAX model, the Prophet model, the LSTM model, and the exponential smoothing model.
[0058] XGBoost Model: XGBoost is a gradient boosting tree model and belongs to a type of ensemble learning. The XGBoost model performs well in processing structured data and regression and classification problems.
[0059] SARIMAX Model: SARIMAX is a time series prediction model, which is an extension of the ARIMA model and allows integrating the influence of exogenous factors. The SARIMAX model is suitable for time series data with obvious seasonality and trends. ARIMA (AutoRegressive Integrated Moving Average) is a classic time series prediction model suitable for stationary or differenced stationary data. It includes autoregressive (AR) and moving average (MA) components.
[0060] Prophet Model: Prophet is a time series prediction model developed by Facebook, especially suitable for data with strong seasonality and holiday effects. It has good tolerance for missing data.
[0061] LSTM Model: LSTM is a variant of the recurrent neural network (RNN) suitable for sequence data and can capture long-term dependencies.
[0062] Exponential Smoothing Model: The weighted average of past values from a fixed time series is used to predict future values of the series iteratively. The simplest model (simple exponential smoothing) calculates the next level value or smoothed value from the previous actual value and the previous level value. This method is called an exponential method because the value of each level is affected by the previous actual value, and the degree of influence decreases exponentially, that is, the newer the value, the greater the weight.
[0063] Each model is an independent logical block and can be processed in parallel as a parallel unit. Input feature data, and each model can execute in parallel to output prediction results. The extracted feature data is used to divide the dataset, and each model is trained.
[0064] Enterprise operating income is affected by various internal and external factors, such as economic cycles, market changes, seasonal fluctuations, etc. Traditional prediction methods often cannot comprehensively consider these influencing factors. This step uses five models, and the design takes into account the characteristics of different time series, including seasonality, trends, and the influence of external factors. This flexibility enables the model to adapt to various business scenarios and provide reasonable predictions whether in peak sales seasons or off-seasons.
[0065] Step S5, Model Evaluation and Optimization: Adjust the hyperparameters through grid search and cross-validation to optimize the performance of each model, select the best-performing parameter combination, and obtain the best model for each model.
[0066] After the model is trained, it needs to be further evaluated and optimized. After optimizing the parameters, train again. Through repeated training and optimization in steps S4 and S5, obtain the optimal hyperparameter combination. The model under this hyperparameter combination is the best model. The model can be optimized for specific features to ensure its adaptability and accuracy under different conditions.
[0067] In the model optimization stage, in this embodiment, the hyperparameters are adjusted through grid search and cross-validation to optimize the performance of each model. For each model, set a parameter grid, and systematically try multiple parameter combinations through grid search, traversing all possible parameter combinations. For each parameter combination, evaluate the performance of each combination through cross-validation and select the best-performing parameter combination. Cross-validation is mainly used in modeling applications, such as in PCR and PLS regression modeling. In a given modeling sample, take most of the samples for modeling, leave a small part of the samples to predict with the just-established model, and calculate the prediction error of this small part of the samples, and record their sum of squares.
[0068] Cross-validation divides the dataset into multiple subsets, evaluates the model performance on different train-test set partitions, and through K-fold cross-validation, ensures the consistency and robustness of the model under different data partitions, and finally obtains the best model.
[0069] For the specific content of adjusting hyperparameters for different models, it can be set according to the specific actual situation. For example, for the XGBoost model, the learning rate, tree depth, subsample ratio, etc. can be adjusted. The SARIMAX model can adjust the orders of autoregressive and moving average terms. The LSTM model can adjust the number of optimization layers, number of units, learning rate, etc.
[0070] The setting of hyperparameters has a direct impact on the performance of the model. Through systematic optimization, the optimal parameter combination can be found, thus significantly improving the model performance. This step uses optimization methods such as grid search and cross-validation. Grid search traverses the parameter space to ensure finding the best hyperparameter combination; while cross-validation evaluates the performance of the model on different datasets by splitting the training data multiple times. This process ensures the stability of the model, reduces the prediction errors caused by overfitting or underfitting, and thus effectively improves the accuracy and reliability of the prediction.
[0071] After all model training and optimization processes are completed, the best models of each model are obtained. The feature data is input into each best model, and each best model outputs the corresponding predicted values. Finally, the predicted values of each model are fused through the following step S6 to obtain the final predicted result of the operating income.
[0072] Step S6, Model Fusion Prediction: Obtain the performance metrics of each best model, assign corresponding weights to each best model according to the performance metric results, and calculate the weighted average of the predicted values of the best models. The result obtained is used as the predicted result of the enterprise's operating income.
[0073] In this step, by evaluating the performance of each model (XGBoost, SARIMAX, Prophet, LSTM, exponential smoothing model) on the validation set, weights are assigned to the models accordingly. The specific process of step S6 is as follows:
[0074] S61. Obtain the performance metrics of each best model to assign a weight to each model. Usually, it is based on the performance of the model on the validation set. For example, the mean squared error MSE or R 2 value. In this embodiment, the performance metric uses the mean squared error MSE, and the calculation method is as follows:
[0075]
[0076] where n is the total number of samples, Y i is the true value of the i-th sample, is the predicted value of the i-th sample, is the square of the prediction error of the i-th sample;
[0077] S62. Calculate the corresponding weight w k assigned to the best model:
[0078]
[0079] where MSE k is the mean squared error of the k-th best model, and m is the total number of best models, which is 5;
[0080] S63. For each time point, calculate the weighted average of the predicted values of each best model as the final predicted value of the enterprise's operating income:
[0081]
[0082] Where is the final predicted value of the fusion, w k is the weight of the k-th best model, is the predicted value of the k-th best model.
[0083] In this step, the weights are calculated based on the MSE values of each model. Take the reciprocal of the MSE value. The smaller the error, the larger the corresponding weight. This can ensure that the better the model performs, the higher its weight, and the sum of all weights is 1. For each time point, calculate the weighted average of the predicted values of each model as the final predicted value.
[0084] For example, assume that in one month, the predicted values of each model are [100, 120, 115, 105, 110], and the corresponding weights are [0.14, 0.23, 0.36, 0.18, 0.09]. Then:
[0085]
[0086] Different prediction models perform differently when dealing with specific types of data. For example, XGBoost is effective in dealing with structured data, while LSTM can effectively capture long-term dependencies in time series. Therefore, a single model may not be able to fully utilize all the characteristics of the data. In this step, by integrating multiple time series models, the complementarity of the models is achieved. The superiority of each model under specific conditions can be compensated for by the deficiencies of other models, thus overall improving the accuracy and robustness of the prediction. This integration method effectively reduces the risk brought by the error of a single model and makes the prediction results more reliable.
[0087] Step S7. Result display: Visually display the prediction results and generate a report for decision-making use.
[0088] After the prediction is generated, use a visualization tool to intuitively display the results for easy understanding and use by decision-makers. Generate the final prediction results and a visualization report.
[0089] In summary, through multi-model fusion, the present invention optimizes the model performance, fully utilizes the advantages of each model through the weight allocation method, enhances the model robustness, improves the model stability, and the deficiencies of a single model can be compensated by other models. Different models are sensitive to different features of the data, and after fusion, it is more representative. The present invention adopts automated design, which not only improves the efficiency but also ensures that the model can quickly respond in a rapidly changing data environment. In addition, the system can be easily scaled as the data volume increases, adapt to more complex prediction requirements, and reduce the operation and maintenance costs of enterprises.
[0090] The present invention closely follows the development trend of data-driven, and uses modern machine learning and time series analysis technologies to provide enterprises with in-depth data insights. This not only improves the scientificity of decision-making but also creates new business value for enterprises. Through continuous optimization and iteration, enterprises can continuously improve their operation levels and enhance their market competitiveness. The enterprise operating income prediction model of the present invention demonstrates significant advantages in multiple aspects such as multi-model integration, adaptability, prediction accuracy, missing value processing, automation and scalability, decision support, and compliance with the data-driven trend. These advantages provide strong data support for decision-makers and help them achieve sustainable development in a complex and changing market environment.
[0091] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for predicting business income based on multiple time series models, characterized in that: The enterprise operating income forecasting method comprises the following steps: Step S1, data collection: collecting historical data for enterprise operating income forecasting, including historical input data, external economic indicators and external influencing factors; Step S2, data preprocessing: preprocess the historical data and convert the data into a format that can be used by the model; Step S3, feature engineering: extracting the time characteristics and external factor characteristics of the data; Step S4, model selection and training: select multiple time series models and train each model, and process each model in parallel; Step S5, model evaluation and optimization: adjust hyperparameters through grid search and cross-validation to optimize the performance of each model, select the best performing parameter combination, and obtain the best model for each model; Step S6, model fusion prediction: obtain the performance indicators of each best model, assign corresponding weights to each best model according to the performance indicator results, and calculate the weighted average of the predicted values of the best model to obtain the result as the prediction result of the company's operating income.
2. The method for predicting business revenue based on multiple time series models as claimed in claim 1, characterized in that: The enterprise operating income forecasting method also includes the following steps: Step S7: Result display Visualize the forecast results and generate reports for decision-making.
3. The method for predicting business revenue based on multiple time series models as claimed in claim 2, characterized in that: In step S2, data preprocessing includes missing value processing, outlier detection, and data cleaning.
4. The method for predicting business revenue based on multiple time series models as claimed in claim 3, characterized in that: In step S4, the selected time series models include XGBoost model, SARIMAX model, Prophet model, LSTM model, and exponential smoothing model.
5. The method for predicting business revenue based on multiple time series models as claimed in claim 4, characterized in that: The specific process in step S5 is: for each model, set the parameter grid, traverse all possible parameter combinations, evaluate the performance of each combination through cross-validation, and select the best performing parameter combination; cross-validation is to divide the data set into multiple subsets, evaluate the model performance on different training sets and validation sets, and use K-fold cross-validation to ensure the consistency and robustness of the model under different data partitions, and finally obtain the best model.
6. The method for predicting business revenue based on multiple time series models as claimed in claim 5, characterized in that: The specific process of step S6 is: S61. Obtain the performance index of each optimal model, the performance index is mean square error MSE: Where n is the total number of samples, Y i is the true value of the i-th sample, is the predicted value of the ith sample, is the square of the prediction error of the i-th sample; S62, calculate the corresponding weight w assigned by the best model k : Among them, MSE k is the mean square error of the kth best model, and m is the total number of best models, which is 5; S63. For each time point, the weighted average of the prediction values of the best models is calculated as the final prediction value of the enterprise's operating income: in, is the final prediction value of the fusion, w k is the weight of the kth best model, is the predicted value of the kth best model.
Citation Information
Cited By
Semiconductor factory energy data prediction method based on fusion of multiple time sequence models
CN120952267A
A Semiconductor Factory Energy Data Prediction Method Based on the Fusion of Multiple Time Series Models
CN120952267B