Deep learning-based sales prediction method and system

Through deep learning-based methods, the characteristics of sales data and external environment data are collected and extracted, and a prediction model that integrates deep residual networks and long-term memory networks is constructed, which solves the problem of insufficient accuracy of traditional sales forecasting methods in complex market environments, and achieves more efficient and accurate sales forecasts.

CN120013590AInactive Publication Date: 2025-05-16BEIJING MINGYA INSURANCE BROKERS

Patent Information

Application Number
CN202510490918.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional sales forecasting methods have limitations when dealing with complex and changing market environments, and it is difficult to accurately capture complex factors such as macroeconomic fluctuations, seasonal changes, policy and regulatory adjustments, and market competition trends, resulting in a large deviation from the forecast results and actual sales situation.

Method used

A sales prediction method based on deep learning is adopted to collect historical sales data and associated external environment data, build a model training data set, and perform multi-dimensional feature extraction to generate a time series feature set and context feature set. Then, a prediction model based on the fusion of deep residual networks and long and short-term memory networks is constructed, and a phased training strategy and incremental learning mechanism are adopted to improve the adaptability and accuracy of the model.

Benefits of technology

It improves the accuracy and adaptability of sales forecasts, can better capture changes in complex market environments, reduce the deviation between the forecast results and actual sales situation, and enhances the generalization ability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013590A_ABST
    Figure CN120013590A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sales market prediction, and discloses a sales prediction method and system based on deep learning. The method comprises the steps of collecting historical sales data and external environment data to construct a training data set, and performing multi-dimensional feature extraction to generate a time sequence and a context feature set; constructing a prediction model fusing the deep residual network and the long-short-term memory network, wherein the prediction model comprises a convolution module and a circulation module; the model is trained in stages, initial parameters are optimized through unsupervised pre-training, and then supervised learning is combined with self-adaptive learning rate fine tuning; and predicting real-time input data by using the trained model, and if the confidence coefficient of a prediction result is low, triggering an incremental learning mechanism to update the model. The system comprises a data acquisition module, a feature extraction module, a model construction module and the like. Through multiple technical means, the accuracy, adaptability and reliability of sales prediction are improved, and powerful support is provided for enterprise sales decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sales market forecasting, and specifically to a sales forecasting method and system based on deep learning. Background Art

[0002] In today's highly competitive business environment, accurate sales forecasting is crucial to the survival and development of enterprises. It not only helps enterprises to rationally plan production, inventory and procurement plans and reduce operating costs, but also supports enterprises to formulate scientific market strategies and enhance market competitiveness. However, traditional sales forecasting methods have exposed many limitations when dealing with complex and changing market environments.

[0003] Early sales forecasts mainly relied on simple statistical analysis methods, such as the moving average method and exponential smoothing method. These methods are based on simple trend analysis of historical sales data, assuming that future sales will continue the past pattern, and ignore many complex factors that affect sales, such as macroeconomic fluctuations, seasonal changes, policy and regulatory adjustments, and market competition. For example, during periods of economic instability, consumers' purchasing power and willingness to buy will change significantly. It is difficult to accurately capture such changes by relying solely on simple statistics of historical sales data, resulting in a large deviation between the forecast results and actual sales.

[0004] With the development of data analysis technology, some prediction methods based on machine learning have gradually been applied, such as linear regression, decision trees, support vector machines, etc. These methods can handle multivariate data to a certain extent, and considering some factors that affect sales, they have made certain progress compared to traditional statistical methods. However, machine learning methods usually require manual feature engineering, have high requirements for domain knowledge, and are difficult to automatically mine complex nonlinear relationships in data. In actual sales scenarios, there are often complex nonlinear correlations between sales data and various influencing factors. For example, the launch of a new product may trigger a chain reaction in the market and change consumers' purchasing preferences for similar products. This complex relationship is difficult to accurately model using traditional machine learning methods.

[0005] In addition, traditional forecasting methods are not capable of handling dynamic changes in data. The market environment is constantly changing, new products, competitors, and consumer trends are constantly emerging, and sales data is also changing dynamically. Traditional methods have difficulty updating models in real time to adapt to these changes, causing the model's forecasting accuracy to decrease over time. For example, during e-commerce promotions, sales data will experience explosive growth and drastic fluctuations, and traditional models cannot be adjusted in time to accurately predict sales during such special periods.

[0006] The emergence of deep learning technology has brought new opportunities for sales forecasting. Deep learning has powerful automatic feature extraction and complex relationship modeling capabilities, and can automatically learn deep feature representations from massive data. However, directly applying deep learning models to sales forecasting also faces many challenges. Sales data usually has time series characteristics and is affected by a variety of external factors. How to effectively process these complex data and give full play to the advantages of deep learning models is an urgent problem to be solved. At the same time, deep learning models often require a large amount of data for training. When the data is sparse, the model is prone to overfitting or insufficient generalization. In addition, the model's training efficiency, prediction accuracy, and adaptability to new data also need to be further optimized and improved. Therefore, it is of great practical significance to develop an efficient, accurate, and adaptable sales forecasting method and system based on deep learning. Summary of the invention

[0007] The purpose of the present invention is to provide a sales forecasting method and system based on deep learning to solve the problems raised in the above background technology.

[0008] To achieve the above object, the present invention provides the following technical solution: a sales forecasting method based on deep learning, the method comprising: Collect historical sales data and related external environment data to build a model training data set; Performing multi-dimensional feature extraction on the model training data set to generate a time series feature set and a context feature set; Constructing a prediction model based on the fusion of a deep residual network and a long short-term memory network, wherein the prediction model includes a convolution module for extracting local spatiotemporal features and a recurrent module for capturing long-term dependencies; According to the time series feature set and the context feature set, the prediction model is trained in stages, wherein the first stage uses unsupervised pre-training to optimize initial parameters, and the second stage uses supervised learning combined with an adaptive learning rate adjustment strategy for fine-tuning; Receive real-time input of target sales scenario data and generate sales forecast results using the trained forecast model; If it is detected that the confidence of the current prediction result is lower than a preset threshold, the incremental learning mechanism is triggered to perform online parameter updates on the prediction model based on the data subset updated by the sliding window.

[0009] Preferably, the multi-dimensional feature extraction includes: Perform multi-scale sliding window segmentation on the original sales data to generate multiple time segments; The following operations are performed on each time segment: frequency domain features are extracted through discrete wavelet transform, one-dimensional sequence is converted into two-dimensional image features through Gram angular field coding, and key statistical features are filtered through attention weight allocation mechanism; Performing tensor splicing on the frequency domain features, two-dimensional image features and key statistical features to form the time series feature set; The calculation formula of the Gram angular field coding to convert a one-dimensional sequence into a two-dimensional image feature is:

[0010] in, They are respectively and The data vector at each time point, For time point and The angle between is the phase offset, is the generated two-dimensional Gram angular field matrix.

[0011] Preferably, the construction of the prediction model further includes: A dilated convolution layer and a channel attention mechanism are deployed in the convolution module, where the dilation rate of the dilated convolution layer is dynamically configured according to the data periodicity:

[0012] in, For the The dilation rate of the convolution kernel, is the data cycle length, is the maximum number of convolution kernels; The channel attention mechanism calculates the weights using the following formula:

[0013] in, For Channel The attention weight value, For Channel The global pooling feature, and is a learnable parameter, is the Sigmoid function.

[0014] Preferably, the staged training includes: In the unsupervised pre-training stage, a contrastive learning framework is used to construct positive and negative sample pairs, and the model parameters are optimized through the following loss function:

[0015] in, is the similarity score of the positive sample pair, For the The similarity score of negative samples is is the temperature coefficient, is the total number of negative sample pairs; In the supervised fine-tuning stage, the dynamic weight allocation strategy calculates the loss weight by the following formula:

[0016] in, For the The loss weight of the dimension feature, For the The importance score of the dimension feature, is the scaling factor, is the total number of feature dimensions.

[0017] Preferably, the method further comprises: Build a cross-domain knowledge transfer framework to extract potential correlation features from non-sales data related to the target sales scenario; Aligning the feature distribution of the non-sales data with the historical sales data through a domain adversarial training method; The aligned cross-domain features are injected into the prediction model as auxiliary input to enhance the generalization ability of the model in data-sparse scenarios.

[0018] Preferably, the incremental learning mechanism specifically includes: Keep the latest information based on the sliding window strategy The data of a time period is collected, and redundant historical data is eliminated based on the feature contribution evaluation results; The elastic weight solidification algorithm is used to record the importance index of the model parameters, and the parameter update is constrained by the following formula:

[0019] in, For parameters The Fisher information matrix value of is the penalty coefficient, is the smoothing constant, The loss function parameter The gradient of is the basic learning rate, For parameters The update amount; Perform adversarial sample generation operations on the newly added data.

[0020] Preferably, the method further comprises: Deploy a reinforcement learning agent module that generates a policy gradient based on the prediction results and feedback from real sales data; Dynamically adjusting hyperparameters in the prediction model based on the policy gradient, including a learning rate decay rate, a batch normalization coefficient, and a fusion weight of a feature pyramid; When a sudden change in external environment data is detected, the Monte Carlo tree search algorithm is enabled to generate a temporary parameter adjustment plan.

[0021] Preferably, the method further comprises: The probability calibration module is placed after the output layer of the prediction model, and the temperature scaling and quantile regression methods are used to calibrate the original prediction results; Construct an uncertainty quantification subnetwork and generate confidence intervals of prediction results through Monte Carlo Dropout sampling; If the confidence interval width exceeds the preset range, the manual review process is triggered and the abnormal case is recorded in the feedback knowledge base; The temperature scaling adjusts the predicted distribution by the following formula:

[0022] in, After calibration Class prediction probability, is the temperature parameter, is the total number of categories, The original output of the model Class logical value.

[0023] Preferably, the method further comprises: Establish a federated learning architecture that allows multiple parties to collaboratively train predictive models under data privacy isolation conditions; Homomorphic encryption technology is used to encrypt and transmit local model gradients, and secure multi-party computing is performed on the aggregation server. Design a differentiated contribution evaluation function for each participant, and dynamically allocate model usage rights based on contribution.

[0024] Preferably, the present invention also includes a sales forecasting system based on deep learning, the system comprising: Data collection module, used to collect historical sales data and related external environment data to build model training data sets; A feature extraction module performs multi-dimensional feature extraction on the model training data set to generate a time series feature set and a context feature set; A model building module, which builds a prediction model based on the fusion of a deep residual network and a long short-term memory network, wherein the prediction model includes a convolution module for extracting local spatiotemporal features and a loop module for capturing long-term dependencies; A model training module, which trains the prediction model in stages according to the time series feature set and the context feature set, wherein the first stage uses unsupervised pre-training to optimize initial parameters, and the second stage uses supervised learning combined with an adaptive learning rate adjustment strategy for fine-tuning; The prediction execution module receives the target sales scenario data input in real time and generates sales prediction results using the trained prediction model; The model update module triggers the incremental learning mechanism if it detects that the confidence of the current prediction result is lower than a preset threshold, and performs online parameter updates on the prediction model based on the data subset updated by the sliding window.

[0025] Compared with the prior art, the present invention has the following beneficial effects: In terms of data processing and feature extraction, the model training data set is constructed by collecting historical sales data and related external environment data, and multi-dimensional feature extraction is performed to generate time series feature sets and context feature sets. The original sales data is segmented into multi-scale sliding windows, combined with discrete wavelet transform, Gram angular field coding and attention weight allocation mechanism, which can mine potential information in the data from different angles and different granularities. Discrete wavelet transform extracts frequency domain features to help capture the fluctuation characteristics of sales data at different frequencies, such as the frequency characteristics of seasonal sales fluctuations; Gram angular field coding converts one-dimensional sales data into two-dimensional image features, which makes it possible to mine features using image processing technology in the future; the attention weight allocation mechanism selects key statistical features according to the importance of the features to sales forecasts, avoiding the interference of redundant information, making the features of the model input more accurate and effective, thus laying a solid foundation for improving forecast accuracy.

[0026] In terms of model construction and training, the prediction model is constructed based on the fusion of deep residual network and long short-term memory network, and the convolution module and the loop module cooperate with each other. The dilated convolution layer and the channel attention mechanism are deployed in the convolution module. The dilated convolution layer dynamically configures the expansion rate according to the periodicity of the data, which can expand the receptive field of the convolution kernel without increasing too much calculation and effectively capture local spatiotemporal features; the channel attention mechanism calculates the channel attention weight, allowing the model to pay more attention to the channel features that contribute significantly to sales forecasting, thereby improving model performance. The staged training strategy further optimizes the model. In the unsupervised pre-training stage, the contrastive learning framework is used to construct positive and negative sample pairs to optimize the initial parameters, so that the model can initially learn the potential features and structure of the data; in the supervised fine-tuning stage, the adaptive learning rate adjustment strategy and the dynamic weight allocation strategy are combined to reasonably allocate loss weights according to the importance of features, accelerate model convergence, and improve prediction accuracy.

[0027] In order to solve the problems of data sparsity and model generalization, a cross-domain knowledge transfer framework is constructed. Potential correlation features are extracted from non-sales data related to the target sales scenario, and the feature distribution of non-sales data is aligned with historical sales data through domain adversarial training method. The aligned cross-domain features are then injected into the prediction model as auxiliary input, which enhances the generalization ability of the model in data-sparse scenarios and enables the model to better adapt to various complex sales scenarios.

[0028] In terms of model updating and adaptability, the incremental learning mechanism plays an important role. When the confidence of the prediction result is lower than the preset threshold, the latest data is retained and redundant historical data is eliminated based on the sliding window strategy to ensure that the data used by the model is always timely; the elastic weight solidification algorithm is used to record the importance indicators of the model parameters, and the parameter updates are reasonably constrained to avoid the model forgetting important information during the update process; adversarial sample generation operations are performed on the newly added data to enhance the robustness of the model against adversarial attacks, so that the model can adapt to the changes in new data in a timely manner and continuously improve the accuracy and reliability of predictions.

[0029] In addition, the reinforcement learning agent module is deployed to dynamically adjust the hyperparameters in the prediction model according to the feedback of the prediction results and the actual sales data, including the learning rate decay rate, the batch normalization coefficient, and the fusion weight of the feature pyramid, etc., to further optimize the model performance; the probability calibration module and the uncertainty quantification subnetwork are placed after the output layer of the prediction model to calibrate the original prediction results and generate a confidence interval. If the width of the confidence interval exceeds the preset range, the manual review process is triggered and the abnormal case is recorded to the feedback knowledge base, which improves the credibility and reliability of the prediction results. A federated learning architecture is established to allow multiple participants to collaboratively train the prediction model under data privacy isolation conditions, and homomorphic encryption technology and secure multi-party computing are used to ensure data security. The model usage rights are dynamically allocated based on the contribution degree, which not only protects data privacy, but also promotes multi-party cooperation, and improves the overall model training effect and prediction accuracy. In summary, the present invention innovates and optimizes the sales forecasting method and system from multiple links, comprehensively improves the accuracy, adaptability and reliability of sales forecasts, provides strong support for the sales decision-making of enterprises, and has significant economic benefits and application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a working principle diagram of the sales forecasting method based on deep learning described in the present invention; Figure 2 Flowchart constructed for the prediction model; Figure 3 A step diagram for cross-domain knowledge transfer; Figure 4 Flowchart for calibration and evaluation of prediction results. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0032] See also Figure 1-4 The present invention provides a technical solution: a sales forecasting method based on deep learning, and the specific implementation scheme is as follows: Data collection and data set construction: Through various data collection channels, such as internal sales databases of enterprises, data from market research institutions, and public industry data resources, historical sales data are collected. These data cover key information such as product sales quantity, sales amount, and sales time. At the same time, external environmental data related to sales are collected, such as macroeconomic indicators (such as GDP growth rate, inflation rate), seasonal information, regional policies and regulations, etc. The collected historical sales data and external environmental data are sorted and cleaned, and data with duplicates, errors, or serious missing values ​​is removed to build a model training data set to provide a reliable data foundation for subsequent model training.

[0033] Multi-dimensional feature extraction: Perform multi-dimensional feature extraction on the constructed model training data set to mine potential valuable information in the data. This process generates a time series feature set and a context feature set. The time series feature set focuses on capturing the law of sales data changing over time, while the context feature set contains more relevant information such as external environment data, providing rich input features for the model from different angles.

[0034] Prediction model construction: Construct a prediction model based on the fusion of deep residual network and long short-term memory network. The prediction model contains two key modules, namely convolution module and loop module. The convolution module is used to extract local spatiotemporal features and can effectively capture the feature information of data in local time and space; the loop module is used to capture long-term dependencies and can remember data associations over a longer time span, so that the model can better understand the long-term trends and patterns of sales data.

[0035] Training the prediction model in stages: Based on the generated time series feature set and context feature set, the prediction model is trained in stages. In the first stage, unsupervised pre-training is used to optimize the initial parameters of the model, allowing the model to initially learn some potential features and patterns in the data. In the second stage, supervised learning combined with an adaptive learning rate adjustment strategy is used for fine-tuning. By comparing with known real sales data, the model parameters are further optimized, and the learning rate is adaptively adjusted to make the model converge faster and more stably during the training process, thereby improving the accuracy of the model prediction.

[0036] Sales forecast execution: After the model training is completed, the system receives real-time input of target sales scenario data, which may include current market environment information, recent sales data trends, etc. The trained forecast model is used to analyze and process these real-time data to generate corresponding sales forecast results, providing important reference for the company's sales decision-making.

[0037] Model update: During the prediction process, the system will detect the confidence of the current prediction result in real time. If the confidence of the current prediction result is detected to be lower than the preset threshold, it means that the prediction accuracy of the model may be affected, and the incremental learning mechanism is triggered. The prediction model is updated online based on the data subset updated by the sliding window, so that the model can adapt to new data changes in a timely manner and continuously improve the accuracy and reliability of the prediction.

[0038] The present invention will be further described below in conjunction with Examples 1 to 6:

[0039] Embodiment 1: When extracting multi-dimensional features, we first perform multi-scale sliding window segmentation on the original sales data. The purpose of setting a multi-scale sliding window is to observe the changing trend of sales data from different time granularities. For example, setting a smaller sliding window, such as a window with a day as the unit, can capture the short-term fluctuation of sales data; setting a larger sliding window, such as a window with a month as the unit, can grasp the long-term trend of sales data. Through this multi-scale segmentation method, multiple data of different time segments are generated.

[0040] For each generated time segment, a series of operations are performed to extract features. Frequency domain features are extracted through discrete wavelet transform. Discrete wavelet transform is a signal processing technique that can decompose time series data into subsequences of different frequencies, thereby revealing the changing characteristics of data at different frequencies. For example, some seasonal sales fluctuations may be obvious in a specific frequency segment, and these features can be extracted through discrete wavelet transform.

[0041] Next, the one-dimensional sequence is converted into two-dimensional image features through Gram angular field coding. The calculation formula is: In this formula, They are respectively and Data vectors at different time points, which represent the specific values ​​of sales data at different times; For time point and The angle between them reflects the relative relationship between the data vectors at two time points; is the phase offset, which can adjust the phase of the generated two-dimensional image features and increase the diversity of the features; The generated two-dimensional Gram angular field matrix converts the one-dimensional sales data sequence into a two-dimensional image form, which is convenient for subsequent feature extraction and analysis using image processing methods.

[0042] Key statistical features must also be screened through the attention weight allocation mechanism. The attention weight allocation mechanism assigns corresponding weights to each feature based on its importance to sales forecasting. For example, in some sales scenarios, recent sales data is more important for predicting future sales, so the features corresponding to this part of the data will be assigned higher weights, while some data features with a long history and little relevance to current sales trends will be assigned lower weights.

[0043] Finally, the extracted frequency domain features, two-dimensional image features, and key statistical features are tensor-concatenated to form a time series feature set. The tensor concatenation operation combines different types of features in a certain dimensional order, so that these features can work together in subsequent model training, providing the model with more comprehensive information, thereby improving the model's predictive ability.

[0044] Embodiment 2: When building a prediction model, the design of the convolution module is crucial. The dilated convolution layer and channel attention mechanism are deployed in the convolution module. The expansion rate of the dilated convolution layer is dynamically configured according to the periodicity of the data, and the formula is: .here, For the The expansion rate of the convolution kernel, which determines the sampling interval of the convolution kernel during the convolution operation; The length of the data cycle. For example, if the sales data shows obvious seasonality, then the length of this seasonal cycle is the length of the data cycle. For example, the peak sales season for some commodities is in certain months of each year, and the cycle formed by these months is the data cycle. It is the maximum number of convolution kernels, which limits the upper limit of the number of convolution kernels in the convolution module and should be set reasonably according to factors such as model complexity and computing resources.

[0045] By dynamically configuring the dilation rate, the atrous convolution layer can expand the receptive field of the convolution kernel and capture a wider range of spatiotemporal information without increasing the number of parameters and computation. For example, when analyzing sales data, different time periods have different effects on sales, and dynamically adjusting the dilation rate can make the convolution layer better adapt to these different periodic characteristics.

[0046] The channel attention mechanism calculates the weights using the following formula: .in, For Channel The attention weight value, which indicates the importance of the channel in the entire feature representation; For Channel The global pooling feature of the channel is compressed through the global pooling operation to obtain a value representing the overall feature of the channel. and They are learnable parameters that are continuously optimized during model training to adjust the way channel attention weights are calculated; is a Sigmoid function, which maps the calculation result to between 0 and 1 as the final attention weight value.

[0047] The role of the channel attention mechanism is to weight the features according to the importance of each channel, so that the model can pay more attention to the channel features that contribute significantly to sales forecasting and suppress the less important channel information, thereby improving the performance of the model and prediction accuracy.

[0048] Embodiment 3: When training the prediction model in stages, the unsupervised pre-training stage uses a contrastive learning framework to construct positive and negative sample pairs. The contrastive learning framework aims to allow the model to learn the similarities and differences between data. The model parameters are optimized using the following loss function: In this formula, The similarity score of the positive sample pairs is usually selected from similar sales scenarios or data with similar trends. This score indicates the degree of similarity between them. For the The similarity scores of negative sample pairs are selected from sales scenarios or data with large differences, so that the model can learn the differences between different data. The temperature coefficient controls the degree of smoothness in the calculation of the similarity score. A smaller temperature coefficient will make the model pay more attention to distinguishing the difference between positive and negative samples, while a larger temperature coefficient will make the calculation result smoother. is the total number of negative sample pairs, which reflects the number of negative samples used in the contrastive learning process.

[0049] In this way, during the unsupervised pre-training stage, the model can learn some common features and structures of the data, laying a good foundation for subsequent supervised learning.

[0050] In the supervised fine-tuning stage, a dynamic weight allocation strategy is adopted. The dynamic weight allocation strategy calculates the loss weight through the following formula: .in, For the The loss weight of the dimension feature, which determines the contribution of the dimension feature to the overall loss when calculating the loss function; For the Importance score of dimension features, which can be determined by various methods such as feature correlation analysis and the degree of influence of features in historical sales data; is a scaling factor, which is used to adjust the influence of the importance score on the loss weight. By adjusting the scaling factor, the model can pay more attention to important features or balance the influence of different features; is the total number of feature dimensions, which indicates the total number of dimensions of the model input features.

[0051] Through the dynamic weight allocation strategy, the model can reasonably allocate loss weights according to the importance of different features, thereby more effectively optimizing model parameters during fine-tuning and improving the model's prediction accuracy.

[0052] Embodiment 4: In order to further improve the generalization ability of the model in data sparse scenarios, this embodiment introduces a cross-domain knowledge transfer framework. In actual sales forecasting scenarios, non-sales data related to the target sales scenario contains a large amount of potentially valuable information. By mining this information, more dimensional references can be provided for the sales forecasting model.

[0053] The first is the extraction of potential correlation features. Taking a company engaged in fashion clothing sales as an example, non-sales data may include fashion trend reports, data on the popularity of fashion topics on social media, and popular culture dynamics. Text mining technology is used to analyze fashion trend reports, extract keywords such as popular clothing styles, colors, and materials of the season, and convert them into numerical features. For the popularity data of fashion topics on social media, the discussion volume, number of likes, number of shares, and other data of related topics are obtained through the data interface, and statistics and normalization are performed according to a certain time period to obtain features that reflect the changes in the popularity of fashion topics over time. Popular culture dynamics can be analyzed by analyzing clothing elements that appear in movies, TV series, music, etc., combined with their dissemination range and influence, to generate corresponding feature vectors. These features extracted from non-sales data may have potential correlations with clothing sales data. For example, the surge in popularity of a certain popular style on social media may indicate that its sales will increase in the future.

[0054] Then, domain adversarial training is performed. The purpose of using the domain adversarial training method is to make the feature distribution of non-sales data match that of historical sales data. A domain adversarial network is constructed, which includes a generator and a discriminator. The function of the generator is to transform the features of non-sales data so that they are as close as possible to the feature distribution of historical sales data. The discriminator is responsible for distinguishing whether the input features come from historical sales data or non-sales data converted by the generator. During the training process, the generator and the discriminator are constantly in confrontation. The generator adjusts its own parameters to try to make the converted features deceive the discriminator; the discriminator is constantly optimized to improve its ability to identify true and false features. Through such adversarial training, the feature distribution of non-sales data gradually becomes consistent with the historical sales data.

[0055] Finally, the aligned cross-domain features are injected into the prediction model. The non-sales data features processed by domain adversarial training are added as auxiliary inputs to the previously constructed prediction model based on the fusion of deep residual network and long short-term memory network. In the model training stage, these cross-domain features are trained together with the original historical sales data and external environment data features to help the model learn richer patterns and rules. In the prediction stage, the model can comprehensively consider various information, so that it can still make relatively accurate sales forecasts when the data is sparse. For example, in the off-season for clothing sales, when historical sales data is relatively small, cross-domain features such as popular culture trends can provide additional information for the model, enabling the model to better grasp market demand and improve the accuracy of predictions.

[0056] Embodiment 5:

[0057] When the confidence level of the results output by the prediction model is lower than the preset threshold, the system will automatically trigger the incremental learning mechanism to optimize the model performance so that it can better adapt to the changing data.

[0058] Data processing based on sliding window strategy. The size of the sliding window is set to retain the latest N time periods of data. Assume that the sales data of a supermarket is processed in weekly time periods, and N is set to 4. This means that the model will always retain the sales data of the last 4 weeks. As time goes by, new sales data will enter the window every week, and the data of the earliest week will be removed. The advantage of this is that the model can obtain the latest sales trends in a timely manner. For example, during holidays, the sales of goods will fluctuate significantly. By retaining the latest data through sliding windows, the model can quickly capture this change. While retaining new data, it is also necessary to eliminate redundant historical data based on the feature contribution evaluation results. The feature contribution is evaluated by calculating the degree of influence of each feature on the model prediction results. For example, the feature importance evaluation method in the random forest algorithm is used to score the importance of features in the sales data, such as product categories, sales time, promotional activities, etc. The historical data corresponding to the features with low scores are regarded as redundant data and deleted. This can reduce the amount of calculation for model training, avoid the interference of redundant information on the model, and improve the training efficiency and prediction accuracy of the model.

[0059] The elastic weight solidification algorithm is used to record the importance index of the model parameters, and the parameter update is constrained by a specific formula. The formula is: .in, Representative parameters The Fisher information matrix value of , which reflects the importance of the parameter in the model. The larger the Fisher information matrix value, the greater the influence of the parameter on the model output results; is the penalty coefficient, which controls the protection of important parameters. A larger penalty coefficient will make the model more cautious about important parameters when updating parameters. is a smoothing constant, which is used to avoid the situation where the denominator is zero and ensure the stability of the formula calculation; Represents the loss function for the parameter The gradient of , which determines the direction and magnitude of parameter updates; is the basic learning rate, which controls the step size of parameter update; The parameter In practical applications, these parameters can be adjusted reasonably according to the model training situation and data characteristics. For example, in the early stage of model training, the basic learning rate can be appropriately increased to speed up the convergence of the model; as the training progresses, the basic learning rate can be gradually reduced to make the model more stable.

[0060] Perform adversarial sample generation on the new data. Adversarial samples are samples generated by making small perturbations to the original data, which can cause the model to produce incorrect predictions. By generating adversarial samples, the robustness of the model to adversarial attacks can be enhanced. The Fast Gradient Sign Method (FGSM) is used to generate adversarial samples. For the new input data, the gradient of the model loss function with respect to the input data is calculated, and then the data is perturbed according to the direction of the gradient. For example, for an input data vector representing the number of goods sold, a small perturbation value is added to each dimension according to the calculated gradient to generate adversarial samples. These adversarial samples are input into the model together with the original new data for training, so that the model can learn how to deal with adversarial attacks, thereby improving the robustness and prediction accuracy of the model.

[0061] Embodiment 6: This embodiment will further optimize the sales forecasting model from multiple aspects to improve the performance and adaptability of the system.

[0062] Deploy the reinforcement learning agent module. The reinforcement learning agent module generates a policy gradient based on the feedback of the prediction results and the actual sales data. Take the sales forecast of a certain e-commerce platform as an example. During a period of time, the model predicts that the sales volume of a certain mobile phone is X, while the actual sales volume is Y. The agent module evaluates the accuracy of the prediction based on the difference between the two (Y - X). If the predicted sales volume is lower than the actual sales volume, it means that the model may underestimate the market demand, and the agent module will give a positive feedback signal; otherwise, it will give a negative feedback signal. By continuously accumulating these feedback information, the agent module calculates the policy gradient. Based on this policy gradient, the agent module dynamically adjusts the hyperparameters in the prediction model. For example, the learning rate decay rate. When the model prediction results are continuously inaccurate, the learning rate decay rate can be appropriately accelerated to enable the model to adapt to new data changes more quickly; for the batch normalization coefficient, it is adjusted according to the distribution of the data to ensure that the distribution of the input data of the model is relatively stable during the training process; the fusion weight of the feature pyramid will also be optimized according to the feedback, so that the model can better integrate feature information at different levels and improve the accuracy of the prediction.

[0063] When a mutation in the external environment data is detected, the Monte Carlo tree search algorithm is enabled to generate a temporary parameter adjustment plan. Suppose a region suddenly introduces a new consumer subsidy policy, which is a mutation in the external environment data. The Monte Carlo tree search algorithm uses the current model state as the root node and simulates different parameter adjustment schemes to predict future sales. Each simulation selects an action (i.e., parameter adjustment plan) according to a certain strategy, then executes this action and observes the results. After a large number of simulations, the optimal parameter adjustment plan is selected and applied to the model. For example, the algorithm may try to adjust the weights of different layers in the model, change the learning rate, etc. By simulating the sales forecast results under different schemes, the scheme that makes the forecast results closest to the actual situation is selected, so that the model can quickly adapt to changes in the external environment and maintain a high prediction accuracy.

[0064] The probability calibration module is placed after the output layer of the prediction model, and the temperature scaling and quantile regression method are used to calibrate the original prediction results. The temperature scaling is calculated by the formula Adjust the forecast distribution. After calibration Class prediction probability, for example, when predicting the sales share of different brands of mobile phones, it represents the predicted probability of the sales share of a certain brand of mobile phones; is the temperature parameter, which is used to control the smoothness of the predicted distribution. A smaller temperature parameter will make the predicted distribution sharper, and vice versa. is the total number of categories, that is, the number of predicted categories, such as the number of mobile phone brands in the above example; The original output of the model Logical value. Temperature scaling can make the model's predictions more reasonable.

[0065] Construct an uncertainty quantification subnetwork and generate confidence intervals for prediction results through Monte Carlo Dropout sampling. During the model training process, Dropout is used for random inactivation multiple times, and different prediction results are obtained each time. Based on these results, the mean and standard deviation of the predicted values ​​are calculated to obtain the confidence interval of the prediction results. If the width of the confidence interval exceeds the preset range, it means that the model has a large uncertainty in the prediction results. At this time, the manual review process is triggered and the abnormal cases are recorded in the feedback knowledge base. Manual reviewers can analyze these abnormal cases to find out the possible causes of uncertainty, such as data anomalies, unreasonable model parameters, etc., and feedback the relevant information to the model optimizer to further improve the model.

[0066] Establish a federated learning architecture to allow multiple participants to collaboratively train prediction models under data privacy isolation conditions. Take multiple retailers jointly conducting sales forecasting as an example. Each retailer has its own sales data, but due to data privacy issues, the data cannot be directly shared. With the federated learning architecture, each participant uses its own data to train the model locally, and then encrypts the gradient information of the model. Homomorphic encryption technology is used to encrypt and transmit the local model gradient to ensure the security of data during transmission. Secure multi-party computation is performed on the aggregation server to aggregate the gradient information of each participant and update the global model. Differentiated contribution evaluation functions are designed for each participant to evaluate their contribution to the global model based on factors such as the data quality, data volume, and model training effect provided by the participant. Model usage rights are dynamically allocated based on contribution. For example, participants with high contribution can obtain more advanced model usage rights, such as using the model for prediction more frequently and obtaining more detailed prediction analysis reports. This can motivate each participant to actively participate in federated learning and improve the overall model training effect and prediction accuracy.

[0067] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0068] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A sales forecasting method based on deep learning, characterized in that: include: Collect historical sales data and related external environment data to build a model training data set; Performing multi-dimensional feature extraction on the model training data set to generate a time series feature set and a context feature set; Constructing a prediction model based on the fusion of a deep residual network and a long short-term memory network, wherein the prediction model includes a convolution module for extracting local spatiotemporal features and a recurrent module for capturing long-term dependencies; According to the time series feature set and the context feature set, the prediction model is trained in stages, wherein the first stage uses unsupervised pre-training to optimize initial parameters, and the second stage uses supervised learning combined with an adaptive learning rate adjustment strategy for fine-tuning; Receive real-time input of target sales scenario data and generate sales forecast results using the trained forecast model; If it is detected that the confidence of the current prediction result is lower than a preset threshold, the incremental learning mechanism is triggered to perform online parameter updates on the prediction model based on the data subset updated by the sliding window.

2. The sales forecasting method according to claim 1, characterized in that: The multi-dimensional feature extraction includes: Perform multi-scale sliding window segmentation on the original sales data to generate multiple time segments; The following operations are performed on each time segment: frequency domain features are extracted through discrete wavelet transform, one-dimensional sequence is converted into two-dimensional image features through Gram angular field coding, and key statistical features are filtered through attention weight allocation mechanism; Performing tensor splicing on the frequency domain features, two-dimensional image features and key statistical features to form the time series feature set; The calculation formula of the Gram angular field coding to convert a one-dimensional sequence into a two-dimensional image feature is: ; in, They are respectively and The data vector at each time point, For time point and The angle between is the phase offset, is the generated two-dimensional Gram angular field matrix.

3. The sales forecasting method according to claim 1, characterized in that: The construction of the prediction model further includes: A dilated convolution layer and a channel attention mechanism are deployed in the convolution module, where the dilation rate of the dilated convolution layer is dynamically configured according to the data periodicity: ; in, For the The dilation rate of the convolution kernel, is the data cycle length, is the maximum number of convolution kernels; The channel attention mechanism calculates the weights using the following formula: ; in, For Channel The attention weight value, For Channel The global pooling feature, and is a learnable parameter, is the Sigmoid function.

4. The sales forecasting method according to claim 1, characterized in that: The phased training includes: In the unsupervised pre-training stage, a contrastive learning framework is used to construct positive and negative sample pairs, and the model parameters are optimized through the following loss function: ; in, is the similarity score of the positive sample pair, For the The similarity score of negative samples is is the temperature coefficient, is the total number of negative sample pairs; In the supervised fine-tuning stage, the dynamic weight allocation strategy calculates the loss weight by the following formula: ; in, For the The loss weight of the dimension feature, For the The importance score of the dimension feature, is the scaling factor, is the total number of feature dimensions.

5. The sales forecasting method according to claim 1, characterized in that: Also includes: Build a cross-domain knowledge transfer framework to extract potential correlation features from non-sales data related to the target sales scenario; Aligning the feature distribution of the non-sales data with the historical sales data through a domain adversarial training method; The aligned cross-domain features are injected into the prediction model as auxiliary input to enhance the generalization ability of the model in data-sparse scenarios.

6. The sales forecasting method according to claim 1, characterized in that: The incremental learning mechanism specifically includes: Keep the latest information based on the sliding window strategy The data of a time period is collected, and redundant historical data is eliminated based on the feature contribution evaluation results; The elastic weight solidification algorithm is used to record the importance index of the model parameters, and the parameter update is constrained by the following formula: ; in, For parameters The Fisher information matrix value of is the penalty coefficient, is the smoothing constant, The loss function parameter The gradient of is the basic learning rate, For parameters The amount of updates; Perform adversarial sample generation operations on the newly added data.

7. The sales forecasting method according to claim 1, characterized in that: Also includes: Deploy a reinforcement learning agent module that generates a policy gradient based on the prediction results and feedback from real sales data; Dynamically adjusting hyperparameters in the prediction model based on the policy gradient, including a learning rate decay rate, a batch normalization coefficient, and a fusion weight of a feature pyramid; When a sudden change in external environment data is detected, the Monte Carlo tree search algorithm is enabled to generate a temporary parameter adjustment plan.

8. The sales forecasting method according to claim 1, characterized in that: Also includes: The probability calibration module is placed after the output layer of the prediction model, and the temperature scaling and quantile regression methods are used to calibrate the original prediction results; Construct an uncertainty quantification subnetwork and generate confidence intervals of prediction results through Monte Carlo Dropout sampling; If the confidence interval width exceeds the preset range, the manual review process is triggered and the abnormal case is recorded in the feedback knowledge base; The temperature scaling adjusts the predicted distribution by the following formula: ; in, After calibration Class prediction probability, is the temperature parameter, is the total number of categories, The original output of the model Class logical value.

9. The sales forecasting method according to claim 1, characterized in that: Also includes: Establish a federated learning architecture that allows multiple parties to collaboratively train predictive models under data privacy isolation conditions; Homomorphic encryption technology is used to encrypt and transmit local model gradients, and secure multi-party computing is performed on the aggregation server. Design a differentiated contribution evaluation function for each participant, and dynamically allocate model usage rights based on contribution.

10. A sales forecasting system based on deep learning, characterized in that: include: Data collection module, used to collect historical sales data and related external environment data to build model training data sets; A feature extraction module performs multi-dimensional feature extraction on the model training data set to generate a time series feature set and a context feature set; A model building module, which builds a prediction model based on the fusion of a deep residual network and a long short-term memory network, wherein the prediction model includes a convolution module for extracting local spatiotemporal features and a loop module for capturing long-term dependencies; A model training module, which trains the prediction model in stages according to the time series feature set and the context feature set, wherein the first stage uses unsupervised pre-training to optimize initial parameters, and the second stage uses supervised learning combined with an adaptive learning rate adjustment strategy for fine-tuning; The prediction execution module receives the target sales scenario data input in real time and generates sales prediction results using the trained prediction model; The model update module triggers the incremental learning mechanism if it detects that the confidence of the current prediction result is lower than a preset threshold, and performs online parameter updates on the prediction model based on the data subset updated by the sliding window.

Citation Information

Patent Citations

  • Identity authentication method using microphone of intelligent ear-mounted device

    CN116628658A

  • Electrocardio multi-mode contrast learning method for single-lead arrhythmia diagnosis

    CN118333130A

  • Deep learning-based smart market product demand prediction method and system

    CN118761805A

  • Indoor daily activity identification method based on comparative learning and few-sample learning

    CN119377770A

  • Unmanned aerial vehicle return detection method and system based on AI identification

    CN119649257A

Cited By

  • Mark reinforcement learning-based multi-dimensional preference alignment method and system for large language model

    CN120196748A

  • A multi-dimensional preference alignment method and system for large language models based on labeled reinforcement learning

    CN120196748B

  • Digital modeling method for landslide surge disaster prediction

    CN120633431A

  • Electronic product sales data prediction method and system based on artificial intelligence

    CN121329490A

  • An electronic product sales data prediction method and system based on artificial intelligence

    CN121329490B