Food and beverage network sales trend prediction model construction system and method based on multi-source data fusion and deep learning

Through the method of multi-source data fusion and deep learning, and the use of an improved LSTM-Transformer algorithm, the problems of single data fusion dimension and lack of deep learning architecture in the model in the prediction of food and beverage online sales trends are solved, and accurate prediction of food and beverage online sales trends is achieved, improving the prediction accuracy and stability.

CN120655340AActive Publication Date: 2025-09-16BEIJING TAOMI TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511021641.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-09-16
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

The existing technology for predicting food and beverage online sales trends has a single dimension of multi-source data fusion and fails to achieve cross-domain feature correlation. The prediction model lacks a deep learning architecture and has difficulty processing nonlinear mapping relationships, resulting in insufficient prediction accuracy.

Method used

A method based on multi-source data fusion and deep learning is adopted to collect food and beverage online sales data through web crawlers and API interfaces. Combined with data cleaning, feature extraction and an improved LSTM-Transformer algorithm, nonlinear mapping modeling of food and beverage sales trends is realized, integrating time series features with global dependency analysis.

Benefits of technology

It achieves accurate prediction of food and beverage online sales trends, improves the stability and accuracy of the prediction model, can handle nonlinear correlations in complex scenarios, eliminate abnormal data and unify feature scales, and provide high-quality input data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655340A_ABST
    Figure CN120655340A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of food and beverage, in particular to a food and beverage network sales trend prediction model construction system and method based on multi-source data fusion and deep learning. Comprising a data acquisition unit; a data processing unit; the model construction unit is used for constructing a deep learning prediction model, and an improved LSTM-Transform fusion algorithm is adopted to realize nonlinear mapping modeling of the food and beverage sales trend by integrating time sequence feature modeling and a global dependency relationship analysis technology; a model training verification unit; and a prediction output unit. According to the method, multi-source data such as network sales platform data, social media emotion texts, weather information and industry information are integrated, cross-correlation features such as time dimension features, text emotion features and price elasticity-weather influence are extracted in combination with a feature engineering technology, influence factors of food and beverage sales are comprehensively covered, and the sales quality is improved. The problem that a traditional scheme is single in data dimension is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of food and beverage technology, and in particular to a system and method for constructing a food and beverage network sales trend prediction model based on multi-source data fusion and deep learning. Background Art

[0002] In the food and beverage industry, with the widespread adoption of internet technology, online sales have become a crucial sales channel. The market environment is complex and volatile, and online food and beverage sales trends are influenced by numerous factors, including rapidly evolving consumer preferences, seasonal factors, unexpected social events, and various marketing campaigns. Accurately forecasting sales trends helps companies rationally plan production and precisely allocate marketing resources, gaining an advantage in the fiercely competitive market.

[0003] In the relevant technical field, there have been many studies and practices on sales trend prediction. For example, Chinese patent CN202310005840.6 discloses a method, device, medium, and equipment for predicting the online sales trend of agricultural products. The method includes: crawling the product information of preset agricultural products on the Internet sales platform; extracting characteristic keywords of the preset agricultural products from the product information, forming characteristic tags based on the characteristic keywords, and using the characteristic tags to mark the preset agricultural products sold on the Internet sales platform; determining the sales volume ratio of the preset agricultural products corresponding to each type of characteristic tags; obtaining the comment information of the preset agricultural products for each type of characteristic tags, and determining the positive or negative nature of the comment information corresponding to the characteristic tags; judging whether there is a sales volume ratio greater than the preset ratio value among the sales volume ratios corresponding to each type of characteristic tags; if so, using the characteristic tag corresponding to the sales volume ratio greater than the preset ratio value as the overall representative tag of the preset agricultural products to determine the overall sales volume change trend of the preset agricultural products. The present invention realizes the prediction of the sales volume change trend of agricultural products. For example, Chinese patent CN202411278350.4 discloses a product precision sales analysis system and analysis method based on big data. The system and method obtain product sales-related data, including sales volume, sales revenue and market feedback data, from multiple data sources through a data acquisition module. The data storage module stores the collected product sales data in a structured manner. The data processing module uses big data processing technology to clean, integrate and analyze the stored sales data. The prediction analysis module uses a prediction model based on historical data and existing data trends to predict product sales trends and analyze market demand. The visualization display module presents the analysis results to users in the form of intuitive charts and reports to help users understand market dynamics and product performance, analyze product sales data efficiently and accurately, enhance the company's market insight and decision-making ability, optimize product promotion and sales strategies, and thus improve sales performance and market competitiveness.

[0004] Although the above technical solution has its design advantages, it still has the following technical defects: First, the multi-source data fusion dimension is single and cross-domain feature correlation is not achieved. Chinese patent CN202310005840.6 relies only on the product's own labels and comment data, and does not incorporate external variables such as meteorological information (such as the impact of temperature on cold drink sales) and industry policies. In addition, feature extraction remains at the keyword statistics level and cannot capture cross-correlation features such as "price elasticity-weather fluctuations"; although Chinese patent CN202411278350.4 mentions multi-source data collection, it does not clarify the processing logic of unstructured data (such as social media sentiment text) and lacks a collaborative modeling mechanism for time series features (such as sales cycles) and global features (such as public opinion heat). Second, the prediction model lacks deep learning architecture innovation and has difficulty processing nonlinear mapping relationships. Chinese patent CN202310005840.6 achieves trend prediction through "label ratio threshold determination," but this is essentially a regularized statistical method that cannot capture the complex nonlinear relationships between sales volume and influencing factors. While Chinese patent CN202411278350.4 mentions a "forecasting model," it does not disclose the specific algorithm architecture. Traditional time series models (such as ARIMA) struggle to simultaneously handle long-range dependencies (such as seasonality) and sudden factors (such as promotions), and they do not incorporate an attention mechanism to distribute cross-feature weights, resulting in insufficient prediction accuracy in complex scenarios. Therefore, we propose a system and method for constructing a food and beverage online sales trend forecasting model based on multi-source data fusion and deep learning. Summary of the Invention

[0005] The purpose of the present invention is to provide a system and method for constructing a food and beverage online sales trend prediction model based on multi-source data fusion and deep learning, so as to solve the problems proposed in the above background technology that the multi-source data fusion dimension is single, cross-domain feature association and prediction model lack deep learning architecture innovation, and it is difficult to handle nonlinear mapping relationships.

[0006] To solve the above technical problems, one of the objectives of the present invention is to provide a food and beverage online sales trend prediction model construction system based on multi-source data fusion and deep learning, including: A data collection unit, which is used to collect multi-source data related to food and beverage online sales, and obtains data from online sales platforms, social media platforms, industry information websites, and meteorological information databases based on web crawler technology and API interface calling technology; A data processing unit, which is used to preprocess and extract features from the multi-source data collected by the data collection unit, remove duplicate and erroneous data through a data cleaning algorithm, use a median filter algorithm for denoising, combine the minimum-maximum normalization method and the random forest algorithm to complete data normalization and missing value filling, and use feature engineering technology to extract multi-dimensional prediction features; A model building unit, which is used to build a deep learning prediction model, using an improved LSTM-Transformer fusion algorithm to achieve nonlinear mapping modeling of food and beverage sales trends by integrating time series feature modeling and global dependency analysis technology; A model training and verification unit is used to train and verify the deep learning model. The deep learning model is trained based on a historical data set, the model parameters are adjusted using the Adam optimization algorithm, and the learning rate, number of iterations, and batch size hyperparameters are set to optimize the model. The accuracy, mean square error, and mean absolute error of the validation data set independent of the training data are used to evaluate the model's prediction performance. The prediction output unit is used to output the prediction results of food and beverage network sales trends, receive real-time data processed by the data processing unit, input the trained deep learning model for prediction, and display the prediction results in the form of multi-dimensional charts through Echarts visualization technology.

[0007] As a further improvement of this technical solution, the data collection unit includes an online sales platform data collection module, a social media data collection module, an industry information data collection module, and a meteorological information data collection module, wherein: The online sales platform data collection module obtains sales order information, user purchase behavior data and product price change logs from the e-commerce platform database through the RESTful API interface; The social media data collection module uses the Scrapy framework to build a web crawler to collect user evaluation texts, topic tags, and interaction data on food and beverages from the public API interfaces of multiple e-commerce platforms; The industry information data collection module accesses the industry report website through a scheduled task, parses HTML pages to obtain policy documents, new product release announcements and market analysis white papers; The meteorological information data collection module is based on the OpenWeatherMap API interface and obtains real-time temperature, precipitation probability and humidity data in batches according to region codes, with a time granularity of hourly level.

[0008] As a further improvement of this technical solution, the data processing unit includes a data cleaning module, a data denoising module and a data standardization module, wherein: The data cleaning module generates a unique identifier for the collected sales order data using the MD5 hash algorithm, identifies and deletes duplicate records based on the order ID and timestamp; and uses the 3σ principle to detect price anomalies, marking price fluctuations exceeding ±3 times the historical mean as erroneous data. The data denoising module is used to apply a median filter algorithm with a window size of 5 to the sales time series data, calculate the median within the sliding window to replace the outliers; and use the TF-IDF algorithm to identify and remove spam comments with keyword frequencies exceeding a threshold for user evaluation texts; The data normalization module is used to process numerical features using the minimum-maximum normalization method. The calculation formula is: ;in, is the original eigenvalue, and are the minimum and maximum values ​​of the feature in the training dataset respectively; is the normalized eigenvalue; and the classification features are processed using one-hot encoding.

[0009] As a further improvement of the present technical solution, the data processing unit further includes a feature engineering module, which uses feature engineering technology to extract multi-dimensional prediction features, including the following steps: S240.1. Time dimension feature extraction: Using a time parsing algorithm, extract the week number feature, holiday feature, and seasonal classification feature based on the sales timestamp; specifically, the following: Week number feature: Extract the week number corresponding to the date, and map Monday to Sunday to values ​​1-7 (e.g. Monday = 1, Sunday = 7); Holiday feature: By comparing the statutory holiday schedule issued by the State Council, a binary feature is generated to indicate whether it is a holiday (yes = 1, no = 0); Seasonal classification characteristics: divided into four seasons based on the month (March-May = Spring = 1, June-August = Summer = 2, September-November = Autumn = 3, December-February = Winter = 4).

[0010] S240.2. Text Sentiment Feature Extraction: Generate sentiment features based on user evaluation text using natural language processing technology; specifically, Segment the text and remove stop words; The sentiment polarity score is calculated using the BERT pre-trained model and mapped to the interval [-1, 1] (-1 is strongly negative and 1 is strongly positive). Divide into three emotional levels: negative, neutral, and positive according to the preset threshold; S240.3. Cross-correlation feature extraction: Using data correlation analysis methods, generate price elasticity features and weather impact features based on multi-source data; specifically, including: Price elasticity: Based on the periodic fluctuations in price and sales volume, the ratio of the price change rate to the sales change rate is calculated using the following formula: ; Weather impact characteristics: Through linear regression analysis, a functional relationship between historical temperature data and sales on corresponding dates is established to generate the temperature-sales elasticity coefficient; S240.4. Feature selection and dimensionality reduction: Use the random forest algorithm to select key features based on feature importance: Calculate the importance scores of the time dimension features, text sentiment features, and cross-correlation features extracted by S240.1-S240.3; Arrange in descending order by score and add them up, retain the feature subset with cumulative importance reaching 80%-90%, and eliminate low-contribution features.

[0011] As a further improvement of the present technical solution, the model building unit adopts an improved LSTM-Transformer fusion algorithm to implement nonlinear mapping modeling of food and beverage sales trends, including the following steps: S300.1, Bidirectional encoding of temporal features: A bidirectional long short-term memory network (Bi-LSTM) is used to encode the time dimension features (such as the time series of the week ordinal features and seasonal classification features) extracted in step S240.1, specifically including: Input processing: Arrange the time features into time steps ;in, is the overall representation of time series feature data, is the batch size, is the historical time step (e.g. 14 days), is the time feature dimension (such as 128 dimensions), is the set of real numbers; Network configuration: The number of unidirectional LSTM hidden units is set to 256, and the feature dimension after bidirectional output splicing is 512. The calculation formula is: ; ; ; in, express The hidden state is forward at all times; Represents the forward LSTM calculation function; express Temporal input features at each moment; express The hidden state of the forward LSTM; express The hidden state of the backward LSTM at time t; Represents the backward LSTM calculation function; express The hidden state of the backward LSTM at time t; Represents the temporal encoding features of the final output of the bidirectional LSTM; Represents a splicing operation; Represents the feature dimension description; S300.2. Global Feature Multi-Head Attention Modeling: Perform cross-feature correlation analysis on the global features (e.g., sentiment level, price elasticity) extracted in steps S240.2-S240.3, specifically including: Dimension alignment: The global features are mapped to 512 dimensions through linear transformation to form ;in, is the number of global features (e.g. 10); It is the global feature input; Multi-head Query, Key, Value generation: Perform multi-head projection to generate query vector Query, key vector Key, and value vector Value; No. The projection formula for the 8 heads is: ; ; ; in, Indicates the The query vector of the head; Indicates the The query projection matrix of each head; Indicates the The key vector of the head; Indicates the The key projection matrix of the head; Indicates the A vector of values ​​for each head; Indicates the The Value projection matrix of the head; Self-attention weight calculation and feature aggregation: Calculate the The self-attention weights of each head and the aggregate value vector; ; ; in, Indicates the Among the heads, Features and The association weight of each feature; represents the normalization function; Indicates the In the query vector of the first The feature vector of each sample; Indicates the In the Key vector of each head, The transpose of the feature vector of the samples; Represents the scaling factor, 64 is the dimension of a single attention head (corresponding to the dimension of the preceding projection matrix). Scaling avoids excessive inner product values ​​that lead to the disappearance of the Softmax gradient and makes the weight distribution more reasonable; Indicates the Characteristics after individual aggregation; Indicates the The attention weight matrix of each head; Indicates the The Value vector of the head; Multi-head result splicing and layer normalization: Splice the attention results of 8 heads and perform linear transformation and layer normalization to output global correlation features , ; in, Representation layer normalization operation; It represents the concatenation of the outputs of 8 attention heads to fuse multi-dimensional correlation information; Represents the output projection matrix, which is used to map the concatenated features to the specified dimension; S300.3. Dual-dimensional feature fusion and prediction: The temporal encoding features are fused with the global correlation features and mapped to the prediction space: Feature fusion: Generate 1024-dimensional fusion features by channel splicing , the dimension after aligning the time steps is ;in, Represents the final generated fusion features; Function representing the splicing operation; Represents the temporal features obtained by bidirectional LSTM encoding; Represents the global correlation features obtained by Transformer encoding; Indicates the batch size; represents the time step; Indicates the number of global features; Indicates taking and The larger value in is used to align the time steps; represents the channel dimension of the fused feature, because and Each of them is 512 dimensions, so together they are dimension; Forecast output: mapped to the sales trend forecast dimension (such as sales volume in the next 7 days) through the fully connected layer. The calculation formula is: ; in, Represents the predicted value of the model output; represents a fully connected layer; represents fusion features; Represents the weight matrix of the fully connected layer; represents the bias term of the fully connected layer; Indicates the dimensional form of the predicted value, is the batch size, is the prediction step size; S300.4. Model training and optimization: To improve prediction accuracy, an adaptive training strategy is adopted, including: Define the damage function and use the mean square error to measure the difference between the predicted value and the true value. Quantify the model prediction error by calculating the mean of the squared difference between the predicted value and the true value. The formula is: ; in, is the batch size, is the prediction step length, It is Sample No. The predicted value of the step, is the corresponding true value; Optimization configuration: Adam algorithm is used to optimize model parameters, and the initial learning rate is set to 0.001. In order to balance the convergence speed in the early stage of model training and the stability in the later stage, the learning rate is set to 0.001 after every 10 training epochs. , performs exponential decay, where is the initial learning rate, which allows the model to quickly approach the optimal solution in the early stage and fine-tune the parameters in the later stage; Set the batch size to 64, meaning that 64 samples are fed into the model for each training run. The maximum number of iterations is 100 to control training duration. When dividing the dataset, use 30% of the data as a validation set to monitor the model training process. If the MSE of the validation set increases for three consecutive rounds, indicating that the model may be overfitting, the early stopping mechanism is triggered to terminate training, avoiding ineffective iterations and improving training efficiency.

[0012] As a further improvement of this technical solution, the model training verification unit includes a training parameter configuration module, a parameter iterative adjustment module and a training process monitoring module, wherein: The training parameter configuration module is used to set the hyperparameters of the Adam optimization algorithm, wherein the initial learning rate is selected in the range of 0.0001 to 0.001, the number of iterations is set to 50 to 200 times, and the batch size is selected from 32 to 128; The parameter iterative adjustment module is based on the historical data set and uses the Adam optimization algorithm to iteratively update the model parameters in rounds, covering core parameters such as the fully connected layer weights, bias terms, and LSTM hidden layer states, and uses the gradient descent mechanism to achieve parameter optimization; The training process monitoring module is used to calculate the mean square error and loss function value of the training set in real time during training iterations and draw a loss curve; when the loss decreases by less than 0.001 to 0.01 for 3 to 5 consecutive rounds, the learning rate decay is triggered, and the decay coefficient is 0.1 times the initial learning rate.

[0013] As a further improvement of this technical solution, the model training and verification unit further includes a verification set partitioning module, a performance evaluation calculation module, and a model performance determination module, wherein: The validation set partitioning module splits an independent validation set from the historical data set at a ratio of 20% to 30%, ensuring that the validation set and the training set are consistent in data distribution characteristics (such as seasonal sales trends and promotion frequency); The performance evaluation calculation module uses the validation data set to calculate the model prediction performance, including accuracy (applicable to sales range classification scenarios), mean square error, and mean absolute error, quantifying the model accuracy by the difference between the predicted value and the true value; The model performance determination module calculates the mean square error, mean absolute error, and accuracy (if accuracy evaluation is required in applicable scenarios) based on the prediction results of the validation set. When the mean square error, mean absolute error, and accuracy are used for determination, if the mean square error is less than 0.05-0.2, the mean absolute error is less than 0.03-0.15, and the accuracy (if applicable) is greater than 70%-90%, the model prediction performance is determined to be up to standard; otherwise, the model parameter readjustment or structure optimization process is triggered.

[0014] As a further improvement of this technical solution, the prediction output unit includes a real-time data adaptation module, a model reasoning execution module and an Echarts visualization rendering module, wherein: The real-time data adaptation module is used to receive real-time data from the data processing unit and automatically verify the data dimensions (such as sales volume, price, promotion logo) and time consistency (daily / weekly / monthly); The model inference execution module is used to input the verified data into the trained deep learning prediction model and output the sales trend forecast (including sales volume and category share) for the next 1-30 days; The Echarts visualization rendering module generates visualization charts through Echarts, including at least a line chart comparing historical and predicted sales and a pie chart of product category sales percentage, to intuitively display the prediction results.

[0015] As a further improvement of this technical solution, the Echarts visualization rendering module includes a basic chart generation submodule, an abnormality warning submodule, and an interactive response submodule, wherein: The basic chart generation submodule generates a historical and predicted sales comparison line chart and a product category sales percentage pie chart through Echarts; The abnormal warning submodule uses red markers to identify predicted abnormal points in the line graph based on preset fluctuation thresholds (such as sales growth ≥ 50% or sales decline ≤ 30%); The interactive response submodule is used to respond to category click operations and separately display the history and forecast trend subgraphs of the selected category (such as beverages).

[0016] A second object of the present invention is to provide a method for constructing a food and beverage online sales trend prediction model based on multi-source data fusion and deep learning. The method comprises the following steps: S100. Collect multi-source data related to online food and beverage sales. Using web crawler technology and API calls, obtain data from online sales platforms, social media, industry information websites, and meteorological information databases. S200 pre-processes and extracts features from the multi-source data collected by S100, performing data cleaning, denoising, and normalization, and applying feature engineering techniques to extract multi-dimensional prediction features (such as time dimension, text sentiment, and cross-correlation). S300: Build a deep learning prediction model using an improved LSTM-Transformer fusion algorithm. This algorithm uses bidirectional LSTM to encode temporal features and Transformer to model global features, fusing two-dimensional features and mapping them to the sales trend prediction space. S400: Train a deep learning prediction model based on a historical dataset, using the Adam optimization algorithm, learning rate, batch size, and number of training rounds as hyperparameters. Use a validation set to evaluate accuracy, mean square error, and mean absolute error, and iteratively optimize the deep learning prediction model parameters. S500 inputs the real-time processed data into a trained deep learning prediction model to generate future sales trend forecast results, which are then displayed in the form of multi-dimensional charts using visualization technology.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention integrates multi-source data, including online sales platform data, social media sentiment text, meteorological information, and industry news. It then uses feature engineering techniques to extract time-related features, text sentiment features, and cross-correlated features such as price elasticity and weather impact. This comprehensive approach addresses the factors influencing food and beverage sales, addressing the single-dimensionality issue of traditional solutions. 2. This paper adopts an improved LSTM-Transformer fusion algorithm, using a bidirectional LSTM to capture the temporal dependencies of sales data and a Transformer multi-head attention mechanism to model global feature associations. This achieves nonlinear mapping of sales trends, breaking through the expressive limitations of traditional statistical models or simple time series models. 3. This paper uses the Adam optimization algorithm to dynamically adjust parameters, exponential decay of the learning rate, and early stopping mechanism to avoid model overfitting. Combined with multi-index evaluation of the validation set, it improves the prediction stability of the model in complex scenarios. 4. This invention uses preprocessing operations such as data cleaning, median filtering denoising, and minimum-maximum normalization to eliminate abnormal data and unify feature scales, providing high-quality input for deep learning models and solving the problems of high noise and inconsistent dimensions in multi-source data. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a system framework diagram of the present invention; Figure 2 Schematic diagram of the method steps of the present invention; The meaning of each number in the figure is: 100. Data collection unit; 110. Online sales platform data collection module; 120. Social media data collection module; 130. Industry information data collection module; 140. Meteorological information data collection module; 200, data processing unit; 210, data cleaning module; 220, data denoising module; 230, data standardization module; 240, feature engineering module; 300, model building unit; 400, model training and verification unit; 410, training parameter configuration module; 420, parameter iteration adjustment module; 430, training process monitoring module; 440, verification set partitioning module; 450, performance evaluation calculation module; 460, model performance determination module; 500, prediction output unit; 510, real-time data adaptation module; 520, model inference execution module; 530, Echarts visualization rendering module; 531, basic chart generation submodule; 532, abnormal warning submodule; 533, interactive response submodule. DETAILED DESCRIPTION

[0019] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention. Example 1

[0020] like Figure 1 As shown, this embodiment provides a food and beverage online sales trend prediction model construction system based on multi-source data fusion and deep learning, including: The data collection unit 100 is used to collect multi-source data related to food and beverage online sales. The data is obtained from online sales platforms, social media platforms, industry information websites, and meteorological information databases based on web crawler technology and API interface calling technology; In this step, the data collection unit 100 includes an online sales platform data collection module 110, a social media data collection module 120, an industry information data collection module 130, and a meteorological information data collection module 140, wherein: The online sales platform data collection module 110 obtains sales order information, user purchase behavior data and product price change logs from the e-commerce platform database through the RESTful API interface; As a further illustration of this embodiment, the online sales platform data collection module 110 in this embodiment is connected to mainstream e-commerce platforms (such as Taobao and JD.com). After being authenticated by the API key provided by the platform, the order / get interface is called to obtain order data for the past 12 months, including fields such as order number, product SKU, purchase quantity, and transaction price; user browsing, add-to-cart, and favorite behavior data are collected through the user / behavior interface; at the same time, for unstructured price change logs, regular expressions are used to extract price change time, price difference before and after, and promotion type (such as full discount, discount).

[0021] The social media data collection module 120 uses the Scrapy framework to build a web crawler to collect user evaluation texts, topic tags, and interaction data on food and beverages from the public API interfaces of multiple e-commerce platforms; As a further illustration of this embodiment, the social media data collection module 120 in this embodiment develops a distributed crawler based on the Scrapy framework, sets a User-Agent pool to simulate different browser environments, and avoids being intercepted by the platform's anti-crawling mechanism. For example, when collecting data from the Douyin platform, the JSON data returned by the / api / v1 / comment / list interface is parsed to extract user comment content, number of likes, release time, and associated product topic tags; at the same time, the collected text data is preliminarily cleaned to remove emoticons, HTML tags, and duplicate content, and stored in a MongoDB database using UTF-8 encoding format.

[0022] The industry information data collection module 130 accesses the industry report website through a scheduled task, parses the HTML page to obtain policy documents, new product release announcements and market analysis white papers; As a further illustration of this embodiment, the industry information data collection module 130 in this embodiment uses Python's BeautifulSoup library to parse the page structure of industry websites (such as the official website of the China Food Industry Association) and locates policy document titles, release dates, and text content using XPath expressions. For example, for new product release announcements, key information such as product name, launch date, and target consumer group is extracted. Simultaneously, for market research reports in PDF format, the PyMuPDF library is used to extract text content and locate key data sections through keyword matching (such as "sales volume" and "growth rate").

[0023] The meteorological information data collection module 140 is based on the OpenWeatherMap API interface and obtains real-time temperature, precipitation probability and humidity data in batches according to region codes, with a time granularity of hourly level.

[0024] To further illustrate this embodiment, the meteorological information data collection module 140 in this embodiment divides the collection scope by provincial administrative region, calls the / data / 2.5 / forecast interface of OpenWeatherMap, passes in the latitude and longitude coordinates and API key, and obtains hourly meteorological data for the next seven days. For example, when collecting data for Guangzhou, Guangdong Province, the returned fields include temperature (°C), precipitation probability (%), wind speed (m / s), etc. The collected meteorological data is also spatially and temporally aligned, and indexed by region-time dimension to facilitate subsequent correlation analysis with sales data.

[0025] It should be added that to address the legality, stability, and privacy compliance issues of data acquisition, this embodiment optimizes the collection process as follows: Using the Scrapy framework, we built a rotating pool of User-Agents for over 50 major browsers, randomly switching the identity every three requests. We set a dynamic request interval: initially 1.5 seconds. If the platform's 429 throttling code is triggered, we use an exponential backoff strategy (the first retry interval is 2 seconds, and each subsequent retry interval is doubled based on the previous one). We also integrated a third-party verification code recognition interface (such as Jiyan 3.0) and switched to a highly anonymous proxy IP if a timeout occurred (the proxy pool covers 100+ nodes). For APIs like OpenWeatherMap, a "retry + downgrade" strategy is adopted: retry three times (with 2-second intervals), and if any failure occurs, cached data from the past 4 hours is used. API keys are encrypted and stored in the cloud key management system and dynamically decrypted when called. Social media comment texts are processed through regular desensitization + differential privacy: masking mobile phone numbers (such as 138****1234), blurring addresses (retaining provinces / cities); adding values ​​such as the number of likes =0.6 Laplace noise, which meets the requirements of GDPR and the Personal Information Protection Law.

[0026] The data processing unit 200 is used to preprocess and extract features from the multi-source data collected by the data collection unit 100, remove duplicate and erroneous data through a data cleaning algorithm, use a median filter algorithm for denoising, combine the minimum-maximum normalization method and the random forest algorithm to complete data normalization and missing value filling, and use feature engineering technology to extract multi-dimensional prediction features; Furthermore, the data processing unit 200 in this embodiment receives JSON format data output by the data acquisition unit 100, and the data integrates three types of business dimensions: the sales order dimension includes order ID (unique identifier), timestamp (time series association), price (transaction amount), and sales volume (quantity indicator); the user evaluation dimension includes text content (original user feedback) and evaluation time (timestamp format, supporting cross-dimensional time association); the meteorological data dimension includes region code (regional positioning), temperature (numeric type), and humidity (numeric type), and the time granularity is unified at the hourly level (to ensure consistency in time dimension analysis).

[0027] At the same time, the operating environment of the data processing unit 200 in this embodiment is based on Python 3.8 and higher versions, and relies on the four core libraries of Pandas, Scikit-learn, Transformers, and Numpy to divide the work and cooperate: Pandas is responsible for data cleaning, JSON parsing and tabular operations, Scikit-learn completes feature engineering processing, Transformers relies on the BERT model to parse user evaluation text, and Numpy accelerates numerical calculations, forming a functional layered processing system.

[0028] Furthermore, for batch data processing scenarios (e.g., processing 1,000 orders and associated user reviews and weather data at once), a multi-threaded parallel cleaning mechanism is employed: The batch data is first logically divided into subtasks (e.g., by data volume or time interval), and then multiple threads simultaneously execute these subtasks, including format verification, field normalization, and cross-data type association verification, fully utilizing CPU multi-core resources. Finally, the results from each thread are aggregated to generate a unified cleaned dataset. This strategy, targeted at I / O-intensive or divisible cleaning processes, shortens the processing time for single batches of data and ensures real-time processing of large-scale data.

[0029] In this step, the data processing unit 200 includes a data cleaning module 210, a data denoising module 220 and a data normalization module 230, wherein: The data cleaning module 210 generates a unique identifier for the collected sales order data using the MD5 hash algorithm, identifies and deletes duplicate records based on the order ID and timestamp; and uses the 3σ principle to detect price anomalies, marking price fluctuations exceeding ±3 times the historical mean as erroneous data. As a further explanation of this step, the data cleaning module 210 in this embodiment generates a unique sales order identifier through the MD5 hash algorithm. The specific process is as follows: the order data is hashed using the md5.hexdigest() method of the hashlib library. The generated hash value is stored in the Redis database. Every morning, a scheduled task is used to compare and delete the order records corresponding to the duplicate hash values. The price anomaly detection adopts the 3σ principle, with the price data of the past 30 days as the calculation period. When the current price fluctuation exceeds the historical mean ± 3 times the standard deviation, it is marked as erroneous data - where the historical mean is The arithmetic mean of the price in the past 30 days, the standard deviation is the sample standard deviation of the price in this period, and the calculation formula is: ;in, For the Day price, is the number of days in the calculation cycle, and =30 days; The data denoising module 220 is used to apply a median filter algorithm with a window size of 5 to the sales time series data, calculate the median within the sliding window to replace the outliers; use the TF-IDF algorithm to identify and remove spam comments with keyword frequencies exceeding a threshold for user evaluation texts; As a further illustration of this step, the data denoising module 220 in this embodiment applies a median filter algorithm with a window size of 5 to the sales time series. The core formula is: ;in, represents the original sales value at time t; Represents the value after median filtering. sales value at the moment; indicates; indicates; indicates; The data normalization module 230 is used to process numerical features using the minimum-maximum normalization method, and the calculation formula is: ;in, is the original eigenvalue, and are the minimum and maximum values ​​of the feature in the training dataset respectively; is the normalized eigenvalue; and the classification features are processed using one-hot encoding.

[0030] As a further illustration of this step, the categorical features (such as season) in this embodiment are processed using one-hot encoding, which is implemented by sklearn.preprocessing.OneHotEncoder. For example, the seasonal feature "summer" (value 2) is encoded into a 4-dimensional vector [0, 1, 0, 0].

[0031] In this step, the data processing unit 200 further includes a feature engineering module 240. The feature engineering module 240 uses feature engineering technology to extract multi-dimensional prediction features, including the following steps: S240.1. Time dimension feature extraction: Using a time parsing algorithm, extract the week number feature, holiday feature, and seasonal classification feature based on the sales timestamp; specifically, the following: Week number feature: Extract the week number corresponding to the date, and map Monday to Sunday to values ​​1-7 (e.g. Monday = 1, Sunday = 7); Holiday feature: By comparing the statutory holiday schedule issued by the State Council, a binary feature is generated to indicate whether it is a holiday (yes = 1, no = 0); Seasonal classification characteristics: divided into four seasons based on the month (March-May = Spring = 1, June-August = Summer = 2, September-November = Autumn = 3, December-February = Winter = 4).

[0032] S240.2. Text Sentiment Feature Extraction: Generate sentiment features based on user evaluation text using natural language processing technology; specifically, Segment the text and remove stop words; The sentiment polarity score is calculated using the BERT pre-trained model and mapped to the interval [-1, 1] (-1 is strongly negative and 1 is strongly positive). Divide into three emotional levels: negative, neutral, and positive according to the preset threshold; It should be noted that the BERT pre-training model in this embodiment uses the bert-base-chinese pre-training model and is fine-tuned on 80,000 food and beverage e-commerce reviews and 30,000 social media posts in the same field (labeled as "positive / negative / neutral"): batch size 64, learning rate 5×10⁻ 5 , trained for 3 epochs; optimized the threshold by the F1 score of the validation set, set to “negative (probability > 0.7), positive (probability > 0.7), neutral (otherwise)”, and marked mixed emotions as “-”.

[0033] S240.3. Cross-correlation feature extraction: Using data correlation analysis methods, generate price elasticity features and weather impact features based on multi-source data; specifically, including: Price elasticity: Based on the periodic fluctuations in price and sales volume, the ratio of the price change rate to the sales change rate is calculated using the following formula: ; Weather impact characteristics: Through linear regression analysis, a functional relationship between historical temperature data and sales on corresponding dates is established to generate the temperature-sales elasticity coefficient; the formula is: ;in, is the temperature-sales elasticity coefficient, is the error term, The input is the historical 7-day daily average temperature; is the intercept term of the regression model; It should be added that the cross-feature validation in this embodiment constructs a multiple linear regression model based on 18 months of hourly data: ,in is the sales volume change rate; Screened by 5-fold cross validation A 30-day rolling window is used to monitor the coefficient stability. If the coefficient of adjacent windows changes by >15%, the training data is repartitioned.

[0034] S240.4. Feature selection and dimensionality reduction: Use the random forest algorithm to select key features based on feature importance: Calculate the importance scores of the time dimension features, text sentiment features, and cross-correlation features extracted by S240.1-S240.3; Arrange in descending order by score and add them up, retain the feature subset with cumulative importance reaching 80%-90%, and eliminate low-contribution features.

[0035] As a further illustration of this step, in this embodiment, the random forest algorithm is used to screen key features based on feature importance. This method calculates the contribution of each feature to the model prediction based on the Gini impurity reduction principle. The specific steps include: First, calculate the importance scores of time dimension features, text sentiment features, and cross-correlation features. This process is based on the Gini impurity reduction principle, and the core formula is: ; in, is the number of decision trees in the random forest (set to 100), For the The number of samples processed by a decision tree, is the total number of training samples; is the Gini impurity of the node before splitting (given by calculate, is the number of target variable categories, in the regression task ), is the number of child nodes, and Respectively The number of samples and Gini impurity of each child node. This calculation quantifies the contribution of each feature to reducing sales forecast error by integrating the split gains of multiple decision trees; Then, the calculated feature importance scores are sorted in descending order and the scores are accumulated. The cumulative importance calculation formula is: ; in, After sorting The score of a feature, is the number of features currently filtered, is the total number of features; Next, when the cumulative importance reaches 80%-90%, the screening stops, the corresponding feature subset is retained, and low-contribution features are eliminated. For example, if there are 20 total features and the cumulative importance of the first 12 features reaches 85%, these 12 features (such as season, temperature, and sentiment level) are retained, and redundant features such as "order delivery method" are eliminated. This threshold range is based on practices in the food retail industry, balancing model complexity and prediction accuracy. Finally, the feature subset obtained through the above screening can not only retain key information such as time series, text sentiment and cross-correlation, but also reduce the data dimension, providing efficient input for subsequent deep learning model training.

[0036] A model building unit 300 is used to build a deep learning prediction model, using an improved LSTM-Transformer fusion algorithm to achieve nonlinear mapping modeling of food and beverage sales trends by integrating time series feature modeling and global dependency analysis technology; In this step, the model building unit 300 uses an improved LSTM-Transformer fusion algorithm to implement nonlinear mapping modeling of food and beverage sales trends, including the following steps: S300.1, Bidirectional encoding of temporal features: A bidirectional long short-term memory network (Bi-LSTM) is used to encode the time dimension features (such as the time series of the week ordinal features and seasonal classification features) extracted in step S240.1, specifically including: Input processing: Arrange the time features into time steps ;in, is the overall representation of time series feature data, is the batch size, is the historical time step (e.g. 14 days), is the time feature dimension (such as 128 dimensions), is the set of real numbers; Network configuration: The number of unidirectional LSTM hidden units is set to 256, and the feature dimension after bidirectional output splicing is 512. The calculation formula is: ; ; ; in, express The hidden state is forward at all times; Represents the forward LSTM calculation function; express Temporal input features at each moment; express The hidden state of the forward LSTM; express The hidden state of the backward LSTM at time t; Represents the backward LSTM calculation function; express The hidden state of the backward LSTM at time t; Represents the temporal encoding features of the final output of the bidirectional LSTM; Represents a splicing operation; Represents the feature dimension description; As a further explanation of this step, this embodiment adds a bidirectional structure to the classic LSTM to capture bidirectional temporal dependencies. The specific improvements are as follows: The original LSTM only includes forward calculations. This embodiment adds a backward LSTM. The formula is: ; ; After bidirectional output splicing, the feature dimension is doubled (256-512), which is used to enhance the ability to express temporal features; in, , is the batch size (set to 64), is the historical time step (taken as 14 days), The time feature dimension (128 dimensions, including week number, season, etc.); , the forward hidden state and the backward hidden state Each is 256-dimensional, and after splicing, it is 512-dimensional. The final output dimension is ; S300.2. Global Feature Multi-Head Attention Modeling: Perform cross-feature correlation analysis on the global features (e.g., sentiment level, price elasticity) extracted in steps S240.2-S240.3, specifically including: Dimension alignment: The global features are mapped to 512 dimensions through linear transformation to form ;in, is the number of global features (e.g. 10); It is the global feature input; Multi-head Query, Key, Value generation: Perform multi-head projection to generate query vector Query, key vector Key, and value vector Value; No. The projection formula for the 8 heads is: ; ; ; in, Indicates the The query vector of the head; Indicates the The query projection matrix of each head; Indicates the The key vector of the head; Indicates the The key projection matrix of the head; Indicates the A vector of values ​​for each head; Indicates the The Value projection matrix of the head; Self-attention weight calculation and feature aggregation: Calculate the The self-attention weights of each head and the aggregate value vector; ; ; in, Indicates the Among the heads, Features and The association weight of each feature; represents the normalization function; Indicates the In the query vector of the first The feature vector of each sample; Indicates the In the Key vector of each head, The transpose of the feature vector of the samples; Represents the scaling factor, 64 is the dimension of a single attention head (corresponding to the dimension of the preceding projection matrix). Scaling avoids excessive inner product values ​​that lead to the disappearance of the Softmax gradient and makes the weight distribution more reasonable; Indicates the The characteristics after the aggregation of the heads; The attention weight matrix of each head; Indicates the The Value vector of the head; Multi-head result splicing and layer normalization: Splice the attention results of 8 heads and perform linear transformation and layer normalization to output global correlation features , ; in, Representation layer normalization operation; It represents the concatenation of the outputs of 8 attention heads to fuse multi-dimensional correlation information; Represents the output projection matrix, which is used to map the concatenated features to the specified dimension; As a further illustration of this step, this embodiment makes the following adjustments for the food and beverage sales scenario: The number of attention heads in the original model is usually 8-16. In this embodiment, it is set to 8 heads, adapting 10 global features ( ); The dimension of a single attention head is set to 64 (originally 64 or 128), and the total dimension is 512 (8×64), which is aligned with the LSTM output dimension; And for the The projection formula of 8 heads (total 8 heads) is as follows: , this embodiment clarifies the projection matrix dimensions: ,make sure .

[0037] S300.3. Dual-dimensional feature fusion and prediction: The temporal encoding features are fused with the global correlation features and mapped to the prediction space: Feature fusion: Generate 1024-dimensional fusion features by channel splicing , the dimension after aligning the time steps is ;in, Represents the final generated fusion features; Function representing the splicing operation; Represents the temporal features obtained by bidirectional LSTM encoding; Represents the global correlation features obtained by Transformer encoding; Indicates the batch size; represents the time step; Indicates the number of global features; Indicates taking and The larger value in is used to align the time steps; represents the channel dimension of the fused feature, because and Each of them is 512 dimensions, so together they are dimension; Forecast output: mapped to the sales trend forecast dimension (such as sales volume in the next 7 days) through the fully connected layer. The calculation formula is: ; in, Represents the predicted value of the model output; represents a fully connected layer; represents fusion features; Represents the weight matrix of the fully connected layer; represents the bias term of the fully connected layer; Indicates the dimensional form of the predicted value, is the batch size, is the prediction step size; this design collaboratively models time series features (such as historical sales cycles) and global features (such as weather-sales correlation) to improve prediction capabilities in complex scenarios.

[0038] S300.4. Model training and optimization: To improve prediction accuracy, an adaptive training strategy is adopted, including: Define the damage function and use the mean square error to measure the difference between the predicted value and the true value. Quantify the model prediction error by calculating the mean of the squared difference between the predicted value and the true value. The formula is: ; in, is the batch size, is the prediction step length, It is Sample No. The predicted value of the step, is the corresponding true value; Optimization configuration: Adam algorithm is used to optimize model parameters, and the initial learning rate is set to 0.001. In order to balance the convergence speed in the early stage of model training and the stability in the later stage, the learning rate is set to 0.001 after every 10 training epochs. , performs exponential decay, where is the initial learning rate, which allows the model to quickly approach the optimal solution in the early stage and fine-tune the parameters in the later stage; The batch size was set to 64, meaning that 64 samples were simultaneously input to update the model during each training run. The maximum number of iterations was set to 100 to control training duration. When the dataset was divided, 30% of the data was used as a validation set to monitor the model training progress. When the mean square error (MSE) of the validation set increased for three consecutive rounds, indicating potential overfitting, the early stopping mechanism was triggered to terminate training, avoiding ineffective iterations and improving training efficiency. This strategy ensured stable model convergence despite the non-stationary nature of food and beverage sales data by dynamically adjusting the learning rate and preventing overfitting.

[0039] As a further illustration of this step, the calculation of the loss function in this embodiment relies on strict alignment of the model prediction output with the true label, specifically including: Predicted value : Generated by the model forward propagation, dimension is (like =64 samples, =7-day forecast step), corresponding to "future sales forecast for each sample in the batch" (e.g., sales of 64 stores over the next 7 days); True value :Acquired through the annotation of historical sales data, constructed using the sliding window method - with length The historical features (such as 14-day sales data) are input samples, and the window The actual sales volume of the next 7 days is the label (such as the next 7 days), ensuring that it matches the model output dimension. Exact match.

[0040] As a further illustration of this step, the batch size in this embodiment is (set to 64) is used to balance computational efficiency and gradient stability: If B is too small (such as 16), the gradient fluctuates greatly and the training time increases; If B is too large (e.g., 256), excessive video memory usage and reduced gradient representativeness will occur. A value of 64 is suitable for the "multi-region, multi-SKU" sample size in food and beverage sales forecasting (daily sales data for a single brand can reach tens of thousands, and a batch size of 64 can cover typical combination training). At the same time, the prediction step size in this embodiment is (Set to 7 days) directly corresponds to business characteristics: Sales volume fluctuates periodically during the week (such as weekend peaks). Promotional activities and supply chain replenishment are usually planned on a weekly basis. A 7-day forecast can support actual business decisions.

[0041] Furthermore, a single iteration (Epoch) of the model training in this embodiment includes the following process: Forward propagation: batch input (such as 14-day historical feature matrix, dimension ) is fed into the model to generate predicted values ); Label extraction: Extract the "real sales volume in the next 7 days" corresponding to the batch input from the training data set to form a label ; Loss calculation: Substitute into the mean square error formula: ; in the formula Normalize the "number of samples × time steps" to decouple the loss value from the batch size and prediction step size, ensuring that the losses at different training stages are comparable; Parameter update: The optimizer (such as Adam) adjusts the model parameters according to the gradient, the learning rate =0.001, the parameter is renew( are model parameters, is the loss gradient).

[0042] It should be added that in order to ensure the accuracy of loss function calculation and the stability of the model training process, this embodiment needs to adopt the following constraints to avoid data anomalies and gradient runaway risks: Dimension consistency check: passed before training Logical judgments such as , forced verification of the consistency of model output dimensions and label dimensions, to avoid calculation anomalies caused by data preprocessing errors (such as predicting 7 days but passing in 14-day labels); Gradient stability guarantee: average processing of loss function (divided by ), naturally avoids the risk of "gradient explosion when the number of samples or time steps is too large", making the training process more stable (such as =64, = 7, the gradient scale is only related to the error of a single prediction point).

[0043] A model training and verification unit 400 is used to train and verify the deep learning model. The deep learning model is trained based on a historical data set, the model parameters are adjusted using the Adam optimization algorithm, and the learning rate, number of iterations, and batch size hyperparameters are set to optimize the model. The accuracy, mean square error, and mean absolute error of the validation data set independent of the training data are used to evaluate the model prediction performance. In this step, the model training verification unit 400 includes a training parameter configuration module 410, a parameter iteration adjustment module 420 and a training process monitoring module 430, wherein: The training parameter configuration module 410 is used to set the hyperparameters of the Adam optimization algorithm, wherein the initial learning rate is selected in the range of 0.0001 to 0.001, the number of iterations is set to 50 to 200 times, and the batch size is selected from 32 to 128; As a further explanation of this step, the core formula of the Adam optimization algorithm in this embodiment is: , is the first-order momentum attenuation coefficient; , ,for ; , used for deviation correction and elimination of the impact of initial values; ;in, is the learning rate, ; This embodiment uses the default algorithm parameters ( 、 、 Unchanged), only the following hyperparameters are bounded: Initial learning rate Adapting the feature scale of food and beverage sales data (e.g., 128-dimensional time features + 512-dimensional fusion features) to avoid gradient oscillation caused by excessively large learning rates and convergence stagnation caused by excessively small learning rates. Number of iterations : Covering the complete cycle of the model from rapid convergence to fine-tuning parameters, balancing training efficiency and accuracy; Batch size : 32 is the "GPU memory-friendly" minimum batch size (to ensure gradient stability), 128 is the conventional maximum batch size (to avoid GPU memory overflow), and it is suitable for the sample size of "multi-region, multi-SKU" (a single batch can cover the combined training of 64 stores × 2 SKUs).

[0044] Furthermore, to ensure training stability, this embodiment refines the hyperparameters and verification design: Hyperparameter selection: The initial learning rate [0.0001, 0.001] was determined via grid search; the learning rate decayed to 1 / 10 of its current value after every 10 epochs (balancing efficiency and convergence); the batch size was 64 and the iterations were 100, optimized based on 8GB of GPU memory and 120 million model parameters.

[0045] Validation set construction: Time stratified sampling (training in the first 10 months and validation in the last 2 months) was used, and the distribution consistency was verified by Kolmogorov-Smirnov test: the maximum absolute difference of the empirical distribution of the core features was <0.1 (satisfying =0.05), the distributions are considered aligned.

[0046] The parameter iterative adjustment module 420 updates the model parameters in rounds based on the historical data set through the Adam optimization algorithm, covering core parameters such as the fully connected layer weights, bias terms, and LSTM hidden layer states, and uses the gradient descent mechanism to achieve parameter optimization; The training process monitoring module 430 is used to calculate the mean square error and loss function value of the training set in real time during the training iteration and draw the loss curve; when the loss decreases by less than 0.001 to 0.01 for 3 to 5 consecutive rounds, the learning rate decay is triggered, and the decay coefficient is 0.1 times the initial learning rate.

[0047] As a further explanation of this step, in this embodiment, the training process monitoring module 430 realizes dynamic control of the training process through loss calculation and learning rate decay: the training set mean square error reuse formula , calculate the loss in real time, and draw a curve with the iteration round as the horizontal axis and the loss value as the vertical axis to assist in judging the convergence state of the model; when the loss decreases for 3-5 consecutive rounds ( , adjusted according to the convergence speed), it is judged as "loss stagnation", triggering the step-by-step learning rate decay ( ), adapt to the needs of fine-tuning parameters (such as the learning rate decaying from 0.001 to 0.0001).

[0048] In this step, the model training and verification unit 400 further includes a verification set partitioning module 440, a performance evaluation calculation module 450, and a model performance determination module 460, wherein: The validation set partitioning module 440 splits an independent validation set from the historical data set at a ratio of 20% to 30%, ensuring that the validation set and the training set are consistent in data distribution characteristics such as seasonal sales trends and promotion activity frequency; As a further explanation of this step, the validation set partitioning module 440 in this embodiment adopts a time-stratified sampling strategy to divide the historical data set into a training set (the first 80% time window) and a validation set (the last 20% time window) according to the time series, ensuring that the two are consistent in terms of temporal characteristics such as seasonal sales trends and promotion frequency, thereby avoiding evaluation distortion caused by "future data information leakage"; at the same time, the KS test verifies that there is no significant difference in sales distribution, promotion frequency and seasonal proportion between the training set and the validation set, thereby ensuring the representativeness of the verification results for the generalization ability of the model.

[0049] The performance evaluation calculation module 450 uses the validation data set to calculate the model prediction performance, including accuracy (applicable to sales range classification scenarios), mean square error, and mean absolute error, quantifying the model accuracy by the difference between the predicted value and the true value; As a further explanation of this step, in this embodiment, the performance evaluation calculation module 450 selects adaptation indicators for different prediction scenarios, specifically including: In the continuous sales forecast scenario, the mean square error is quantified using the mean square error, and the formula is: ,in is the number of samples in the validation set; Combined with the mean absolute error, it makes up for the defect of MSE being sensitive to outliers. The formula is: ; In the sales range classification scenario, the accuracy indicator is introduced, and the formula is: , used to measure the accuracy of the model's judgment on the "low / medium / high sales range" and achieve multi-dimensional quantification of prediction performance.

[0050] The model performance determination module 460 calculates the mean square error, mean absolute error, and accuracy (if accuracy evaluation is required in the applicable scenario) based on the prediction results of the validation set. When the mean square error, mean absolute error, and accuracy are used for determination, if the mean square error is less than 0.05-0.2, the mean absolute error is less than 0.03-0.15, and the accuracy (if applicable) is greater than 70%-90%, the model prediction performance is determined to be up to standard; otherwise, the model parameter readjustment or structure optimization process is triggered.

[0051] As a further explanation of this step, this embodiment targets the continuous sales forecast scenario, refers to the magnitude characteristics of the daily sales of a single food and beverage store (500-5000 cups), and sets the mean square error (MSE) threshold [0.05, 0.2] (corresponding to the prediction error being within 4.5% of the true value, such as , with an error of approximately 450 cups when the true mean is 1000). Mean absolute error thresholds of [0.03, 0.15] are used to enhance the MSE's tolerance for outliers. For sales range classification scenarios, accuracy thresholds of [70% to 90%] are set based on the business's requirements for accuracy in "range judgment" (70% is the minimum standard for better than random guessing, and 90% represents high-quality model performance). If the validation set metrics meet the thresholds, the model is saved and deployed. If not, parameter adjustments (such as adjusting the learning rate or number of iterations) or model structure optimization (such as modifying the number of LSTM layers or Transformer heads) are triggered, and the training process is restarted to iteratively upgrade performance.

[0052] The prediction output unit 500 is used to output the food and beverage network sales trend prediction results, receive real-time data processed by the data processing unit 200, input the trained deep learning model for prediction, and display the prediction results in the form of multi-dimensional charts through Echarts visualization technology.

[0053] In this step, the prediction output unit 500 includes a real-time data adaptation module 510, a model reasoning execution module 520, and an Echarts visualization rendering module 530, wherein: The real-time data adaptation module 510 is used to receive real-time data from the data processing unit 200 and automatically verify the data dimensions (such as sales volume, price, promotion logo) and time consistency (daily / weekly / monthly); As a further explanation of this step, after receiving the real-time data stream from the data processing unit 200, the real-time data adaptation module 510 in this embodiment performs a three-level process of dimension verification, time alignment, and pre-processing multiplexing: Dimension verification: Verify the time step of real-time data through logical judgment ( , such as the 14-day historical window during model training) and feature dimensions ( , such as 128-dimensional fusion features), to ensure consistency with the input dimension during the model training phase; Time alignment: Analyze data timestamps and verify that the time granularity (day / week / month) matches the time characteristics during model training to avoid prediction bias caused by time scale differences. Preprocessing reuse: Call the feature engineering process of the data processing unit 200 to perform standardization and encoding on the real-time data consistent with the training set: Numerical feature normalization: reuse the mean of the training set and standard deviation , which is converted by the formula , ensuring consistent feature distribution; among them, is the normalized numerical feature; is the original numerical feature; Category feature encoding: Reuse the encoding mapping of the training set (such as the promotion identifier "yes-1, no-0") to ensure the consistency of feature semantics.

[0054] The model inference execution module 520 is used to input the verified data into the trained deep learning prediction model and output the sales trend forecast (including sales volume and category share) for the next 1-30 days; As a further explanation of this step, the model reasoning execution module 520 in this embodiment implements the mapping of real-time data to sales trend forecasts through the following process: First, the real-time data is reshaped into a single-sample time series tensor (dimension is ),in For a historical time step such as 14 days, For feature dimensions such as 128 dimensions), ensure the consistency of time dependencies; Then call the trained deep learning model, the output dimension is The prediction tensor ( For forecasting steps such as 1–30 days, is the number of food and beverage categories, and the "+1" dimension corresponds to the global total sales); Finally, the output tensor is parsed to directly extract the global total sales forecast of the last dimension and use the formula , calculate the proportion of each category and complete the multi-dimensional prediction output of "total sales + category structure".

[0055] The Echarts visualization rendering module 530 generates visualization charts through Echarts, including at least a line chart comparing historical and predicted sales and a pie chart of product category sales percentage, to intuitively display the prediction results.

[0056] In this step, the Echarts visualization rendering module 530 includes a basic chart generation submodule 531, an abnormality warning submodule 532, and an interactive response submodule 533, wherein: The basic chart generation submodule 531 generates a historical and predicted sales comparison line chart and a product category sales percentage pie chart through Echarts; As a further explanation of this step, for the historical and forecasted sales comparison line chart, first connect the forecast data output by the model (such as sales in the next 1-30 days) with the historical real data (such as sales in the past 14 days) along the time axis to construct a continuous time series (ensuring a seamless connection between the last day of history and the first day of forecast). Through Echarts configuration, use solid lines (such as gray) to render historical data and dotted lines (such as blue) to distinguish forecast data, achieving a visual comparison of sales trends. Furthermore, for the product category sales share pie chart, the sales volume of each category is predicted based on the model output, using the formula: , calculate the category proportion, where For the Forecast sales of each category, The total number of food and beverage categories; a ring pie chart layout is used to avoid crowding of the central data, and the label format is set to "Category Name: Percentage %". Through mouse hover interaction, the predicted sales volume and proportion details of the category are displayed in the prompt box.

[0057] Furthermore, the basic chart generation submodule 531 in this embodiment also supports time range synchronization: when the line chart adjusts the time range through interaction, the pie chart automatically updates the category proportion of the corresponding time period to ensure data consistency; through adaptive rendering strategies (such as screen size adaptation) and large data set sampling optimization (such as enabling mean sampling), visualization performance is guaranteed under different data scales and display terminals.

[0058] In addition, for real-time data verification, this embodiment sets the following fault-tolerant process: When missing dimensions (such as null values ​​in weather data) or out-of-order timestamps are detected: If there are ≤3 consecutive missing time points, linear interpolation is used (fitting the 4 valid points before and after); if there are >3 missing points, the historical mean of the same day and time period in the past 3 years is used; If the verification failure rate is greater than 15%, the system sends an email alert containing the failed sample and error code, automatically suspends inference, and switches to the "historical model" (last week's snapshot) output until the data source is restored.

[0059] The abnormal warning submodule 532 uses a red marker to identify the predicted abnormal points in the line graph based on a preset fluctuation threshold (e.g., sales increase ≥ 50% or sales decrease ≤ 30%); As a further explanation of this step, the abnormal warning in this embodiment dynamically sets the threshold based on historical fluctuation statistics: first calculate the daily month-on-month increase in the training set:

[0060] , For the sales volume of the day, The standard deviation of the sales volume of the previous day and the statistical increase Then, set the abnormal threshold to (Thresholds are adjusted based on business characteristics); Forecast points that meet the threshold are highlighted with a flashing red mark, and a prompt box is configured to display the abnormal magnitude (such as "sales growth +52%") to assist in quickly identifying abnormal fluctuations.

[0061] The interactive response submodule 533 is used to respond to a category click operation and separately display the history and forecast trend subgraphs of the selected category (such as beverages).

[0062] As a further explanation of this step, the interactive response submodule 533 in this embodiment relies on visual interactive events to achieve data linkage: monitoring the category click event of the pie chart, obtaining the identifier of the selected category (such as "beverage"); extracting the prediction sequence of the corresponding category from the model output tensor (such as the first dimensional data), combined with historical data to reuse the line chart configuration logic, dynamically generate a "single category history - forecast trend sub-graph", achieve seamless switching between "global view - category details", and improve data exploration efficiency. Example 2

[0063] like Figure 2 As shown, this embodiment also provides a method for constructing a food and beverage network sales trend prediction model based on multi-source data fusion and deep learning. The method includes the following steps: S100. Collect multi-source data related to online food and beverage sales. Using web crawler technology and API calls, obtain data from online sales platforms, social media, industry information websites, and meteorological information databases. S200 pre-processes and extracts features from the multi-source data collected by S100, performing data cleaning, denoising, and normalization, and applying feature engineering techniques to extract multi-dimensional prediction features (such as time dimension, text sentiment, and cross-correlation). S300: Build a deep learning prediction model using an improved LSTM-Transformer fusion algorithm. This algorithm uses bidirectional LSTM to encode temporal features and Transformer to model global features, fusing two-dimensional features and mapping them to the sales trend prediction space. S400: Train a deep learning prediction model based on a historical dataset, using the Adam optimization algorithm, learning rate, batch size, and number of training rounds as hyperparameters. Use a validation set to evaluate accuracy, mean square error, and mean absolute error, and iteratively optimize the deep learning prediction model parameters. S500 inputs the real-time processed data into a trained deep learning prediction model to generate future sales trend forecast results, which are then displayed in the form of multi-dimensional charts using visualization technology.

[0064] Those skilled in the art will appreciate that the process of implementing all or part of the steps of the above embodiments may be accomplished by hardware, or by instructing related hardware through a program.

[0065] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A food and beverage online sales trend prediction model construction system based on multi-source data fusion and deep learning, characterized by: include: A data collection unit (100), the data collection unit (100) is used to collect multi-source data related to food and beverage online sales, and obtain data from online sales platforms, social media platforms, industry information websites, and meteorological information databases based on web crawler technology and API interface calling technology; A data processing unit (200) is used to pre-process and extract features from the multi-source data collected by the data collection unit (100), remove duplicate and erroneous data through a data cleaning algorithm, use a median filter algorithm for denoising, combine a minimum-maximum normalization method and a random forest algorithm to complete data normalization and missing value filling, and use feature engineering technology to extract multi-dimensional prediction features; A model building unit (300), wherein the model building unit (300) is used to build a deep learning prediction model, adopt an improved LSTM-Transformer fusion algorithm, and realize nonlinear mapping modeling of food and beverage sales trends by integrating time series feature modeling and global dependency analysis technology; A model training and verification unit (400) is used to train and verify a deep learning model, train the deep learning model based on a historical data set, adjust model parameters using an Adam optimization algorithm, set learning rate, number of iterations, and batch size hyperparameters for model optimization, and use a verification data set independent of the training data to calculate accuracy, mean square error, and mean absolute error to evaluate model prediction performance; A prediction output unit (500) is used to output food and beverage network sales trend prediction results, receive real-time data processed by the data processing unit (200), input the trained deep learning model for prediction, and display the prediction results in the form of multi-dimensional charts using Echarts visualization technology.

2. The food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to claim 1 is characterized in that: The data collection unit (100) includes an online sales platform data collection module (110), a social media data collection module (120), an industry information data collection module (130), and a meteorological information data collection module (140), wherein: The online sales platform data acquisition module (110) obtains sales order information, user purchase behavior data and product price change logs from the e-commerce platform database through a RESTful API interface; The social media data collection module (120) uses the Scrapy framework to build a web crawler to collect user evaluation texts, topic tags and interaction data on food and beverages from the public API interfaces of multiple e-commerce platforms; The industry information data collection module (130) accesses the industry report website through a scheduled task, parses the HTML page to obtain policy documents, new product release announcements and market analysis white papers; The meteorological information data acquisition module (140) is based on the OpenWeatherMap API interface and obtains real-time temperature, precipitation probability and humidity data in batches according to region codes.

3. The food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to claim 1 is characterized in that: The data processing unit (200) includes a data cleaning module (210), a data denoising module (220) and a data standardization module (230), wherein: The data cleaning module (210) generates a unique identifier for the collected sales order data using the MD5 hash algorithm, identifies and deletes duplicate records based on the order ID and timestamp; and uses the 3σ principle to detect price anomalies, marking them as erroneous data when the price fluctuation exceeds ±3 times the historical mean; The data denoising module (220) is used to apply a median filter algorithm with a window size of 5 to the sales time series data, calculate the median in the sliding window to replace the outliers; use the TF-IDF algorithm to identify spam comments with keyword frequencies exceeding a threshold value for user evaluation texts and remove them; The data normalization module (230) is used to process numerical features using the minimum-maximum normalization method, and the calculation formula is: ;in, is the original eigenvalue, and are the minimum and maximum values ​​of the feature in the training dataset respectively; is the normalized eigenvalue; and the classification features are processed using one-hot encoding.

4. The food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to claim 3 is characterized in that: The data processing unit (200) further includes a feature engineering module (240), wherein the feature engineering module (240) extracts multi-dimensional prediction features using feature engineering technology, including the following steps: S240.

1. Time dimension feature extraction: Using a time parsing algorithm, extract the week number feature, holiday feature, and seasonal classification feature based on the sales timestamp; S240.

2. Text Sentiment Feature Extraction: Generate sentiment features based on user evaluation text using natural language processing technology; S240.

3. Cross-correlation feature extraction: Using data correlation analysis methods, generate price elasticity features and weather impact features based on multi-source data; S240.

4. Feature selection and dimensionality reduction: Use the random forest algorithm to select key features based on feature importance: Calculate the importance scores of the time dimension features, text sentiment features, and cross-correlation features extracted by S240.1-S240.3; Arrange in descending order by score and add them up, retain the feature subset with cumulative importance reaching 80%-90%, and eliminate low-contribution features.

5. The food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to claim 4 is characterized in that: The model building unit (300) uses an improved LSTM-Transformer fusion algorithm to implement nonlinear mapping modeling of food and beverage sales trends, including the following steps: S300.1, Bidirectional encoding of temporal features: A bidirectional long short-term memory network is used to encode the time dimension features extracted in step S240.1, specifically including: Input processing: Arrange the time features into time steps ;in, is the overall representation of time series feature data, is the batch size, is the historical time step, is the time feature dimension, is the set of real numbers; Network configuration: The number of unidirectional LSTM hidden units is set to 256, and the feature dimension after bidirectional output splicing is 512. The calculation formula is: ; ; ; in, express The hidden state is forward at all times; Represents the forward LSTM calculation function; express Temporal input features at each moment; express The hidden state of the forward LSTM; express The hidden state of the backward LSTM at time t; Represents the backward LSTM calculation function; express The hidden state of the backward LSTM at time t; Represents the temporal encoding features of the final output of the bidirectional LSTM; Represents a splicing operation; Represents the feature dimension description; S300.

2. Global Feature Multi-Head Attention Modeling: Perform cross-feature correlation analysis on the global features extracted in steps S240.2-S240.3, specifically including: Dimension alignment: The global features are mapped to 512 dimensions through linear transformation to form ;in, is the number of global features; It is the global feature input; Multi-head Query, Key, Value generation: Perform multi-head projection to generate query vector Query, key vector Key, and value vector Value; No. The projection formula of the head is: ; ; ; in, Indicates the The query vector of the head; Indicates the The query projection matrix of each head; Indicates the The key vector of the head; Indicates the The key projection matrix of the head; Indicates the A vector of values ​​for each head; Indicates the The Value projection matrix of the head; Self-attention weight calculation and feature aggregation: Calculate the The self-attention weights of each head and the aggregate value vector; ; ; in, Indicates the Among the heads, Features and The association weight of each feature; represents the normalization function; Indicates the In the query vector of the first The feature vector of each sample; Indicates the In the Key vector of each head, The transpose of the feature vector of the samples; represents the scaling factor; Indicates the Characteristics after individual aggregation; Indicates the The attention weight matrix of each head; Indicates the The Value vector of the head; Multi-head result splicing and layer normalization: Splice the attention results of 8 heads and perform linear transformation and layer normalization to output global correlation features : ; in, Representation layer normalization operation; It represents the concatenation of the outputs of 8 attention heads to fuse multi-dimensional correlation information; Represents the output projection matrix, which is used to map the concatenated features to the specified dimension; S300.

3. Dual-dimensional feature fusion and prediction: The temporal encoding features are fused with the global correlation features and mapped to the prediction space: Feature fusion: Generate 1024-dimensional fusion features by channel splicing , the dimension after aligning the time steps is ;in, Represents the final generated fusion features; Function representing the splicing operation; Represents the temporal features obtained by bidirectional LSTM encoding; Represents the global correlation features obtained by Transformer encoding; Indicates the batch size; represents the time step; Indicates the number of global features; Indicates taking and The larger value in is used to align the time steps; Indicates the channel dimension of the fused feature; Prediction output: mapped to the sales trend prediction dimension through the fully connected layer, the calculation formula is: ; in, Represents the predicted value of the model output; represents a fully connected layer; represents fusion features; Represents the weight matrix of the fully connected layer; represents the bias term of the fully connected layer; Indicates the dimensional form of the predicted value, is the batch size, is the prediction step size; S300.

4. Model training and optimization: To improve prediction accuracy, an adaptive training strategy is adopted, including: Define the damage function and use the mean square error to measure the difference between the predicted value and the true value. , quantify the model prediction error, the formula is: ; in, is the batch size, is the prediction step length, It is Sample No. The predicted value of the step, is the corresponding true value; Optimization configuration: Adam algorithm is used to optimize model parameters, and the initial learning rate is set to 0.001; in order to balance the convergence speed in the early stage of model training and the stability in the later stage, the learning rate is set to 0.001 after every 10 training epochs. , performs exponential decay, where is the initial learning rate; The batch size is set to 64, meaning that 64 samples are input simultaneously to update the model during each training session. The maximum number of iterations is set to 100 to control the training duration. When dividing the dataset, 30% of the data is used as a validation set to monitor the model training process. When the MSE of the validation set increases for three consecutive rounds, the early stopping mechanism is triggered to terminate training.

6. The food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to claim 1 is characterized in that: The model training verification unit (400) includes a training parameter configuration module (410), a parameter iteration adjustment module (420) and a training process monitoring module (430), wherein: The training parameter configuration module (410) is used to set the hyperparameters of the Adam optimization algorithm, wherein the initial learning rate is selected from a range of 0.0001 to 0.001, the number of iterations is set to 50 to 200 times, and the batch size is selected from 32 to 128; The parameter iterative adjustment module (420) updates the model parameters in rounds based on the historical data set through the Adam optimization algorithm, and optimizes the parameters using the gradient descent mechanism; The training process monitoring module (430) is used to calculate the mean square error and loss function value of the training set in real time during the training iteration and draw the loss curve; when the loss decreases by less than 0.001 to 0.01 for 3 to 5 consecutive rounds, the learning rate decay is triggered, and the decay coefficient is 0.1 times the initial learning rate.

7. The food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to claim 6 is characterized in that: The model training and verification unit (400) further includes a verification set partitioning module (440), a performance evaluation calculation module (450) and a model performance determination module (460), wherein: The verification set division module (440) divides the historical data set into an independent verification data set at a ratio of 20% to 30%, ensuring that the verification set and the training set are consistent in data distribution characteristics; The performance evaluation calculation module (450) uses the validation data set to calculate the model prediction performance, including accuracy, mean square error, and mean absolute error, and quantifies the model accuracy by the difference between the predicted value and the true value; The model performance determination module (460) calculates the mean square error, mean absolute error and accuracy based on the prediction results of the validation set; when the mean square error, mean absolute error and accuracy are used for determination, if the mean square error is less than 0.05-0.2, the mean absolute error is less than 0.03-0.15 and the accuracy is greater than 70%-90%, the model prediction performance is determined to be up to standard; otherwise, the model parameter readjustment or structure optimization process is triggered.

8. The food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to claim 1 is characterized in that: The prediction output unit (500) includes a real-time data adaptation module (510), a model reasoning execution module (520) and an Echarts visualization rendering module (530), wherein: The real-time data adaptation module (510) is used to receive real-time data from the data processing unit (200) and automatically verify the data dimension and time consistency; The model reasoning execution module (520) is used to input the verified data into the trained deep learning prediction model and output the sales trend forecast for the next 1-30 days; The Echarts visualization rendering module (530) generates visualization charts through Echarts, including at least a line chart comparing historical and predicted sales volume and a pie chart of product category sales percentage, to intuitively display the prediction results.

9. The food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to claim 8 is characterized in that: The Echarts visualization rendering module (530) includes a basic chart generation submodule (531), an abnormal warning submodule (532) and an interactive response submodule (533), wherein: The basic chart generation submodule (531) generates a historical and predicted sales comparison line chart and a product category sales percentage pie chart through Echarts; The abnormal warning submodule (532) identifies the predicted abnormal point in the line graph using a red mark based on a preset fluctuation threshold; The interactive response submodule (533) is used to respond to category click operations and separately display the history and forecast trend subgraphs of the selected category.

10. A method for constructing a food and beverage network sales trend prediction model based on multi-source data fusion and deep learning, based on the food and beverage network sales trend prediction model construction system based on multi-source data fusion and deep learning according to any one of claims 1 to 9, characterized in that: The steps include: S100. Collect multi-source data related to online food and beverage sales. Using web crawler technology and API calls, obtain data from online sales platforms, social media, industry information websites, and meteorological information databases. S200 pre-processes and extracts features from the multi-source data collected by S100, performs data cleaning, denoising, and normalization, and uses feature engineering technology to extract multi-dimensional prediction features; S300: Build a deep learning prediction model using an improved LSTM-Transformer fusion algorithm. This algorithm uses bidirectional LSTM to encode temporal features and Transformer to model global features, fusing two-dimensional features and mapping them to the sales trend prediction space. S400: Train a deep learning prediction model based on a historical dataset, using the Adam optimization algorithm, learning rate, batch size, and number of training rounds as hyperparameters. Use a validation set to evaluate accuracy, mean square error, and mean absolute error, and iteratively optimize the deep learning prediction model parameters. S500 inputs the real-time processed data into a trained deep learning prediction model to generate future sales trend forecast results, which are then displayed in the form of multi-dimensional charts using visualization technology.

Citation Information

Patent Citations

  • Agricultural product network sales trend prediction method and device, medium and equipment

    CN116308453A

  • Product precision sales analysis system and analysis method based on big data

    CN119130523A

  • Commodity sales volume prediction method and device based on Transformer + LSTM neural network model

    CN111626764A

  • Ship trajectory prediction method based on improved P-B-T

    CN118861532A

  • Visual processing method and system for sales data

    CN119202064A

Cited By

  • Method and system for detecting legal network traffic in complex environment

    CN121000523A

  • Method for detecting and analyzing content of liquid flavor substances in beverage

    CN121364222A

  • Method and system for constructing multi-source time series data fusion prediction model in BI analysis

    CN121480859A

  • Method and system for constructing multi-source time series data fusion prediction model in bi analysis

    CN121480859B