Tobacco retailer sales data prediction method and device

By adopting two-way LSTM and attention mechanism in the retail sales data prediction model, combined with multi-source data and feature processing technology, the problem of inaccurate data prediction in the existing technology is solved, and higher prediction accuracy and effective early warning functions are achieved.

CN119991195APending Publication Date: 2025-05-13CHINA NAT TOBACCA CORP YUNNAN CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510436023.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately process continuous data when processing cigarette sales data of retailers, resulting in false positive or false negative warning results and inaccurate prediction values.

Method used

A two-way LSTM and attention mechanism are used to build a retail user sales data prediction model, obtain multi-source data such as retail user profile, location information, and geographical characteristics. Through feature screening, batch normalization and embedding layer processing, data analysis and prediction are carried out in combination with two-way LSTM and attention mechanism.

Benefits of technology

It improves the accuracy and robustness of time series prediction, effectively improves prediction accuracy, and achieves more accurate retail sales data prediction and daily supervision and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991195A_ABST
    Figure CN119991195A_ABST
Patent Text Reader

Abstract

The invention discloses a tobacco retailer sales data prediction method and device, and relates to the technical field of data processing, and the method comprises the steps: obtaining retailer sample feature data, carrying out the feature screening, carrying out the batch normalization processing of continuous features, and carrying out the embedded layer processing of discrete features; inputting the normalized data into a bidirectional LSTM layer, and outputting feature representation of each time point; splicing the output of the bidirectional LSTM layer and the output of the embedded layer to form spliced data, calculating an attention weight through an attention mechanism, and carrying out weighted summation on the spliced data to obtain a context vector; and inputting the context vector into a full-connection layer to output retailer sales prediction data, comparing the sales prediction data with real sales data, and giving an early warning when the sales prediction data is higher than a preset upper threshold or lower than a preset lower threshold. According to the method, the retailer sales data prediction model is constructed, the prediction accuracy, robustness and precision are improved, and effective supervision and early warning of the retailer sales data are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and device for predicting sales data of tobacco retailers. Background Art

[0002] In the cigarette sales industry, retailers, as the terminal of cigarette sales, are an important component in determining cigarette sales. Daily supervision and early warning of cigarette sales by retailers are required to prevent retailers from engaging in illegal activities such as cross-selling, which may lead to problems in cigarette sales. The current mainstream early warning algorithms are: 1. Based on the association rule learning model, the Apriori algorithm is used to generate candidate rules, and the rule analysis is performed on the retailer's data. The meaningful rules are screened out through the indicators of Support, Confidence, and Lift. After defining the rules for abnormal product flow, the abnormal indicators of retailers can be obtained through cross-validation of the retailer's data; 2. Based on the anomaly detection model, by collecting statistics on the sales data of retailers and using LsolationForest for analysis and calculation, abnormal sales activities can be identified, and abnormal sales can be warned and identified.

[0003] However, the existing model for early warning of retailers is based on rule association analysis of retailers' behaviors. Although it is possible to predict and judge abnormal behaviors of retailers, it is difficult to process continuous data in terms of data processing, which may lead to deviations in the results of early warnings, resulting in false positive results (false alarms) or false negative results (missed alarms), and the resulting predicted values ​​are inaccurate. In the tobacco industry, there is a set of complex rules and gear requirements for the delivery of products to retailers, which requires the support of multi-faceted, multi-element, and continuous data sources to obtain accurate prediction values. Summary of the invention

[0004] In view of the above-mentioned defects or deficiencies in the prior art, the present invention provides a method and device for predicting tobacco retailer sales data, which obtains multi-source data such as retailer stalls, location information, geographic features, etc., and constructs a retailer sales data prediction model based on bidirectional LSTM and attention mechanism, so as to achieve more accurate sales analysis and prediction of retailer data.

[0005] One aspect of the present invention provides a tobacco retailer sales data prediction method comprising: Obtain sample characteristic data of retail households and perform feature screening, perform batch normalization processing on continuous features in the screened characteristic data to obtain normalized data, and use embedding layer processing on discrete features in the screened characteristic data to obtain embedding layer output data; Input the normalized data into a bidirectional LSTM layer, capture the features of the normalized data from the forward and backward directions respectively, and combine and output the feature representation of each time point; The output of the bidirectional LSTM layer is concatenated with the output of the embedding layer to form concatenated data, an attention weight is calculated through an attention mechanism, and a context vector is obtained by weighted summing the concatenated data according to the attention weight; The context vector is input into the fully connected layer to output retail sales forecast data, and the sales forecast data is compared with the actual sales data. When the sales forecast data is higher than a preset upper threshold or lower than a preset lower threshold, an early warning is issued.

[0006] Furthermore, the feature screening includes stepwise regression method, determination coefficient method and L2 regularization.

[0007] Furthermore, it also includes: the filtered feature data includes batch size, sequence length and feature dimension.

[0008] Furthermore, the batch normalization of the continuous features in the filtered feature data to obtain normalized data specifically includes: Dividing the continuous features into batch data of several batches and storing them in a data table; Calculate the mean of each feature dimension in the batch data of the current batch: in, is the mean of feature dimension d, m is the batch size, is the feature value of the i-th sample in feature dimension d; Calculate the variance of each feature dimension in the batch data of the current batch: in, is the variance of feature dimension d; The features of each sample are normalized using the mean and variance: in, is the eigenvalue of the i-th sample after standardization in the feature dimension d, is a constant; Scale and translate the normalized data: in, is the final output value of the i-th sample in the feature dimension d, is the scaling parameter, is the translation parameter.

[0009] Furthermore, the method of using an embedding layer to process the discrete features in the filtered feature data to obtain embedding layer output data includes: Create a unique integer index for each unique category or word; Initialize an embedding matrix, where each row of the embedding matrix corresponds to an embedding vector of a category or word; Given a new input value, use the index of the input value to find the corresponding embedding vector; During the training process, the embedding matrix is ​​optimized by the back propagation algorithm, and the optimization formula is as follows: Among them, index is the index, is the embedding vector extracted from the embedding matrix E according to the index, To optimize the updated embedding vector, is the learning rate, For loss.

[0010] Furthermore, the normalized data is input into a bidirectional LSTM layer, the features of the normalized data are captured from the forward and backward directions respectively, and the feature representation of each time point is combined and outputted, including: Obtaining the forward hidden state of the normalized data through a forward LSTM; Obtaining a backward hidden state of the normalized data through a backward LSTM; The forward hidden state and the backward hidden state are combined into a bidirectional hidden state by concatenation or weighted summation, and the feature representation of each time point is output.

[0011] Furthermore, the calculating of the attention weight by the attention mechanism and performing weighted summation on the spliced ​​data according to the attention weight to obtain the context vector include: The concatenated data is mapped into a query vector, a key vector, and a value vector. For each output time point t, its correlation score with all input time points j is calculated: in, is the query vector at time point t, is the transpose of the key vector at time point j, is the dimension of the key vector; Convert the relevance scores into probability distributions: in, is the attention weight of the output at time point t to the input at time point j; The context vector is obtained by weighted summing the value vector using the attention weights: in, is the context vector at time point t, is the value vector at time point j.

[0012] Furthermore, the step of inputting the context vector into a fully connected layer to output retail sales forecast data includes: Concatenate the context vector and the query vector to form a feature representation : The feature is represented Input the fully connected layer and output the feature representation Z: in, is the weight matrix of the fully connected layer, is the bias term of the fully connected layer; Feature Representation Apply the ReLU activation function and output the activated feature representation A: The activated feature representation A generates a prediction result through the fully connected layer and the output layer: Among them, Softmax() is the activation function that converts the output of the fully connected layer into a probability distribution. is the weight matrix of the output layer, is the bias term of the output layer.

[0013] Furthermore, it also includes: Compute the boundaries of the uniform distribution of the current linear layer: Among them, a is the interval limit, gain is the scaling factor adjusted according to the activation function, Enter the number of features for the current layer, Output feature count for the current layer; Initialize each element of the weight matrix of the current linear layer to the interval Random value within .

[0014] Another aspect of the present invention provides a tobacco retailer sales data forecasting device, comprising: The first module is configured to obtain sample characteristic data of retail households and perform characteristic screening, perform batch normalization processing on continuous characteristics in the screened characteristic data to obtain normalized data, and use an embedding layer to process discrete characteristics in the screened characteristic data to obtain embedding layer output data; The second module is configured to input the normalized data into a bidirectional LSTM layer, capture the features of the normalized data from the forward and backward directions respectively, and merge and output the feature representation of each time point; A third module is configured to concatenate the output of the bidirectional LSTM layer with the output of the embedding layer to form concatenated data, calculate the attention weight through an attention mechanism, and perform weighted summation on the concatenated data according to the attention weight to obtain a context vector; The fourth module is configured to input the context vector into the fully connected layer to output retail sales forecast data, compare the sales forecast data with the actual sales data, and issue an early warning when the forecast data is higher than a preset upper threshold or lower than a preset lower threshold.

[0015] The present invention provides a method and device for predicting tobacco retailer sales data, which obtain multi-source data such as retailer stalls, location information, and geographic features, and construct a retailer sales data prediction model based on a bidirectional LSTM and an attention mechanism. The model is a time series prediction model, which improves the accuracy and robustness of time series prediction on the basis of data support from multi-faceted, multi-element, and continuous data sources, combines the analysis and processing of continuous data and discrete data, and effectively improves the prediction accuracy, realizes more accurate retailer sales data prediction, and realizes effective daily supervision and early warning of retailer sales data. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 It is a flowchart of a tobacco retailer sales data prediction method provided by an embodiment of the present application; Figure 2 It is a structural diagram of a tobacco retailer sales data prediction device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0018] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings.

[0019] It should be understood that although the terms first, second, third, etc. may be used to describe the acquisition modules in the embodiments of the present invention, the acquisition modules should not be limited to these terms. These terms are only used to distinguish the acquisition modules from each other.

[0020] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0021] It should be noted that the directional words such as "upper", "lower", "left", and "right" described in the embodiments of the present invention are described at the angles shown in the drawings and should not be understood as limiting the embodiments of the present invention. In addition, in the context, it should also be understood that when it is mentioned that an element is formed "on" or "under" another element, it can not only be formed directly "on" or "under" another element, but also be formed "on" or "under" another element indirectly through an intermediate element.

[0022] refer to Figure 1 The embodiment of the present invention provides a tobacco retailer sales data prediction method, comprising: Step S101, obtaining sample characteristic data of retail customers and performing characteristic screening, performing batch normalization processing on continuous characteristics in the screened characteristic data to obtain normalized data, and performing embedding layer processing on discrete characteristics in the screened characteristic data to obtain embedding layer output data; Specifically, data related to retailers in the sample area, such as retail sales data, purchase data, stall information and location information, etc., are obtained from the internal database of the tobacco company; geographic feature data around retailers, such as passenger flow, occupational distribution of the crowd, etc., are obtained from the map manufacturer; the retail samples are divided into a training set, a validation set and a test set; the sample area is divided into multiple plots, and the plots where the retailers are located are determined according to the retailer location information for binding; features are extracted or combined from the retailer related data and the geographic feature data to construct features; features are selected according to priority. Exemplarily, the features selected in this embodiment are in descending order of priority: stall, month, average sales of retailers in the first three months, hourly passenger flow from 0 to 1 o'clock in the whole day, passenger flow in the nine-square grid, the number of passengers whose occupation is commercial service, the number of passengers who have been to high-end restaurants in the nine-square grid, the maximum value of all traffic at 12-13 o'clock in the plot, and the number of commercial service staff in the nine-square grid passenger flow; The eigenvalues ​​of the above features are subjected to feature screening. Exemplarily, stepwise regression method, determination coefficient method and L2 regularization are used for feature screening. The screened feature data are organized into the form of a tensor with a shape of (batch_size, sequence_length, feature_size), where batch_size is the batch size, which indicates the sample size used in one training. The choice of batch size depends on the balance between memory limitation and optimization efficiency; sequence_length is the sequence length, which indicates the time series length of each sample, that is, the time span of the observation or the length of the historical record; feature_size is the feature dimension, which indicates the number of features observed at each time point, such as sales volume, inventory level, price changes, etc.

[0023] The continuous features in the filtered feature data are batch normalized to obtain normalized data, specifically including: Divide the continuous features into several batches of batch data and store them in a data table, call them in array form, and perform batch normalization calculations on each batch of batch data. For example, batch , where each , , is a vector of shape (feature_size), representing the value of a single sample on all feature dimensions, Perform batch normalization calculation on a given batch B. The steps include: Calculate the mean of each feature dimension in batch B: in, is the mean of the feature dimension d, a vector of size feature_size, m is the batch size, that is, the number of samples contained in batch B, is the feature value of the i-th sample in feature dimension d; Calculate the variance of each feature dimension in batch B: in, is the variance of feature dimension d, which is a vector of size feature_size; The features of each sample are normalized using the mean and variance: in, is the eigenvalue of the i-th sample after standardization in the feature dimension d, is a very small constant, such as le-8, which is used to prevent division by zero; Scale and translate the normalized data: in, is the final output value of the i-th sample in the feature dimension d, is the scaling parameter, is the translation parameter, and is a learnable parameter used to restore the expressiveness of the data; By performing batch normalization on the feature data and standardizing the batch data of each batch so that its mean is close to 0 and its variance is close to 1, more accurate and unified feature data values ​​can be obtained, which helps to improve the generalization ability and prediction accuracy of the model.

[0024] The embedding layer is used to map high-dimensional sparse discrete features into low-dimensional dense vector space. It is suitable for processing discrete inputs such as classified data or words in the vocabulary. The embedding layer is used to process the discrete features in the filtered feature data to obtain the embedding layer output data, including: Create a unique integer index for each unique category or word; Initialize the embedding matrix. Each row of the embedding matrix corresponds to the embedding vector of a category or word. These vectors are usually randomly initialized. The discrete values ​​are converted into continuous vectors through the embedding matrix as follows: Among them, index is the index, such as the position of a word in the vocabulary, is the corresponding row extracted from the embedding matrix E according to the index, i.e., the embedding vector; Given a new input value, use the index of the input value to find the corresponding embedding vector; During the training process, in order to minimize the loss function, the gradient of the loss relative to the embedding vector is calculated by the back propagation algorithm, and the embedding matrix is ​​optimized. The optimization formula is as follows: in, To optimize the updated embedding vector, is the learning rate, which determines the size of each optimization update. For loss, is the gradient of the loss with respect to the embedding vector.

[0025] Step S102, inputting the normalized data into a bidirectional LSTM layer, capturing features of the normalized data from forward and backward directions respectively, and merging and outputting feature representations of each time point; Specifically, it includes: obtaining the forward hidden state of the normalized data through the forward LSTM; obtaining the backward hidden state of the normalized data through the backward LSTM; merging the forward hidden state and the backward hidden state into a bidirectional hidden state through splicing or weighted summation, and outputting the feature representation of each time point, with a shape of (batch_size, sequence_length, 2*hidden_size).

[0026] Step S103, concatenating the output of the bidirectional LSTM layer and the output of the embedding layer to form concatenated data, calculating the attention weight through the attention mechanism, and performing weighted summation on the concatenated data according to the attention weight to obtain a context vector; Specifically, the output of the bidirectional LSTM layer is concatenated with the output of the embedding layer to form concatenated data with a shape of (batch_size, sequence_length, 2*hidden_size+emd_size); Map the concatenated data into query vectors, key vectors, and value vectors. For each output time point t, calculate the correlation score between it and all input time points j: in, is the query vector at time point t, is the transpose of the key vector at time point j, is the dimension of the key vector; Convert the relevance scores to a probability distribution, ensuring that they sum to 1 and that each value is between 0 and 1: in, is the attention weight of the output at time point t to the input at time point j; The context vector is obtained by weighted summing the value vector using the attention weights: in, is the context vector at time point t, is the value vector at time point j.

[0027] Step S104: input the context vector into the fully connected layer to output retail sales forecast data, compare the sales forecast data with the actual sales data, and issue a warning when the sales forecast data is higher than a preset upper threshold or lower than a preset lower threshold.

[0028] Specifically, the context vector is concatenated with the query vector to form a feature representation : The feature is represented Input the fully connected layer and output the feature representation Z: in, is the weight matrix of the fully connected layer, is the bias term of the fully connected layer; Feature Representation Apply the ReLU activation function and output the activated feature representation A: According to the requirements of the specific task, the activated feature representation A passes through one or more fully connected layers and finally generates the prediction result through the output layer: Among them, Softmax() is the activation function that converts the output of the fully connected layer into a probability distribution. is the weight matrix of the output layer, is the bias term of the output layer.

[0029] This embodiment also initializes the weights of all linear layers through Xavier uniform distribution to ensure the stability and convergence speed of the model in the early stage of training. The specific steps are as follows: Compute the boundaries of the uniform distribution of the current linear layer: Among them, a is the interval limit, gain is the scaling factor adjusted according to the activation function, Enter the number of features for the current layer, Output feature count for the current layer; Initialize each element of the weight matrix of the current linear layer to the interval Random value within .

[0030] Through steps S101-S104, the retail sales data prediction model is trained according to the training set of retail data, the model parameters and hyperparameters are adjusted according to the validation set, and the final performance of the model is evaluated according to the test set, so as to obtain a trained retail sales data prediction model, which can effectively predict the retail sales data, compare the obtained sales prediction data with the actual sales data, and issue a warning when the sales prediction data is higher than the preset upper threshold or lower than the preset lower threshold.

[0031] For example, this embodiment uses data from 300 retailers as a test data set, and uses MAPE (Mean Absolute Percentage Error) as a measurement indicator to measure the prediction accuracy. The MAPE calculation formula is as follows: Where n is the number of data points, i.e. the number of retailers. is the lth real sales data value, is the lth sales forecast data value; The error percentages between the sales forecast data values ​​obtained by three different models and the actual sales data values ​​are shown in the following table: No. 1: Using SARIMA (Seasonal Autoregressive Integrated Moving Average Model) to conduct predictive analysis on retail data based on maximum likelihood estimation; No. 2: Use XGBoost (extreme gradient boosting) to perform predictive analysis on retail customer data based on gradient boosting decision trees; No. 3: Use the retail sales data prediction model proposed in the present invention to perform predictive analysis on retail data.

[0032] It can be seen from the above table that the retailer sales data prediction model proposed by the tobacco retailer sales data prediction method provided by the present invention has a smaller prediction result error range under the same conditions, thereby achieving more accurate retailer sales data prediction.

[0033] The present invention provides a method for predicting tobacco retailer sales data, obtains multi-source data such as retailer stalls, location information, geographic features, etc., and constructs a retailer sales data prediction model based on a bidirectional LSTM and an attention mechanism. The model is a time series prediction model, which improves the accuracy and robustness of time series prediction on the basis of data support from multi-faceted, multi-element, and continuous data sources, combines the analysis and processing of continuous data and discrete data, effectively improves the prediction accuracy, realizes more accurate retailer sales data prediction, and realizes effective daily supervision and early warning of retailer sales data.

[0034] refer to Figure 2 Another embodiment of the present invention provides a tobacco retailer sales data prediction device 200, comprising: The first module 201 is configured to obtain sample characteristic data of retail customers and perform characteristic screening, perform batch normalization processing on continuous characteristics in the screened characteristic data to obtain normalized data, and perform embedding layer processing on discrete characteristics in the screened characteristic data to obtain embedding layer output data; The second module 202 is configured to input the normalized data into a bidirectional LSTM layer, capture features of the normalized data from forward and backward directions respectively, and combine and output feature representations of each time point; The third module 203 is configured to concatenate the output of the bidirectional LSTM layer with the output of the embedding layer to form concatenated data, calculate the attention weight through the attention mechanism, and perform weighted summation on the concatenated data according to the attention weight to obtain a context vector; The fourth module 204 is configured to input the context vector into the fully connected layer to output retail sales forecast data, compare the sales forecast data with the actual sales data, and issue an early warning when the sales forecast data is higher than a preset upper threshold or lower than a preset lower threshold.

[0035] It should be noted that the tobacco retailer sales data prediction device 200 provided in this embodiment corresponds to a technical solution that can be used to execute each method embodiment, and its implementation principle and technical effect are similar to the method, which will not be repeated here.

[0036] The above description is only a preferred embodiment of the present invention. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solution formed by a specific combination of the above technical features, but also should cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present invention (but not limited to) to form a technical solution.

Claims

1. A tobacco retailer sales data prediction method, characterized in that: include: Obtain sample characteristic data of retail households and perform feature screening, perform batch normalization processing on continuous features in the screened characteristic data to obtain normalized data, and use embedding layer processing on discrete features in the screened characteristic data to obtain embedding layer output data; Input the normalized data into a bidirectional LSTM layer, capture the features of the normalized data from the forward and backward directions respectively, and combine and output the feature representation of each time point; The output of the bidirectional LSTM layer is concatenated with the output of the embedding layer to form concatenated data, an attention weight is calculated through an attention mechanism, and a context vector is obtained by weighted summing the concatenated data according to the attention weight; The context vector is input into the fully connected layer to output the retail sales forecast data, and the sales forecast data is compared with the actual sales data. When the sales forecast data is higher than a preset upper threshold or lower than a preset lower threshold, an early warning is issued.

2. A tobacco retailer sales data prediction method according to claim 1, characterized in that: The feature screening includes stepwise regression method, determination coefficient method and L2 regularization.

3. A tobacco retailer sales data prediction method according to claim 1, characterized in that: Also includes: The filtered feature data includes batch size, sequence length and feature dimension.

4. A tobacco retailer sales data prediction method according to claim 3, characterized in that: The step of performing batch normalization processing on the continuous features in the filtered feature data to obtain normalized data specifically includes: Dividing the continuous features into batch data of several batches and storing them in a data table; Calculate the mean of each feature dimension in the batch data of the current batch: in, is the mean of feature dimension d, m is the batch size, is the feature value of the i-th sample in feature dimension d; Calculate the variance of each feature dimension in the batch data of the current batch: in, is the variance of feature dimension d; The features of each sample are normalized using the mean and variance: in, is the eigenvalue of the i-th sample after standardization in the feature dimension d, is a constant; Scale and translate the normalized data: in, is the final output value of the i-th sample in the feature dimension d, is the scaling parameter, is the translation parameter.

5. A tobacco retailer sales data prediction method according to claim 4, characterized in that: The method of using an embedding layer to process discrete features in the filtered feature data to obtain embedding layer output data includes: Create a unique integer index for each unique category or word; Initialize an embedding matrix, where each row of the embedding matrix corresponds to an embedding vector of a category or word; Given a new input value, use the index of the input value to find the corresponding embedding vector; During the training process, the embedding matrix is ​​optimized by the back propagation algorithm, and the optimization formula is as follows: Among them, index is the index, is the embedding vector extracted from the embedding matrix E according to the index, To optimize the updated embedding vector, is the learning rate, For loss.

6. A tobacco retailer sales data prediction method according to claim 5, characterized in that: The normalized data is input into a bidirectional LSTM layer, the features of the normalized data are captured from the forward and backward directions respectively, and the feature representation of each time point is outputted in a combined manner, including: Obtaining the forward hidden state of the normalized data through a forward LSTM; Obtaining a backward hidden state of the normalized data through a backward LSTM; The forward hidden state and the backward hidden state are combined into a bidirectional hidden state by concatenation or weighted summation, and the feature representation of each time point is output.

7. A tobacco retailer sales data prediction method according to claim 6, characterized in that: The step of calculating the attention weight by the attention mechanism and performing weighted summation on the spliced ​​data according to the attention weight to obtain the context vector includes: The concatenated data is mapped into a query vector, a key vector, and a value vector. For each output time point t, its correlation score with all input time points j is calculated: in, is the query vector at time point t, is the transpose of the key vector at time point j, is the dimension of the key vector; Convert the relevance scores into probability distributions: in, is the attention weight of the output at time point t to the input at time point j; The context vector is obtained by weighted summing the value vector using the attention weights: in, is the context vector at time point t, is the value vector at time point j.

8. A tobacco retailer sales data prediction method according to claim 7, characterized in that: The step of inputting the context vector into a fully connected layer and outputting retail sales forecast data comprises: Concatenate the context vector and the query vector to form a feature representation : The feature is represented Input the fully connected layer and output the feature representation Z: in, is the weight matrix of the fully connected layer, is the bias term of the fully connected layer; Feature Representation Apply the ReLU activation function and output the activated feature representation A: The activated feature representation A generates a prediction result through the fully connected layer and the output layer: Among them, Softmax() is the activation function that converts the output of the fully connected layer into a probability distribution. is the weight matrix of the output layer, is the bias term of the output layer.

9. A tobacco retailer sales data prediction method according to claim 8, characterized in that: Also includes: Compute the boundaries of the uniform distribution of the current linear layer: Among them, a is the interval limit, gain is the scaling factor adjusted according to the activation function, Enter the number of features for the current layer, Output feature count for the current layer; Initialize each element of the weight matrix of the current linear layer to the interval Random value within .

10. A tobacco retailer sales data prediction device, characterized in that: include: The first module is configured to obtain sample characteristic data of retail households and perform characteristic screening, perform batch normalization processing on continuous characteristics in the screened characteristic data to obtain normalized data, and use an embedding layer to process discrete characteristics in the screened characteristic data to obtain embedding layer output data; The second module is configured to input the normalized data into a bidirectional LSTM layer, capture the features of the normalized data from the forward and backward directions respectively, and merge and output the feature representation of each time point; A third module is configured to concatenate the output of the bidirectional LSTM layer with the output of the embedding layer to form concatenated data, calculate the attention weight through the attention mechanism, and perform weighted summation on the concatenated data according to the attention weight to obtain a context vector; The fourth module is configured to input the context vector into the fully connected layer to output retail sales forecast data, compare the sales forecast data with the actual sales data, and issue an early warning when the sales forecast data is higher than a preset upper threshold or lower than a preset lower threshold.