Inventory Optimization Method and System Based on Deep Learning Model

Through deep learning models, process multi-dimensional inventory data, extract short-term and long-term features, and generate accurate sales forecasts and optimized inventory decisions, solving the shortcomings of traditional methods in data diversity, complexity and dynamics, and achieving more efficient inventory management.

CN119338511BActive Publication Date: 2025-06-24QINSILK COM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411889872.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-06-24
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Traditional inventory optimization methods are difficult to effectively process multi-dimensional data, capture complex demand patterns and adapt to dynamically changing market environments, resulting in a decrease in prediction accuracy and timeliness.

Method used

The inventory optimization method based on the deep learning model is adopted, and the short-term timing characteristics are extracted by obtaining multi-dimensional data (historical sales, product attributes, consumer behavior, etc.), and the timing convolutional neural network is used to extract long-term dependency characteristics, combined with the graph neural network and multi-layer perceptron, and finally the sales prediction sequence is generated through the two-way long and short-term memory network, and the inventory replenishment decision is generated through the dynamic planning model.

Benefits of technology

It improves the accuracy and reliability of sales forecasts, optimizes the safe inventory level, realizes intelligent and refined management of inventory replenishment, and reduces inventory costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119338511B_ABST
    Figure CN119338511B_ABST
Patent Text Reader

Abstract

The present invention provides an inventory optimization method and system based on a deep learning model, which relates to the technical field of deep learning. The method includes obtaining multi-dimensional data of commodities, constructing a dynamic convolution kernel using a graph neural network, and combining a multi-layer perceptron to construct a soft clustering assignment matrix to obtain long-term dependence features that integrate spatial and temporal dimensions. After fusing short-term and long-term features, the fused features are input into a bidirectional long short-term memory network to generate a sales volume prediction sequence, and the prediction confidence interval is calculated to obtain the final sales volume prediction result. Based on the sales volume prediction result, the safety inventory level is calculated, and combined with an inventory cost function, a dynamic programming model based on deep reinforcement learning is used to generate a replenishment decision, and finally the decision is sent to the supply chain execution system. The present invention accurately predicts sales volume through a deep learning model and optimizes inventory management, reduces warehousing costs, out-of-stock costs, and overstock costs, and effectively improves the efficiency of the supply chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to deep learning technology, and in particular to an inventory optimization method and system based on a deep learning model. Background Art

[0002] Traditional inventory optimization methods mainly rely on statistical models, such as the moving average method, exponential smoothing method, and ARIMA model, etc. These methods usually assume that the demand pattern is relatively stable and it is difficult to capture complex demand fluctuations. Especially when facing the influence of factors such as promotions, seasonal changes, and emergencies, the prediction accuracy is often insufficient. In addition, these methods usually only consider historical sales volume data and ignore other important information, such as product attributes, consumer behavior, and external market environment, etc., resulting in a decrease in the reliability of the prediction results.

[0003] The existing technologies have the following defects and deficiencies:

[0004] 1. Traditional inventory optimization methods are difficult to effectively process multi-dimensional data. For example, data such as product attributes and consumer behavior contain rich potential information, but traditional statistical models are difficult to effectively integrate and utilize this information, resulting in limited accuracy of the prediction model.

[0005] 2. Traditional inventory optimization methods are difficult to capture complex demand patterns. The actual demand is often affected by multiple factors, showing complex characteristics such as non-linearity and time-variation. Traditional statistical models are usually based on linear assumptions and are difficult to accurately describe and predict such complex demand patterns.

[0006] 3. Traditional inventory optimization methods are difficult to adapt to the dynamically changing market environment. Changes in the market environment, such as the competition pattern and consumer preferences, etc., will have an important impact on demand. Traditional inventory optimization methods often lack an effective capture and response mechanism for these dynamic factors, resulting in a decrease in the timeliness of the prediction results. Summary of the Invention

[0007] Embodiments of the present invention provide an inventory optimization method and system based on a deep learning model, which can solve the problems in the existing technologies.

[0008] In the first aspect of the embodiments of the present invention,

[0009] An inventory optimization method based on a deep learning model is provided, including:

[0010] Obtain multi-dimensional data corresponding to the target product. The multi-dimensional data includes historical sales data, product attribute data, and consumer behavior data. Among them, the historical sales data includes product code, sales time, sales quantity, sales price, and product inventory. The product attribute data includes product material, product style, product seasonal attribute, and product size information. The consumer behavior data includes page views, favorites, add-to-cart quantities, and repurchase rates; perform time series alignment processing on the multi-dimensional data to construct a unified time series feature matrix; input the time series feature matrix into a time series convolutional neural network model to extract short-term time series features;

[0011] Construct a parameterized dynamic convolution kernel through a graph neural network according to the multi-dimensional data, and construct a soft clustering assignment matrix based on a multi-layer perceptron. Combine residual connections to obtain long-term dependence features that fuse the spatial dimension and the time dimension; perform feature fusion on the short-term time series features and the long-term dependence features to obtain a mixed feature vector; input the mixed feature vector into a bidirectional long short-term memory network to generate a sales prediction sequence; calculate a prediction confidence interval based on the sales prediction sequence to form a final sales prediction result;

[0012] Based on the sales prediction result, calculate the safety inventory level, where the safety inventory level is determined by the upper and lower bounds of the prediction confidence interval; construct an inventory cost function, which includes a weighted combination of three dimensions: warehousing cost, out-of-stock cost, and overstock cost; input the safety inventory level and the inventory cost function into a dynamic programming model to generate an inventory replenishment decision, and the inventory replenishment decision includes the replenishment time point and the replenishment quantity; send the inventory replenishment decision to the supply chain execution system, where the dynamic programming model is constructed based on a deep reinforcement learning model.

[0013] Constructing a parameterized dynamic convolution kernel through a graph neural network according to the multi-dimensional data, and constructing a soft clustering assignment matrix based on a multi-layer perceptron. Combining residual connections to obtain long-term dependence features that fuse the spatial dimension and the time dimension includes:

[0014] Perform normalization processing on the input multi-dimensional data to obtain a normalized feature matrix. Use each data in the multi-dimensional data as a graph node, calculate the similarity between graph nodes based on the normalized feature matrix to construct an adaptive adjacency matrix, and perform graph convolution operation on the adaptive adjacency matrix and the normalized feature matrix to obtain initial node features;

[0015] Extract time encoding information from the initial node features, perform splicing operation on the time encoding information and the initial node features and input them into a multi-layer perceptron network to generate dynamic convolution kernel parameters, and perform adaptive convolution operation on the initial node features based on the dynamic convolution kernel parameters to obtain dynamic time series features;

[0016] Construct a feature difference vector between node pairs based on the dynamic temporal features, input the feature difference vector into a multi-layer perceptron to calculate a node similarity matrix, perform softmax normalization on the node similarity matrix to obtain a soft clustering assignment matrix, and use the soft clustering assignment matrix to perform weighted aggregation on the dynamic temporal features to obtain clustering center features;

[0017] Fuse the clustering center features with the dynamic temporal features to obtain fused features, input the fused features into a gated network to generate adaptive gating weights, and perform selective information transmission on the fused features based on the adaptive gating weights to obtain gated output features;

[0018] Input the gated output features into three independent linear mapping layers respectively to obtain a query matrix, a key matrix, and a value matrix, calculate attention scores based on the query matrix and the key matrix, and multiply the attention scores by a time decay factor to obtain temporal attention weights;

[0019] Perform weighted aggregation on the value matrix based on the temporal attention weights to obtain multi-head attention features, input the multi-head attention features into a feed-forward neural network for non-linear transformation to obtain transformed features, perform residual connection on the transformed features and the multi-head attention features and perform layer normalization processing to obtain long-term dependence features with long-term dependence relationships.

[0020] Input the mixed feature vector into a bidirectional long short-term memory network to generate a sales volume prediction sequence; calculate a prediction confidence interval based on the sales volume prediction sequence to form a final sales volume prediction result including:

[0021] Input the mixed feature vector into the forward network of the bidirectional long short-term memory network to obtain forward hidden state features, input the mixed feature vector into the backward network of the bidirectional long short-term memory network to obtain backward hidden state features, and concatenate the forward hidden state features and the backward hidden state features to obtain bidirectional hidden state features;

[0022] Construct an attention weight calculation unit based on the bidirectional hidden state features. The attention weight calculation unit includes a historical state mapping layer and a current state mapping layer. Input the historical bidirectional hidden state features and the bidirectional hidden state features at the current moment into the historical state mapping layer and the current state mapping layer respectively, and obtain temporal attention weights through attention scoring; perform weighted summation on the bidirectional hidden state features at the current moment using the temporal attention weights to obtain a context vector, and fuse the context vector with the bidirectional hidden state features at the current moment to obtain an enhanced feature representation;

[0023] Input the enhanced feature representation into a prediction distribution generation network, which includes a mean prediction layer and a variance prediction layer. The mean prediction layer outputs a predicted mean, and the variance prediction layer outputs a predicted variance; perform Monte Carlo sampling based on the predicted mean and the predicted variance to obtain multiple groups of prediction samples, and calculate the statistical distribution characteristics of the multiple groups of prediction samples to obtain a sales volume prediction sequence;

[0024] Adaptive adjustment is performed on the prediction interval, and the predicted mean and the predicted variance are dynamically updated based on the deviation degree between the actual sales volume and the predicted mean to obtain the upper and lower bounds of the calibrated prediction confidence interval. The sales volume prediction sequence and the upper and lower bounds of the prediction confidence interval form the final sales volume prediction result.

[0025] Based on the sales volume prediction result, calculate the safety inventory level, where the safety inventory level is determined by the upper and lower bounds of the prediction confidence interval and includes:

[0026] Calculate the standard deviation of the prediction interval based on the upper bound and the lower bound of the prediction confidence interval of the sales volume prediction sequence, and combine the standard deviation of the prediction interval with the service level coefficient, the replenishment lead time, and the review period to obtain the initial safety inventory level;

[0027] Calculate the first absolute value of the difference between the upper bound of the prediction confidence interval and the historical maximum predicted value, and the second absolute value of the difference between the historical minimum predicted value and the lower bound of the prediction confidence interval. Select the ratio of the larger value of the first absolute value and the second absolute value to the standard deviation of the prediction interval as the demand fluctuation adjustment factor;

[0028] Take the product of the demand fluctuation adjustment factor and the pre-obtained sensitivity parameter as the dynamic adjustment amount, and combine the initial safety inventory level with the dynamic adjustment amount to obtain the final safety inventory level.

[0029] Construct an inventory cost function, which includes a weighted combination of three dimensions: warehousing cost, stockout cost, and overstock cost, including:

[0030] Construct a warehousing cost function, which is the product of the unit warehousing cost and the current inventory level; construct a stockout cost function, which is the product of the unit stockout cost and the inventory gap; construct an overstock cost function, which is the product of the unit overstock cost and the inventory quantity exceeding the upper bound of the prediction confidence interval;

[0031] Initial weight coefficients are assigned to the warehousing cost function, the stock - out cost function, and the backlog cost function respectively. The weight gradients are calculated based on the contributions of each cost function to the total cost, and the weight coefficients are updated according to the weight gradients and the learning rate. After exponential transformation and normalization of the updated weight coefficients, the final weight coefficients are obtained, and an inventory cost function is constructed by the weighted combination of the final weight coefficients and the corresponding cost functions.

[0032] The safety inventory level and the inventory cost function are input into a dynamic programming model to generate an inventory replenishment decision. The inventory replenishment decision includes the replenishment time point and the replenishment quantity, including:

[0033] A state vector space is constructed, and the state vector space includes the safety inventory levels of each category and the in - transit order quantities. The inventory cost function and the expected value of the future value function adjusted by the discount factor are added together to construct a dynamic programming value function. A state transition equation is constructed based on the state vector space. The state transition equation calculates the inventory status of the next period according to the current inventory level, in - transit order quantity, actual demand quantity, and replenishment quantity of each category, and updates the in - transit order status according to the replenishment lead time of each category.

[0034] The state transition equation is substituted into the dynamic programming value function, and the dynamic programming value function is solved recursively to obtain the optimal value function. The reorder point of each category is calculated based on the optimal value function. The sum of the safety inventory level and the in - transit order quantity of each category is compared with the corresponding reorder point. When the sum of the safety inventory level and the in - transit order quantity is lower than the reorder point, the replenishment time point is determined.

[0035] The resource occupancy of each category and the total available resources are obtained to construct a resource constraint condition, and the unit procurement cost of each category and the available budget are obtained to construct a budget constraint condition. The economic order quantity is calculated based on the fixed order cost, average demand rate, and unit holding cost of each category. On the premise of meeting the resource constraint condition and the budget constraint condition, the economic order quantity is used as the replenishment quantity.

[0036] The state transition equation is substituted into the dynamic programming value function, and the dynamic programming value function is solved recursively to obtain the optimal value function, including:

[0037] ;

[0038] where represents the optimal value function at time t in state S t and A t represents the action at time t, represents the inventory cost function at time t, represents the discount factor, which is used to balance the current cost and future costs, represents the expected value of the value function at time t + 1;

[0039] The inventory cost function is shown in the following formula:

[0040] ;

[0041] where, represents the warehousing cost function, represents the stock - out cost function, represents the overstock cost function, 、 、 respectively represent the final weight coefficients corresponding to the warehousing cost function, the stock - out cost function, and the overstock cost function, represents the inventory level at time t;

[0042] The state transition equation is shown in the following formula:

[0043] ;

[0044] ;

[0045] ;

[0046] where, represents the demand of category i at time t, represents the replenishment lead time of category i at time t + 1, represents the replenishment lead time of category N at time t + 1, represents the in - transit order quantity of category i at time t, represents the order quantity of category i at time t, represents the transpose;

[0047] The reorder point calculation formula is shown as follows:

[0048] ;

[0049] where, represents the optimal reorder point of category i, represents the average demand rate of category i, represents the standard normal distribution quantile corresponding to the service level, represents the standard deviation of the demand of category i.

[0050] In the second aspect of the embodiments of the present invention, an inventory optimization system based on a deep learning model is provided, including:

[0051] The first unit is used to obtain multi-dimensional data corresponding to the target commodity. The multi-dimensional data includes historical sales data, commodity attribute data, and consumer behavior data. The historical sales data includes commodity code, sales time, sales quantity, sales price, and commodity inventory. The commodity attribute data includes commodity material, commodity style, commodity seasonal attribute, and commodity size information. The consumer behavior data includes page views, favorites, add-to-cart quantities, and repurchase rates. Perform time series alignment processing on the multi-dimensional data to construct a unified time series feature matrix. Input the time series feature matrix into a time series convolutional neural network model to extract short-term time series features.

[0052] The second unit is used to construct a parameterized dynamic convolution kernel through a graph neural network based on the multi-dimensional data, and construct a soft clustering assignment matrix based on a multi-layer perceptron. Combine residual connections to obtain long-term dependence features that integrate spatial and temporal dimensions. Perform feature fusion on the short-term time series features and the long-term dependence features to obtain a mixed feature vector. Input the mixed feature vector into a bidirectional long short-term memory network to generate a sales prediction sequence. Calculate a prediction confidence interval based on the sales prediction sequence to form a final sales prediction result.

[0053] The third unit is used to calculate the safety inventory level based on the sales prediction result, where the safety inventory level is determined by the upper and lower bounds of the prediction confidence interval. Construct an inventory cost function, which includes a weighted combination of three dimensions: warehousing cost, stock-out cost, and overstock cost. Input the safety inventory level and the inventory cost function into a dynamic programming model to generate an inventory replenishment decision, where the inventory replenishment decision includes the replenishment time point and the replenishment quantity. Send the inventory replenishment decision to the supply chain execution system, where the dynamic programming model is constructed based on a deep reinforcement learning model.

[0054] The beneficial effects of this application are as follows:

[0055] 1. Improve the accuracy of sales prediction: The present invention uses a time series convolutional neural network to extract short-term time series features, combines a graph neural network and a multi-layer perceptron to extract long-term dependence features, and finally generates a sales prediction sequence through a bidirectional long short-term memory network, which can more comprehensively capture the changing trend of commodity sales, and evaluate the reliability of the prediction result through the prediction confidence interval, thereby improving the accuracy of sales prediction.

[0056] 2. Optimize the safety inventory level: The present invention determines the safety inventory level based on the prediction confidence interval, avoiding the limitations of setting the safety inventory based on empirical values or simple statistical methods in traditional methods, and can more accurately determine the safety inventory level, reducing the inventory cost while meeting the sales demand.

[0057] 3. Intelligent inventory replenishment decision-making: The present invention constructs an inventory cost function that includes warehousing costs, out-of-stock costs, and overstock costs, and combines a deep reinforcement learning model to construct a dynamic programming model, which can automatically generate an optimal replenishment strategy according to the predicted sales volume and cost function, including the replenishment time point and the replenishment quantity, so as to achieve intelligent and refined management of inventory replenishment. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a schematic flowchart of an inventory optimization method based on a deep learning model according to an embodiment of the present invention;

[0059] Figure 2 is a schematic structural diagram of an inventory optimization system based on a deep learning model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0061] The technical solutions of the present invention will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0062] Figure 1 is a schematic flowchart of an inventory optimization method based on a deep learning model according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0063] S101. Obtain multi-dimensional data corresponding to the target commodity. The multi-dimensional data includes historical sales volume data, commodity attribute data, and consumer behavior data. The historical sales volume data includes commodity code, sales time, sales quantity, sales price, and commodity inventory. The commodity attribute data includes commodity material, commodity style, commodity seasonal attribute, and commodity size information. The consumer behavior data includes page views, favorites, add-to-cart quantities, and repurchase rates; perform time series alignment processing on the multi-dimensional data to construct a unified time series feature matrix; input the time series feature matrix into a time series convolutional neural network model to extract short-term time series features;

[0064] S102. Construct a parameterized dynamic convolution kernel through a graph neural network based on the multi-dimensional data, and construct a soft clustering assignment matrix based on a multi-layer perceptron. Combine residual connections to obtain long-term dependence features that fuse the spatial dimension and the time dimension; perform feature fusion on the short-term time series features and the long-term dependence features to obtain a mixed feature vector; input the mixed feature vector into a bidirectional long short-term memory network to generate a sales volume prediction sequence; calculate a prediction confidence interval based on the sales volume prediction sequence to form a final sales volume prediction result;

[0065] S103. Based on the sales volume prediction result, calculate the safety inventory level, where the safety inventory level is determined by the upper and lower bounds of the prediction confidence interval; construct an inventory cost function, which includes a weighted combination of three dimensions: warehousing cost, stock-out cost, and overstock cost; input the safety inventory level and the inventory cost function into a dynamic programming model to generate an inventory replenishment decision, where the inventory replenishment decision includes the replenishment time point and the replenishment quantity; send the inventory replenishment decision to the supply chain execution system, where the dynamic programming model is constructed based on a deep reinforcement learning model.

[0066] In an alternative embodiment, constructing a parameterized dynamic convolution kernel through a graph neural network based on the multi-dimensional data, and constructing a soft clustering assignment matrix based on a multi-layer perceptron. Combining residual connections to obtain long-term dependence features that fuse the spatial dimension and the time dimension includes:

[0067] Perform normalization processing on the input multi-dimensional data to obtain a normalized feature matrix. Use each data in the multi-dimensional data as a graph node, calculate the similarity between graph nodes based on the normalized feature matrix to construct an adaptive adjacency matrix, and perform graph convolution operation on the adaptive adjacency matrix and the normalized feature matrix to obtain initial node features;

[0068] Extract time encoding information from the initial node features, perform splicing operation on the time encoding information and the initial node features and input them into a multi-layer perceptron network to generate dynamic convolution kernel parameters, and perform adaptive convolution operation on the initial node features based on the dynamic convolution kernel parameters to obtain dynamic time series features;

[0069] Construct a feature difference vector between node pairs based on the dynamic time series features, input the feature difference vector into a multi-layer perceptron to calculate a node similarity matrix, perform softmax normalization on the node similarity matrix to obtain a soft clustering assignment matrix, and use the soft clustering assignment matrix to perform weighted aggregation on the dynamic time series features to obtain clustering center features;

[0070] Fuse the clustering center features and the dynamic time series features to obtain fused features, input the fused features into a gating network to generate adaptive gating weights, and perform selective information transmission on the fused features based on the adaptive gating weights to obtain gated output features;

[0071] Input the gated output features into three independent linear mapping layers respectively to obtain a query matrix, a key matrix, and a value matrix, calculate attention scores based on the query matrix and the key matrix, and multiply the attention scores by a time decay factor to obtain time series attention weights;

[0072] Perform weighted aggregation on the value matrix based on the time series attention weights to obtain multi-head attention features, input the multi-head attention features into a feed-forward neural network for non-linear transformation to obtain transformed features, perform residual connection on the transformed features and the multi-head attention features, and perform layer normalization processing to obtain long-term dependence features with long-term dependence relationships.

[0073] A method for extracting long-term dependence features based on graph neural networks and soft clustering assignment matrices, which is used to process multi-dimensional time series data and can effectively capture long-term dependence relationships in spatial and temporal dimensions. The specific steps of this method are as follows:

[0074] First, perform normalization processing on the input multi-dimensional data. For example, assume that the input data is a matrix containing 10 time steps, 20 sensors at each time step, and each sensor collects 3-dimensional data. Min-Max normalization can be performed on the data of each dimension of each sensor to scale the data range to between 0 and 1, obtaining a normalized feature matrix of 10x20x3.

[0075] Next, construct an adaptive adjacency matrix based on the normalized feature matrix. Treat each data point in the multi-dimensional data as a node in the graph, and calculate the similarity between nodes. For example, the Euclidean distance can be used to measure the similarity between nodes. The smaller the distance, the higher the similarity. The calculated similarity values are used to construct an adaptive adjacency matrix, representing the connection relationship between nodes.

[0076] Then, perform graph convolution operations on the adaptive adjacency matrix and the normalized feature matrix to obtain initial node features. Graph convolution operations can be understood as performing weighted averaging on the features of the node itself and the features of its neighbor nodes.

[0077] Extract time encoding information from the initial node features. For example, sine and cosine functions can be used to generate time encoding vectors to incorporate time information into the node features. Concatenate the time encoding information with the initial node features to obtain new features containing time information.

[0078] Input the concatenated features into a multi-layer perceptron network to generate dynamic convolutional kernel parameters. The multi-layer perceptron network can learn the implicit patterns in the data and generate appropriate convolutional kernel parameters based on these patterns. Perform adaptive convolution operations on the initial node features based on the dynamic convolutional kernel parameters to obtain dynamic temporal features.

[0079] Construct a feature difference vector between nodes based on the dynamic temporal features. For example, the difference between two node features can be calculated to obtain the feature difference vector. Input the feature difference vector into a multi-layer perceptron to calculate and obtain the node similarity matrix. Perform Softmax normalization on the node similarity matrix to obtain the soft clustering assignment matrix.

[0080] Use the soft clustering assignment matrix to perform weighted aggregation on the dynamic temporal features to obtain the cluster center features. The soft clustering assignment matrix can be regarded as the probability of each node belonging to each cluster.

[0081] Fuse the cluster center features and the dynamic temporal features to obtain the fused features. For example, the two can be concatenated. Input the fused features into a gating network to generate adaptive gating weights. The gating network can control the flow of information and selectively transmit important information. Perform selective information transmission on the fused features based on the adaptive gating weights to obtain the gated output features.

[0082] Input the gated output features into three independent linear mapping layers respectively to obtain the query matrix, key matrix, and value matrix. Calculate the attention scores based on the query matrix and the key matrix. Multiply the attention scores by the time decay factor to obtain the temporal attention weights. The time decay factor can assign higher weights to nodes in the more recent time steps.

[0083] Perform weighted aggregation on the value matrix based on the temporal attention weights to obtain the multi-head attention features. Input the multi-head attention features into a feed-forward neural network for non-linear transformation to obtain the transformed features. Perform residual connection on the transformed features and the multi-head attention features and perform layer normalization to obtain the long-term dependence features with long-term dependence relationships.

[0084] The beneficial effects of this method are reflected in the following three aspects:

[0085] 1. Improve feature expression ability: Through graph convolution operations and the soft clustering assignment matrix, it is able to effectively capture the dependence relationships of the data in the spatial and temporal dimensions, thereby improving the feature expression ability. For example, in the traffic flow prediction task, it can better capture the spatio-temporal patterns of the road network structure and traffic flow.

[0086] 2. Enhancing the robustness of the model: The dynamic convolution kernel and the adaptive gating mechanism enable the model to better adapt to different data distributions and noise interferences, thereby enhancing the robustness of the model. For example, in the case of abnormal sensor data, it can still maintain good prediction performance.

[0087] 3. Capturing long-term dependencies: By combining temporal encoding information and the temporal attention mechanism, long-term dependencies in the data can be effectively captured, thus improving the model's prediction ability for long-term trends. For example, in the stock price prediction task, it can better capture market trends and periodic fluctuations.

[0088] In an alternative implementation, inputting the mixed feature vector into a bidirectional long short-term memory network to generate a sales volume prediction sequence; calculating a prediction confidence interval based on the sales volume prediction sequence to form the final sales volume prediction result includes:

[0089] Inputting the mixed feature vector into the forward network of the bidirectional long short-term memory network to obtain a forward hidden state feature, inputting the mixed feature vector into the backward network of the bidirectional long short-term memory network to obtain a backward hidden state feature, and splicing the forward hidden state feature and the backward hidden state feature to obtain a bidirectional hidden state feature;

[0090] Constructing an attention weight calculation unit based on the bidirectional hidden state feature, the attention weight calculation unit includes a historical state mapping layer and a current state mapping layer, inputting the historical bidirectional hidden state feature and the bidirectional hidden state feature at the current moment into the historical state mapping layer and the current state mapping layer respectively, obtaining temporal attention weights through attention scoring; using the temporal attention weights to perform weighted summation on the bidirectional hidden state feature at the current moment to obtain a context vector, and performing feature fusion on the context vector and the bidirectional hidden state feature at the current moment to obtain an enhanced feature representation;

[0091] Inputting the enhanced feature representation into a prediction distribution generation network, the prediction distribution generation network includes a mean prediction layer and a variance prediction layer, the mean prediction layer outputs a prediction mean, and the variance prediction layer outputs a prediction variance; performing Monte Carlo sampling based on the prediction mean and the prediction variance to obtain multiple groups of prediction samples, and calculating the statistical distribution characteristics of the multiple groups of prediction samples to obtain a sales volume prediction sequence;

[0092] Adapting the prediction interval, dynamically updating the prediction mean and the prediction variance based on the deviation degree between the actual sales volume and the prediction mean to obtain the upper and lower bounds of the calibrated prediction confidence interval, and combining the sales volume prediction sequence with the upper and lower bounds of the prediction confidence interval to form the final sales volume prediction result.

[0093] A commodity sales prediction method based on bidirectional long short-term memory network and attention mechanism can effectively improve the accuracy and reliability of sales prediction. This method comprehensively considers various factors affecting sales, such as historical sales data, seasonal factors, promotional activities, etc., and uses the attention mechanism to capture the importance of different time steps, and finally generates a prediction result with a confidence interval.

[0094] First, extract and process various features affecting commodity sales. For example, historical sales data, commodity prices, promotional activity information, seasonal factors, holiday information, competitor information, macroeconomic indicators, etc. These features can be selected and adjusted according to specific scenarios. Then, standardize or normalize these different types of features to eliminate the dimensional differences between different features. Finally, splice these processed features together to form a mixed feature vector. For example, the sales volume of a certain commodity in the past 7 days is [10, 12, 15, 13, 16, 18, 20] respectively, the commodity price is 50 yuan, there is a recent promotional activity, the season is summer, and the holiday is a non-holiday. After numerical and standardized processing of these features, they are spliced into a mixed feature vector.

[0095] Next, input the mixed feature vector into the bidirectional long short-term memory network. The bidirectional long short-term memory network contains a forward network and a backward network, which learn time series features from the forward and backward directions respectively. Input the mixed feature vector into the forward network to obtain a forward hidden state feature sequence; input the mixed feature vector into the backward network to obtain a backward hidden state feature sequence. Then, splice the forward hidden state feature and the backward hidden state feature at the corresponding time step to obtain a bidirectional hidden state feature sequence. For example, input the above mixed feature vector into the bidirectional long short-term memory network to obtain 7 time steps of bidirectional hidden state features, and the dimension of each feature vector is 256.

[0096] Then, construct an attention weight calculation unit based on the bidirectional hidden state features. This unit contains a historical state mapping layer and a current state mapping layer. Input the historical bidirectional hidden state features and the bidirectional hidden state features at the current moment into these two mapping layers respectively. By calculating the similarity score between the bidirectional hidden state features at the current moment and the bidirectional hidden state features at each historical moment, the temporal attention weight is obtained. For example, calculate the similarity score between the bidirectional hidden state features at the 7th time step and the bidirectional hidden state features at the previous 6 time steps to obtain a 6-dimensional attention weight vector.

[0097] The obtained temporal attention weights are used to perform weighted summation on the historical bidirectional hidden state features to obtain a context vector. This context vector integrates the contributions of historical information and highlights the historical information that is important for the prediction at the current moment. Then, the context vector and the bidirectional hidden state features at the current moment are feature fused, for example, by concatenating the two, to obtain an enhanced feature representation.

[0098] The enhanced feature representation is input into a prediction distribution generation network, which includes a mean prediction layer and a variance prediction layer. The mean prediction layer outputs a predicted mean, and the variance prediction layer outputs a predicted variance. For example, the enhanced feature representation at the 7th time step is input into the prediction distribution generation network, and the predicted mean is 22 and the predicted variance is 4.

[0099] Monte Carlo sampling is performed based on the predicted mean and the predicted variance to obtain multiple groups of prediction samples. For example, sampling is performed according to a normal distribution with a mean of 22 and a variance of 4 to obtain 1000 prediction samples. The statistical distribution characteristics of these prediction samples are calculated, such as calculating the average value, to obtain a sales volume prediction sequence.

[0100] Finally, the prediction interval is adaptively adjusted. According to the deviation degree between the actual sales volume and the predicted mean, the predicted mean and the predicted variance are dynamically updated to obtain the upper and lower bounds of the calibrated prediction confidence interval. For example, if the actual sales volume continuously exceeds the predicted mean, the predicted mean and the predicted variance are appropriately increased. The sales volume prediction sequence and the upper and lower bounds of the prediction confidence interval form the final sales volume prediction result. For example, the final prediction result is: the sales volume prediction sequence for the next 7 days is [22, 23, 25, 24, 26, 28, 30], and the upper and lower bounds of the confidence interval are [20, 24], [21, 25], [23, 27], [22, 26], [24, 28], [26, 30], [28, 32].

[0101] The beneficial effects of this method are reflected in the following three aspects:

[0102] 1. Improved prediction accuracy: By fusing multiple features and using the attention mechanism, the changing trend of commodity sales volume can be captured more accurately, improving the prediction accuracy.

[0103] 2. Enhanced reliability: Through the prediction confidence interval, the uncertainty of the prediction result can be quantified, providing more reliable decision-making support.

[0104] 3. Strong adaptability: By dynamically adjusting the prediction interval, it can adapt to the changes of different commodities and different sales environments, improving the generalization ability of the model.

[0105] In an alternative embodiment, based on the sales volume prediction result, a safety stock level is calculated, where the safety stock level is determined by the upper and lower bounds of the prediction confidence interval, including:

[0106] Calculate the standard deviation of the prediction interval based on the upper bound and the lower bound of the prediction confidence interval of the sales volume prediction sequence, and combine the standard deviation of the prediction interval with the service level coefficient, the replenishment lead time, and the review period to obtain the initial safety stock level;

[0107] Calculate the first absolute value of the difference between the upper bound of the prediction confidence interval and the historical maximum predicted value, and the second absolute value of the difference between the historical minimum predicted value and the lower bound of the prediction confidence interval, and select the ratio of the larger value of the first absolute value and the second absolute value to the standard deviation of the prediction interval as the demand fluctuation adjustment factor;

[0108] Take the product of the demand fluctuation adjustment factor and the pre-obtained sensitivity parameter as the dynamic adjustment amount, and combine the initial safety stock level with the dynamic adjustment amount to obtain the final safety stock level.

[0109] The safety stock calculation method is used to cope with the uncertainty of sales volume prediction and improve the stability of the supply chain. The core of this method is to use the prediction confidence interval, combine historical data and sensitivity parameters, and dynamically adjust the safety stock level to better cope with demand fluctuations.

[0110] First, obtain the sales volume prediction sequence and the corresponding upper and lower bounds of the prediction confidence interval. For example, through time series analysis method, predict the daily sales volume in the next week, and obtain the prediction confidence interval of the daily sales volume. Assume that the predicted daily sales volume of a certain product in the next week is: [100, 110, 120, 130, 140, 150, 160], the corresponding upper bound of the prediction confidence interval is: [110, 121, 132, 143, 154, 165, 176], and the lower bound is: [90, 99, 108, 117, 126, 135, 144].

[0111] Next, calculate the standard deviation of the prediction interval. Taking the above prediction interval as an example, the fluctuation range of the daily predicted sales volume can be calculated. For example, the fluctuation range on the first day is 110 - 90 = 20, and so on, to obtain the fluctuation range for a week. Then, based on these fluctuation ranges, calculate the standard deviation of the prediction interval. Assume the calculation result is 10.

[0112] Then, considering the service level coefficient, replenishment lead time, and review period, calculate the initial safety stock level. Assume the service level coefficient is 1.65 (corresponding to a 95% service level), the replenishment lead time is 3 days, and the review period is 1 day. Then, the initial safety stock level can be set as Forecast Interval Standard Deviation * Service Level Coefficient * sqrt(Replenishment Lead Time / Review Period) = 10 * 1.65 * sqrt(3 / 1) ≈ 28.6. Here, sqrt represents the square root.

[0113] Next, calculate the demand volatility adjustment factor. First, obtain the historical maximum forecast value and the historical minimum forecast value. Assume the historical maximum forecast value is 200 and the historical minimum forecast value is 50. Then, calculate the absolute value of the difference between the upper bound of the forecast confidence interval and the historical maximum forecast value, and the absolute value of the difference between the historical minimum forecast value and the lower bound of the forecast confidence interval respectively. Taking the upper bound of the above-mentioned forecast interval as an example, the difference on the first day is |200 - 110| = 90, and so on to get the differences for a week. Similarly, the differences between the historical minimum forecast value and the lower bound can be calculated. Then, select the maximum value among these differences, assume it is 100. Finally, take the ratio of this maximum value to the forecast interval standard deviation as the demand volatility adjustment factor, that is, 100 / 10 = 10.

[0114] After that, calculate the dynamic adjustment amount. Multiply the demand volatility adjustment factor by the pre-obtained sensitivity parameter to get the dynamic adjustment amount. Assume the sensitivity parameter is 0.2, then the dynamic adjustment amount is 10 * 0.2 = 2.

[0115] Finally, add the initial safety stock level and the dynamic adjustment amount to get the final safety stock level. That is, 28.6 + 2 = 30.6. Therefore, the final safety stock level of this product is set to 31.

[0116] The beneficial effects of this method can be summarized in the following three aspects:

[0117] Improve the accuracy of inventory management: By combining the forecast confidence interval and historical data, estimate demand volatility more accurately, thus avoiding excessive or insufficient inventory levels.

[0118] Enhance the resilience of the supply chain: Dynamically adjust the safety stock level, which can better respond to market changes and emergencies, and improve the risk resistance ability of the supply chain.

[0119] Reduce inventory costs: On the premise of ensuring the service level, avoid unnecessary inventory backlogs, thereby reducing inventory holding costs.

[0120] In an alternative embodiment, an inventory cost function is constructed. The inventory cost function includes a weighted combination of three dimensions: warehousing cost, stock-out cost, and overstock cost, including:

[0121] Construct a warehousing cost function, which is the product of the unit warehousing cost and the current inventory level; construct a stock-out cost function, which is the product of the unit stock-out cost and the inventory gap; construct an overstock cost function, which is the product of the unit overstock cost and the inventory quantity exceeding the upper bound of the prediction confidence interval.

[0122] Assign initial weight coefficients to the warehousing cost function, the stock-out cost function, and the overstock cost function respectively. Calculate the weight gradient based on the contribution of each cost function to the total cost, and update the weight coefficients according to the weight gradient and the learning rate; perform exponential transformation and normalization on the updated weight coefficients to obtain the final weight coefficients, and construct an inventory cost function with the weighted combination of the final weight coefficients and the corresponding cost functions.

[0123] An inventory cost optimization method that realizes the minimization of inventory cost by dynamically adjusting the weights of warehousing cost, stock-out cost, and overstock cost.

[0124] First, collect historical inventory data, including daily inventory levels, sales data, forecast data, and forecast confidence intervals, etc. For example, collect data for the past year, recording the beginning inventory, sales volume, forecasted sales volume, and the upper and lower bounds of the forecast confidence interval for each day.

[0125] Next, construct a warehousing cost function, a stock-out cost function, and an overstock cost function respectively. The calculation method of the warehousing cost function is: the unit warehousing cost multiplied by the current inventory level. For example, if the unit warehousing cost is 0.5 yuan per item per day and the current inventory level is 100 items, then the warehousing cost is 50 yuan. The calculation method of the stock-out cost function is: the unit stock-out cost multiplied by the inventory gap. The inventory gap is the difference between the demand quantity and the inventory quantity. If this difference is positive, it means a stock-out, otherwise it is 0. For example, if the unit stock-out cost is 2 yuan per item per day, the daily demand quantity is 120 items, and the inventory is only 100 items, then the stock-out cost is 2 * (120 - 100) = 40 yuan. The calculation method of the overstock cost function is: the unit overstock cost multiplied by the inventory quantity exceeding the upper bound of the prediction confidence interval. The inventory quantity exceeding the upper bound of the prediction confidence interval is the part where the actual inventory quantity exceeds the upper bound of the prediction confidence interval. For example, if the unit overstock cost is 0.2 yuan per item per day, the daily inventory quantity is 150 items, and the upper bound of the prediction confidence interval is 130 items, then the overstock cost is 0.2 * (150 - 130) = 4 yuan.

[0126] Then, initial weight coefficients are assigned to the warehousing cost function, the stock - out cost function, and the overstock cost function respectively. For example, the initial weight coefficients are set to 0.3, 0.5, and 0.2 respectively.

[0127] Next, calculate the contribution of each cost function to the total cost and calculate the weight gradient based on this. For example, assume that on a certain day, the warehousing cost is 50 yuan, the stock - out cost is 40 yuan, and the overstock cost is 4 yuan. Then the total cost is 50 * 0.3 + 40 * 0.5 + 4 * 0.2 = 35.8 yuan. Then, according to the value of each cost function and the total cost, calculate the weight gradient. For example, the gradient of the warehousing cost weight can be approximately calculated as the difference between the proportion of the warehousing cost in the total cost and the current weight, and so on.

[0128] Update the weight coefficients according to the calculated weight gradient and the preset learning rate. The learning rate controls the speed of weight update, for example, set to 0.01. The new weight coefficient is equal to the original weight coefficient plus the learning rate multiplied by the weight gradient.

[0129] Perform exponential transformation on the updated weight coefficients and then normalize them to obtain the final weight coefficients. Exponential transformation can amplify the differences in weight coefficients, and normalization can ensure that the sum of weight coefficients is 1.

[0130] Finally, perform weighted combination of the final weight coefficients and the corresponding cost functions to construct the final inventory cost function. For example, if the final weight coefficients are 0.25, 0.6, and 0.15 respectively, the final inventory cost function is 0.25 * warehousing cost+0.6 * stock - out cost + 0.15 * overstock cost.

[0131] The beneficial effects of this inventory cost optimization method are reflected in three aspects:

[0132] First, improve inventory management efficiency. By dynamically adjusting the weights of cost functions, the importance of various costs in different situations can be more accurately reflected, so as to formulate more reasonable inventory strategies.

[0133] Second, reduce inventory costs. This method can effectively balance the warehousing cost, the stock - out cost, and the overstock cost, so as to minimize the total inventory cost.

[0134] Third, enhance the competitiveness of enterprises. By optimizing inventory management, operating costs can be reduced, customer satisfaction can be improved, and ultimately the competitiveness of enterprises can be enhanced.

[0135] In an alternative implementation, input the safety inventory level and the inventory cost function into a dynamic programming model to generate an inventory replenishment decision, and the inventory replenishment decision includes the replenishment time point and the replenishment quantity, including:

[0136] Construct a state vector space, where the state vector space includes the safety inventory levels of each category and the quantity of orders in transit; add the inventory cost function to the expected value of the future value function adjusted by a discount factor to construct a dynamic programming value function; construct a state transition equation based on the state vector space, and the state transition equation calculates the inventory status of the next time period according to the current inventory level, the quantity of orders in transit, the actual demand quantity, and the replenishment quantity of each category, and updates the status of orders in transit according to the replenishment lead time of each category;

[0137] Substitute the state transition equation into the dynamic programming value function, solve the dynamic programming value function recursively, and obtain the optimal value function; calculate the reorder point of each category based on the optimal value function, compare the sum of the safety inventory level and the quantity of orders in transit of each category with the corresponding reorder point, and determine the replenishment time point when the sum of the safety inventory level and the quantity of orders in transit is lower than the reorder point;

[0138] Obtain the resource occupancy of each category and the total available resources to construct a resource constraint condition, and obtain the unit procurement cost of each category and the available budget to construct a budget constraint condition; calculate the economic order quantity based on the fixed order cost, average demand rate, and unit holding cost of each category, and use the economic order quantity as the replenishment quantity on the premise of satisfying the resource constraint condition and the budget constraint condition.

[0139] An inventory replenishment decision-making method based on dynamic programming is used to determine the replenishment time points and replenishment quantities of multiple categories to minimize inventory costs while satisfying resource and budget constraints. The core of this method is to construct a dynamic programming model that takes into account safety inventory levels, orders in transit, demand forecasts, and various cost factors.

[0140] First, construct a state vector space. This space includes the safety inventory levels and the quantities of orders in transit of each category. For example, if there are two categories, the safety inventory of one category is 10 units and the order in transit is 5 units, and the safety inventory of the other category is 20 units and the order in transit is 10 units, then the current state vector can be represented as (10, 5, 20, 10).

[0141] Next, define the inventory cost function. This function includes holding costs, shortage costs, and order costs. For example, the holding cost per unit is 1 yuan, the shortage cost per unit is 10 yuan, and the fixed cost per order is 50 yuan. Then, add the inventory cost function to the expected value of the future value function adjusted by a discount factor to construct a dynamic programming value function. The discount factor is used to measure the impact of future costs on current decisions. Assume the discount factor is 0.9, indicating that the cost of the next period is equivalent to 0.9 times that of the current period.

[0142] Then, construct the state transition equation. This equation describes how the inventory state changes over time. It takes into account the current inventory level, orders in transit, actual demand, and replenishment quantity, and updates the status of orders in transit according to the replenishment lead time for each category. For example, if the current inventory of a certain category is 10 units, the orders in transit are 5 units, the actual demand is 8 units, the replenishment quantity is 15 units, and the replenishment lead time is 2 periods, then the inventory state in the next period will be 10 + 5 - 8 + 15 = 22 units, and the status of orders in transit will be updated according to the lead time.

[0143] Substitute the state transition equation into the dynamic programming value function and solve the dynamic programming value function recursively to obtain the optimal value function. The recursive process starts from the last period and gradually calculates backward until the first period. The output value of the optimal value function corresponds to the reorder point for each category. The reorder point is a threshold at which replenishment is required when the inventory level drops below this threshold.

[0144] Compare the sum of the safety inventory level and the quantity of orders in transit for each category with the corresponding reorder point. When the sum of the safety inventory level and the quantity of orders in transit is lower than the reorder point, determine the replenishment time point.

[0145] Obtain the resource occupancy of each category and the total available resources, and construct the resource constraint conditions. For example, assume that each category occupies one unit of resources and the total available resources are 100 units. At the same time, obtain the unit procurement cost and the available budget for each category, and construct the budget constraint conditions. For example, assume that the unit procurement cost for each category is 2 yuan and the available budget is 200 yuan.

[0146] Calculate the economic order quantity based on the fixed ordering cost, average demand rate, and unit holding cost for each category. The economic order quantity refers to the optimal quantity for each order to minimize the total inventory cost. For example, assume that the fixed ordering cost for a certain category is 50 yuan, the average demand rate is 10 units / period, and the unit holding cost is 1 yuan / period, then the economic order quantity can be calculated.

[0147] On the premise of meeting the resource constraint conditions and budget constraint conditions, use the economic order quantity as the replenishment quantity.

[0148] For example, assume that the economic order quantity calculated according to the above steps is 20 units and it meets the resource constraint and budget constraint, then the replenishment quantity for this category is 20 units.

[0149] The beneficial effects of this method can be summarized in the following three aspects:

[0150] Reduce inventory costs: By optimizing replenishment decisions, inventory holding costs, stock - out costs, and ordering costs can be effectively reduced, thus lowering the total inventory cost.

[0151] Improve inventory management efficiency: This method provides a systematic and scientific approach to inventory management, which can help enterprises better manage inventory and improve inventory management efficiency.

[0152] Enhance customer satisfaction: By ensuring sufficient inventory, customer needs can be met in a timely manner, avoiding stock - out situations, and thus enhancing customer satisfaction.

[0153] In an alternative implementation, substituting the state - transition equation into the dynamic - programming value function and solving the dynamic - programming value function recursively, the optimal value function obtained includes:

[0154] ;

[0155] where, represents the optimal value function at time t in state S t , A t represents the action at time t, represents the inventory - cost function at time t, represents the discount factor, which is used to balance current costs and future costs, represents the expected value of the value function at time t + 1;

[0156] The inventory - cost function is shown by the following formula:

[0157] ;

[0158] where, represents the warehousing - cost function, represents the stock - out - cost function, represents the backlog - cost function, , , respectively represent the final weight coefficients corresponding to the warehousing - cost function, stock - out - cost function, and backlog - cost function, represents the inventory level at time t;

[0159] The state - transition equation is shown by the following formula:

[0160] ;

[0161] ;

[0162] ;

[0163] Among them, represents the demand quantity of category i at time t, represents the replenishment lead time of category i at time t + 1, represents the replenishment lead time of category N at time t + 1, represents the quantity of in - transit orders of category i at time t, represents the order quantity of category i at time t, represents the transpose;

[0164] The reorder point calculation formula is as follows:

[0165] ;

[0166] Among them, represents the optimal reorder point of category i, represents the average demand rate of category i, represents the standard normal distribution quantile corresponding to the service level, represents the standard deviation of the demand of category i.

[0167] A multi - category inventory control method based on dynamic programming and reorder point can effectively reduce inventory costs and improve service levels. The core of this method is to determine the optimal inventory level through dynamic programming and make replenishment decisions in combination with the reorder point strategy.

[0168] First, it is necessary to determine the parameters of the model, including the warehousing cost, stock - out cost, backlog cost, average value and standard deviation of demand, replenishment lead time, and service level target for each category. For example, assume we have two categories. The warehousing cost of category 1 is 1 yuan per unit per day, the stock - out cost is 10 yuan per unit, the backlog cost is 0.5 yuan per unit per day, the average value of demand is 10 units per day, the standard deviation is 2 units, the replenishment lead time is 3 days, and the service level target is 95%. The warehousing cost of category 2 is 2 yuan per unit per day, the stock - out cost is 20 yuan per unit, the backlog cost is 1 yuan per unit per day, the average value of demand is 5 units per day, the standard deviation is 1 unit, the replenishment lead time is 2 days, and the service level target is 90%.

[0169] Next, according to the set service level target, look up the standard normal distribution table to obtain the corresponding quantile. For example, the quantile corresponding to a 95% service level is approximately 1.65, and the quantile corresponding to a 90% service level is approximately 1.28.

[0170] Then, calculate the reorder point for each category. The calculation method of the reorder point is: the average demand rate multiplied by the replenishment lead time, plus the safety stock. The calculation method of the safety stock is: the quantile multiplied by the standard deviation of demand and then multiplied by the square root of the replenishment lead time. For example, the reorder point for Category 1 is 10 * 3 + 1.65 * 2 * sqrt(3) ≈ 35.8, and the reorder point for Category 2 is 5 * 2 + 1.28 * 1 * sqrt(2) ≈ 11.8.

[0171] Next, construct a dynamic programming model. The dynamic programming value function represents the minimum expected value of the total cost from the current to a future period of time under a specific time and inventory status. This function includes the current inventory cost and the expected value of future costs, and the future costs are adjusted by a discount factor. Inventory costs include warehousing costs, stockout costs, and overstock costs, which are multiplied by their respective weight coefficients.

[0172] Then, describe the change of inventory level over time through the state transition equation. The state transition equation takes into account the current inventory level, demand quantity, on-order quantity, and order quantity.

[0173] Finally, solve the dynamic programming value function recursively. Starting from the last time period, gradually backward to the first time period, calculate the optimal value function for each time period and inventory status. At each time period, according to the current inventory level and the reorder point, decide whether to place an order. If the inventory level is lower than the reorder point, an order needs to be placed, and the order quantity is equal to the reorder point minus the current inventory level plus the expected demand for the next period.

[0174] Through the above steps, the optimal ordering strategy for each time period and inventory status can be obtained, thereby minimizing the total inventory cost.

[0175] The beneficial effects of this method are reflected in three aspects:

[0176] Reduce inventory costs: By optimizing the inventory level through dynamic programming, warehousing costs, stockout costs, and overstock costs can be effectively reduced, thereby minimizing the total inventory cost.

[0177] Improve service level: By setting service level targets and calculating safety stock, the order fulfillment rate can be effectively improved, thereby enhancing customer satisfaction.

[0178] Simplify the decision-making process: This method provides a clear ordering strategy, which can help enterprises simplify the inventory management process and improve operational efficiency.

[0179] Figure 2 For the structural schematic diagram of the inventory optimization system based on the deep learning model in the embodiment of the present invention, as Figure 2As shown, the system includes:

[0180] A first unit for obtaining multi-dimensional data corresponding to a target commodity. The multi-dimensional data includes historical sales data, commodity attribute data, and consumer behavior data. The historical sales data includes commodity code, sales time, sales quantity, sales price, and commodity inventory. The commodity attribute data includes commodity material, commodity style, commodity seasonal attribute, and commodity size information. The consumer behavior data includes page views, favorites, add-to-cart quantities, and repurchase rates. Perform time series alignment processing on the multi-dimensional data to construct a unified time series feature matrix. Input the time series feature matrix into a time series convolutional neural network model to extract short-term time series features.

[0181] A second unit for constructing a parameterized dynamic convolution kernel through a graph neural network based on the multi-dimensional data, and constructing a soft clustering assignment matrix based on a multi-layer perceptron, and obtaining long-term dependence features that integrate the spatial dimension and the time dimension through residual connection. Perform feature fusion on the short-term time series features and the long-term dependence features to obtain a mixed feature vector. Input the mixed feature vector into a bidirectional long short-term memory network to generate a sales prediction sequence. Calculate a prediction confidence interval based on the sales prediction sequence to form a final sales prediction result.

[0182] A third unit for calculating a safety inventory level based on the sales prediction result, where the safety inventory level is determined by the upper and lower bounds of the prediction confidence interval. Construct an inventory cost function, which includes a weighted combination of three dimensions: warehousing cost, out-of-stock cost, and overstock cost. Input the safety inventory level and the inventory cost function into a dynamic programming model to generate an inventory replenishment decision, where the inventory replenishment decision includes a replenishment time point and a replenishment quantity. Send the inventory replenishment decision to a supply chain execution system, where the dynamic programming model is constructed based on a deep reinforcement learning model.

[0183] The present invention can be a method, device, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions for performing various aspects of the present invention loaded thereon.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. Inventory optimization method based on deep learning model, characterized in that: include: Acquire multi-dimensional data corresponding to the target product, the multi-dimensional data including historical sales data, product attribute data, and consumer behavior data, wherein the historical sales data includes product code, sales time, sales quantity, sales price, and product inventory; the product attribute data includes product material, product style, product seasonal attribute, and product size information; the consumer behavior data includes page views, collections, purchases, and repurchase rates; perform time series alignment on the multi-dimensional data to construct a unified time series feature matrix; input the time series feature matrix into a time series convolutional neural network model to extract short-term time series features; Initial node features are obtained through graph convolution operation according to the multi-dimensional data, dynamic time series features are obtained by adaptive convolution operation on the initial node features based on the dynamic convolution kernel generated by the multi-layer perceptron network, and cluster center features are obtained by weighted aggregation of the dynamic time series features in combination with the soft clustering allocation matrix; gated output features are obtained through a gating network according to the dynamic time series features and the cluster center features, and multi-head attention features are obtained through a multi-head attention mechanism based on the gated output features, and the multi-head attention features are nonlinearly transformed to obtain transformation features, and the transformation features and the multi-head attention features are residually connected to obtain long-term dependency features; the short-term time series features are feature fused with the long-term dependency features to obtain a mixed feature vector; the mixed feature vector is input into a bidirectional long short-term memory network to generate a sales forecast sequence; the forecast confidence interval is calculated based on the sales forecast sequence to form a final sales forecast result; Based on the sales forecast result, the safety stock level is calculated, wherein the safety stock level is determined by the upper and lower bounds of the forecast confidence interval; an inventory cost function is constructed, wherein the inventory cost function includes a weighted combination of three dimensions: warehousing cost, out-of-stock cost, and backlog cost; the safety stock level and the inventory cost function are input into a dynamic programming model to generate an inventory replenishment decision, wherein the inventory replenishment decision includes a replenishment time point and a replenishment quantity; the inventory replenishment decision is sent to a supply chain execution system, wherein the dynamic programming model is constructed based on a deep reinforcement learning model.

2. The method according to claim 1, characterized in that Methods for generating long-term dependency features include: Normalizing the input multi-dimensional data to obtain a normalized feature matrix, taking each data in the multi-dimensional data as a graph node, calculating the similarity between graph nodes based on the normalized feature matrix to construct an adaptive adjacency matrix, and performing a graph convolution operation on the adaptive adjacency matrix and the normalized feature matrix to obtain initial node features; Extracting time coding information from the initial node features, performing a splicing operation on the time coding information and the initial node features and inputting the splicing operation into a multi-layer perceptron network to generate dynamic convolution kernel parameters, and performing an adaptive convolution operation on the initial node features based on the dynamic convolution kernel parameters to obtain dynamic time series features; Based on the dynamic time series features, a feature difference vector between node pairs is constructed, the feature difference vector is input into a multilayer perceptron to calculate a node similarity matrix, the node similarity matrix is ​​softmax normalized to obtain a soft clustering allocation matrix, and the dynamic time series features are weighted aggregated using the soft clustering allocation matrix to obtain cluster center features; Performing feature fusion on the cluster center feature and the dynamic time series feature to obtain a fused feature, inputting the fused feature into a gating network to generate an adaptive gating weight, and performing selective information transfer on the fused feature based on the adaptive gating weight to obtain a gated output feature; Input the gated output features into three independent linear mapping layers to obtain a query matrix, a key matrix and a value matrix, calculate an attention score based on the query matrix and the key matrix, and multiply the attention score by a time decay factor to obtain a temporal attention weight; Based on the temporal attention weights, a value matrix is ​​weightedly aggregated to obtain a multi-head attention feature, the multi-head attention feature is input into a feedforward neural network for nonlinear transformation to obtain a transformation feature, the transformation feature is residually connected with the multi-head attention feature and subjected to layer normalization processing to obtain a long-term dependency feature with a long-term dependency relationship.

3. The method according to claim 1, characterized in that Inputting the mixed feature vector into a bidirectional long short-term memory network to generate a sales forecast sequence; calculating a forecast confidence interval based on the sales forecast sequence to form a final sales forecast result includes: Input the mixed feature vector into the forward network of the bidirectional long short-term memory network to obtain a forward hidden state feature, input the mixed feature vector into the backward network of the bidirectional long short-term memory network to obtain a backward hidden state feature, and concatenate the forward hidden state feature with the backward hidden state feature to obtain a bidirectional hidden state feature; An attention weight calculation unit is constructed based on the bidirectional hidden state feature, and the attention weight calculation unit includes a historical state mapping layer and a current state mapping layer. The historical bidirectional hidden state feature and the bidirectional hidden state feature at the current moment are respectively input into the historical state mapping layer and the current state mapping layer, and a temporal attention weight is obtained through attention scoring; the context vector is obtained by weighted summing the bidirectional hidden state feature at the current moment using the temporal attention weight, and the context vector is feature-fused with the bidirectional hidden state feature at the current moment to obtain an enhanced feature representation; Input the enhanced feature representation into a prediction distribution generation network, the prediction distribution generation network includes a mean prediction layer and a variance prediction layer, the mean prediction layer outputs a prediction mean, and the variance prediction layer outputs a prediction variance; Monte Carlo sampling is performed based on the prediction mean and the prediction variance to obtain multiple groups of prediction samples, and statistical distribution characteristics of the multiple groups of prediction samples are calculated to obtain a sales forecast sequence; The prediction interval is adaptively adjusted, and the prediction mean and the prediction variance are dynamically updated based on the degree of deviation between the actual sales volume and the predicted mean, so as to obtain the calibrated upper and lower bounds of the prediction confidence interval, and the sales forecast sequence and the upper and lower bounds of the prediction confidence interval are combined to form the final sales forecast result.

4. The method according to claim 3, characterized in that Based on the sales forecast result, the safety stock level is calculated, wherein the safety stock level is determined by the upper and lower bounds of the forecast confidence interval, including: Calculate the forecast interval standard deviation based on the forecast confidence interval upper bound and the forecast confidence interval lower bound of the sales forecast sequence, and combine the forecast interval standard deviation with the service level coefficient, replenishment lead time, and review cycle to obtain the initial safety stock level; Calculate a first absolute value of the difference between the upper bound of the forecast confidence interval and the maximum historical forecast value, and a second absolute value of the difference between the minimum historical forecast value and the lower bound of the forecast confidence interval, and select the ratio of the larger value of the first absolute value and the second absolute value to the standard deviation of the forecast interval as the demand fluctuation adjustment factor; The product of the demand fluctuation adjustment factor and the pre-acquired sensitivity parameter is used as a dynamic adjustment amount, and the initial safety stock level and the dynamic adjustment amount are combined to obtain a final safety stock level.

5. The method according to claim 4, characterized in that Construct an inventory cost function, which includes a weighted combination of three dimensions: storage cost, out-of-stock cost, and backlog cost: Constructing a storage cost function, which is the product of the unit storage cost and the current inventory level; constructing a stockout cost function, which is the product of the unit stockout cost and the inventory gap; constructing a backlog cost function, which is the product of the unit backlog cost and the inventory amount exceeding the upper limit of the prediction confidence interval; Initial weight coefficients are respectively assigned to the warehousing cost function, the out-of-stock cost function, and the backlog cost function, and weight gradients are calculated based on the contribution of each cost function to the total cost. The weight coefficients are updated according to the weight gradients and the learning rate. The updated weight coefficients are normalized after exponential transformation to obtain the final weight coefficients, and the inventory cost function is constructed by weighted combination of the final weight coefficients and the corresponding cost functions.

6. The method according to claim 5, characterized in that The safety stock level and the inventory cost function are input into a dynamic programming model to generate an inventory replenishment decision, wherein the inventory replenishment decision includes a replenishment time point and a replenishment quantity. Constructing a state vector space, wherein the state vector space includes the safety inventory level and the in-transit order quantity of each category; constructing a dynamic programming value function by adding the inventory cost function to the expected value of the future value function adjusted by the discount factor; constructing a state transfer equation based on the state vector space, wherein the state transfer equation calculates the inventory status of the next period according to the current inventory level, in-transit order quantity, actual demand quantity and replenishment quantity of each category, and updates the in-transit order status according to the replenishment lead time of each category; Substituting the state transfer equation into the dynamic programming value function, solving the dynamic programming value function in a recursive manner to obtain an optimal value function; calculating the reorder point of each category based on the optimal value function, comparing the sum of the safety stock level and the in-transit order quantity of each category with the corresponding reorder point, and determining the replenishment time point when the sum of the safety stock level and the in-transit order quantity is lower than the reorder point; The resource occupancy and total available resources of each category are obtained to construct resource constraints, and the unit procurement cost and available budget of each category are obtained to construct budget constraints. The economic order quantity is calculated based on the fixed ordering cost, average demand rate and unit holding cost of each category, and the economic order quantity is used as the replenishment quantity on the premise that the resource constraints and the budget constraints are met.

7. The method according to claim 6, characterized in that Substituting the state transfer equation into the dynamic programming value function, solving the dynamic programming value function in a recursive manner, and obtaining the optimal value function includes: Among them, V t (S t ) represents the state S at time t t The optimal value function under t represents the action at time t, TC(I t ) represents the inventory cost function at time t, γ represents the discount factor, which is used to balance current cost and future cost, represents the expected value of the value function at time t+1; The inventory cost function is shown in the following formula: Among them, C s (I t ) represents the storage cost function, C p (I t ) represents the out-of-stock cost function, C o (I t ) represents the backlog cost function, They represent the final weight coefficients corresponding to the warehousing cost function, out-of-stock cost function, and backlog cost function, respectively. t represents the inventory level at time t; The state transition equation is shown in the following formula: in, represents the demand for category i at time t, represents the replenishment lead time of category i at time t+1, represents the replenishment lead time of category N at time t+1, represents the number of in-transit orders of category i at time t, represents the order quantity of category i at time t, (.) T represents transpose; The reorder point calculation formula is as follows: in, represents the optimal reorder point for category i, μ i represents the average demand rate of category i, z α represents the standard normal distribution quantile corresponding to the service level, σ i represents the standard deviation of demand for category i.

8. An inventory optimization system based on a deep learning model, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain multi-dimensional data corresponding to the target product, wherein the multi-dimensional data includes historical sales data, product attribute data, and consumer behavior data, wherein the historical sales data includes product code, sales time, sales quantity, sales price, and product inventory; the product attribute data includes product material, product style, product seasonal attribute, and product size information; the consumer behavior data includes page views, collections, purchases, and repurchase rates; perform time series alignment processing on the multi-dimensional data to construct a unified time series feature matrix; input the time series feature matrix into a time series convolutional neural network model to extract short-term time series features; The second unit is used to obtain initial node features through graph convolution operations according to the multi-dimensional data, perform adaptive convolution operations on the initial node features based on the dynamic convolution kernel generated by the multi-layer perceptron network to obtain dynamic time series features, and weightedly aggregate the dynamic time series features in combination with the soft clustering allocation matrix to obtain cluster center features; obtain gated output features through a gating network according to the dynamic time series features and the cluster center features, obtain multi-head attention features through a multi-head attention mechanism based on the gated output features, perform nonlinear transformation on the multi-head attention features to obtain transformation features, perform residual connection on the transformation features and the multi-head attention features to obtain long-term dependency features; perform feature fusion on the short-term time series features and the long-term dependency features to obtain a mixed feature vector; input the mixed feature vector into a bidirectional long short-term memory network to generate a sales forecast sequence; calculate the forecast confidence interval based on the sales forecast sequence to form a final sales forecast result; The third unit is used to calculate the safety stock level based on the sales forecast result, wherein the safety stock level is determined by the upper and lower bounds of the forecast confidence interval; construct an inventory cost function, wherein the inventory cost function includes a weighted combination of three dimensions: warehousing cost, out-of-stock cost, and backlog cost; input the safety stock level and the inventory cost function into a dynamic programming model to generate an inventory replenishment decision, wherein the inventory replenishment decision includes a replenishment time point and a replenishment quantity; and send the inventory replenishment decision to a supply chain execution system, wherein the dynamic programming model is constructed based on a deep reinforcement learning model.

Citation Information

Patent Citations

  • Replenishment decision model training and replenishment decision method, system, device and medium

    CN113095745A

  • Traffic flow prediction system based on deep learning and dynamic network analysis and application method thereof

    CN118675324A