Logistics distribution demand prediction analysis method and system based on big data
By dividing the logistics distribution area into hexagonal grid units, constructing the order density and flow direction association matrix, and combining wavelet decomposition and graph attention network, the optimal distribution path is generated, which solves the accuracy and efficiency problems of logistics distribution prediction, realizes intelligent allocation and cross-grid scheduling, and improves distribution efficiency and user experience.
Patent Information
- Application Number
- CN202510936052.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-12
AI Technical Summary
Existing logistics and delivery prediction methods are unable to fully exploit the complex spatiotemporal characteristics and dynamic evolution laws in order data, resulting in low accuracy and timeliness of prediction results. It is difficult to effectively integrate multidimensional factors such as order flow, rider movement, and road network status, affecting delivery efficiency and user experience.
The logistics distribution area is divided into multiple regular hexagonal grid units, and the order density feature matrix and order flow correlation intensity matrix are constructed. Through wavelet decomposition and soft threshold denoising, combined with graph attention network and kernel density estimation method, the optimal distribution path is generated and cross-grid scheduling and intelligent allocation are performed.
It improves the accuracy and reliability of logistics distribution demand forecasts, optimizes resource allocation and route planning, improves distribution efficiency and reduces costs.
Smart Images

Figure CN120634173A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to logistics distribution technology, and in particular to a logistics distribution demand forecasting and analysis method and system based on big data. Background Art
[0002] The modern logistics and distribution industry is rapidly developing. With the advancement of e-commerce and urbanization, order volumes are growing exponentially, and order demand is highly uneven across time and space. This poses a significant challenge to accurate demand forecasting for logistics and distribution. Traditional forecasting methods, often based on simple statistical models or rule-based algorithms, fail to fully exploit the complex spatiotemporal characteristics and dynamic evolution patterns hidden in order data. This results in low accuracy and timeliness of forecasts, which in turn impacts delivery efficiency and user experience.
[0003] Furthermore, multidimensional factors such as order flow, rider movement, and road network conditions intertwine within logistics distribution networks, forming a complex dynamic system. Existing research often struggles to effectively integrate order timing characteristics, flow correlations, and the impact of external factors (such as weather, traffic, and emergencies) on order demand. This lacks scientific methods for dynamically analyzing order flows and resource allocation across grid cells from a global perspective, hindering the rational scheduling of distribution resources and real-time planning of optimized routes.
[0004] Therefore, there is an urgent need for a logistics distribution demand forecasting and analysis method and system based on big data to accurately predict order demand and optimize resource allocation, thereby improving the overall efficiency and service quality of logistics distribution. Summary of the Invention
[0005] The embodiments of the present invention provide a logistics distribution demand forecasting and analysis method and system based on big data, which can solve the problems in the prior art.
[0006] According to a first aspect of the embodiments of the present invention, A logistics distribution demand forecasting and analysis method based on big data is provided, including: The logistics distribution area is divided into multiple regular hexagonal grid cells, the order density feature matrix of the grid cells is calculated, the grid cells are dynamically reconstructed and feature extraction is performed to construct a multi-dimensional feature vector of the grid cells, and the order flow correlation strength matrix between the grid cells is calculated based on the multi-dimensional feature vector. The order flow correlation strength matrix is subjected to wavelet decomposition and soft threshold denoising to obtain a denoised correlation strength matrix. The entropy of the order flow within the reconstructed grid cells is calculated based on the denoised correlation strength matrix, and an information flow uncertainty index is constructed. Extract time-series order data to construct the order evolution pattern feature vector. Concatenate the order evolution pattern feature vector with the multidimensional feature vector to obtain an enhanced feature vector. Input the denoised association strength matrix, information flow uncertainty index, and enhanced feature vector into the graph attention network to obtain order propagation characteristics. External information is obtained to calculate the influencing factor vector, and a nonlinear combination is performed with the order propagation characteristics to obtain the order demand forecast value. The kernel density estimation method is used for correction to obtain the demand forecast distribution. The demand forecast distribution gradient between grid units is calculated and converted into a motion guidance vector. The optimal delivery path is generated in combination with the real-time traffic status. When the backlog of orders exceeds the threshold, cross-grid scheduling and intelligent allocation are performed.
[0007] In an optional embodiment, The logistics distribution area is divided into multiple regular hexagonal grid cells. The order density feature matrix of the grid cells is calculated as follows: Establish a coordinate system based on the geometric center of the logistics distribution area, calculate the length of the regular hexagon based on the area and the number of target grids, generate a grid center coordinate mapping table, use the ray method to determine the boundary point set, identify the adjacent grid identifier set, extract the path to construct a road network connection diagram, establish the mapping relationship between the grid and the road network, and generate a grid structure containing the grid center coordinates, boundary point set, adjacent grid identifiers and the road network connection diagram; Set a time window to collect the starting point coordinate sequence and timestamp sequence of orders within the grid unit. Calculate the order density per unit area based on the ratio of the number of orders to the grid coverage area. Obtain the GPS positioning trajectory of the rider and the delivery range polygon to calculate the rider density per unit area. Calculate the delivery capacity coefficient based on the load limit and historical order completion rate. Collect the real-time traffic volume and average speed of the road section to calculate the road congestion index. The order density per unit area, rider density per unit area, delivery capacity coefficient and road congestion index collected from each grid unit are normalized to their maximum and minimum values to obtain the corresponding standardized eigenvalue sequence. The information entropy is calculated based on each standardized eigenvalue sequence to determine the weight coefficient of each feature. The order density eigenvalue is obtained by weighted summation of the standardized eigenvalue and the corresponding weight coefficient, and organized into a grid unit order density feature matrix according to the grid unit identifier.
[0008] In an optional embodiment, After dynamic reconstruction of the grid cells, feature extraction is performed to construct a multi-dimensional feature vector of the grid cells. Based on the multi-dimensional feature vector, the order flow correlation strength matrix between the grid cells is calculated, including: Obtain historical order data and the initial grid unit order density feature matrix, construct a grid unit adjacency relationship graph, use a graph neural network to extract structural features, calculate the dynamic weight coefficients of adjacent grid units, and weightedly fuse them with the order density feature value to obtain grid similarity; Calculating a grid reconstruction effect evaluation value based on historical order distribution, inputting the grid similarity and reconstruction effect evaluation value into a deep reinforcement learning model to obtain a grid reconstruction strategy; performing a grid reconstruction operation according to the grid reconstruction strategy, and constructing the order density distribution of the reconstructed grid unit into a spatiotemporal feature sequence; Based on the spatiotemporal feature sequence, the order change trend characteristics, grid spatial dependency characteristics, and distribution resource characteristics of each grid unit are extracted respectively, and the deep feature vector of the grid unit is obtained by adaptive weight fusion. Based on the distribution patterns of order starting points and destinations in historical order flow data, a knowledge base containing the spatiotemporal patterns of order flows is constructed; the comprehensive correlation weights of the time dimension, space dimension, and resource dimension are calculated based on the deep feature vectors of the grid units; based on the order flow patterns in the knowledge base and the comprehensive correlation weights between the grid units, the order flow correlation strength between the grid units is calculated, and an order flow correlation strength matrix is constructed.
[0009] In an optional embodiment, The order flow correlation strength matrix is subjected to wavelet decomposition and soft threshold denoising to obtain the denoised correlation strength matrix. The entropy of the order flow in the reconstructed grid unit is calculated based on the denoised correlation strength matrix, and the information flow uncertainty index is constructed, including: The order flow correlation intensity matrix is segmented into multiple sliding windows. The number of wavelet decomposition layers is calculated based on the order peak distribution characteristics within each window after segmentation. A two-dimensional discrete wavelet transform is performed to obtain a multi-scale frequency coefficient matrix, and abnormal fluctuation components are identified through energy cluster analysis. Based on the abnormal fluctuation component, a bidirectional shrinkage function having a positive fluctuation shrinkage parameter and a negative fluctuation shrinkage parameter is constructed, a dynamic noise estimate is generated by combining the historical noise variance and the current observation noise variance, and the dynamic noise estimate is input into an adaptive threshold function to generate a threshold parameter; The bidirectional shrinkage function and the threshold parameter are combined to construct a hybrid threshold processing operator, and soft threshold denoising is performed on the multi-scale frequency coefficient matrix. The multi-scale order flow features are extracted from the multi-scale frequency coefficient matrix, and the features are cascaded with the denoised frequency coefficient matrix. The denoised association strength matrix is reconstructed using the attention weight mechanism; A multi-layer probability transfer network is constructed based on the denoised correlation intensity matrix. The order flow entropy of each grid unit is calculated separately. A hierarchical entropy index is formed through inter-layer information transmission and integration. A recursive neural network is used to extract time scale features and perform nonlinear combination to construct an information flow uncertainty index.
[0010] In an optional embodiment, Extract time-series order data to construct the order evolution pattern feature vector. Concatenate the order evolution pattern feature vector with the multi-dimensional feature vector to obtain an enhanced feature vector. Input the denoised association strength matrix, information flow uncertainty index, and enhanced feature vector into the graph attention network to obtain the order propagation features including: Extracting the time series order data of the reconstructed grid cells, calculating an adaptive penalty factor based on the local fluctuation characteristics of the time series order data, constructing an optimization objective function with a penalty term using the adaptive penalty factor, and iteratively solving the optimization objective function to obtain a trend term, a period term, and a random term; The slope feature and acceleration feature of the trend item are calculated and weightedly combined to obtain the trend fluctuation intensity feature; the period item is subjected to a multi-scale Fourier transform to obtain a frequency domain feature sequence, the correlation coefficients between different frequency components are calculated to obtain a frequency correlation matrix, and the main period feature, periodic intensity feature, and periodic stability feature are extracted; and the feature combination forms the order evolution pattern feature vector; Calculating the importance score of each feature component in the order evolution pattern feature vector and the multidimensional feature vector, and selecting feature components that are higher than a preset feature selection threshold for splicing to obtain an enhanced feature vector; The grid units are used as graph network nodes. The spatial adjacency relationship and semantic association relationship are constructed based on the denoised association strength matrix. The spatial distance attention coefficient and feature similarity attention coefficient are calculated. The enhanced feature vector is combined with the information flow uncertainty index to form the initial feature of the node. The order propagation feature is obtained through multi-level information transmission and residual connection fusion.
[0011] In an optional embodiment, Obtain external information to calculate the impact factor vector, and perform nonlinear combination with the order propagation characteristics to obtain the order demand forecast value. Then use the kernel density estimation method to perform correction, and obtain the demand forecast distribution including: The influence factor vector is concatenated with the order propagation feature to obtain an initial feature vector, the neighbor clustering of the local features in the initial feature vector is calculated to obtain the self-attention weight, the feature importance is calculated to obtain the cross-attention weight, and the initial feature vector is weightedly combined according to the self-attention weight and the cross-attention weight to obtain the weighted feature; Calculating the mutual information and conditional entropy of the weighted features, determining the branch merging weights and performing nonlinear combination to obtain the initial order demand forecast value of the grid unit; Performing time-frequency decomposition on the historical forecast deviation sequence to obtain multiple frequency components, calculating the fluctuation patterns of the multiple frequency components, calculating information entropy based on the fluctuation patterns, determining the optimal time window size, and performing stratified sampling on the historical forecast deviation sequence to obtain a forecast deviation sample set; Perform kernel density estimation based on the temporal correlation constraint and information divergence of the prediction deviation sample set to obtain an error probability density function; Calculate the statistical characteristics of the error probability density function, extract the local fluctuation characteristics of the initial order demand forecast value, calculate the corrected forecast value based on the statistical characteristics and the local fluctuation characteristics, repeatedly sample the error probability density function at the same time to obtain a multi-component site, evaluate the stability of the multi-component site to obtain a confidence interval, and combine the corrected forecast value with the confidence interval to obtain an order demand forecast distribution with a confidence interval.
[0012] In an optional embodiment, Calculate the demand forecast distribution gradient between grid cells, convert it into a motion guidance vector, and generate the optimal delivery path based on real-time traffic status. When the backlog exceeds the threshold, cross-grid scheduling and intelligent allocation are performed, including: Extract the multi-order statistical moments of the order demand forecast distribution of the grid cells, calculate the information entropy difference and kernel density distance between adjacent grid cells, construct an asymmetric migration probability matrix, and obtain the benchmark flow direction characteristics through tensor decomposition; Obtaining the geographic connectivity between grids and historical order flow sequences, extracting periodic flow patterns, and combining them with spatial autocorrelation coefficients to construct an adaptive weight network. The baseline flow features are input into the adaptive weight network to obtain the rider movement guidance vector. The historical traffic time series is obtained and subjected to wavelet packet decomposition to obtain multiple frequency band signals. The fluctuation characteristics and phase information of the multiple frequency band signals are extracted. A recursive state equation is constructed based on the fluctuation characteristics. The phase information is used to design an adaptive observation matrix. The recursive state equation and adaptive observation matrix are used to predict the real-time traffic status. The time-varying road network cost matrix is constructed in combination with the road network topology characteristics. The rider motion guidance vector is mapped into the driving force field to divide the sub-blocks. The overall optimal delivery path is obtained through collaborative optimization. Based on the long-range correlation and scale-dependence characteristics of the historical data of grid backlog orders, a memory decay function and an adaptive bandwidth kernel function are constructed to generate a dynamic threshold surface. When it is detected that the number of grid backlog orders exceeds the dynamic threshold surface, a cross-grid collaborative objective function is constructed based on the rider's historical performance and current load, and an optimization solution is performed based on the propagation entropy matrix and the grid reachability tensor to obtain a cross-grid scheduling strategy for intelligent order allocation.
[0013] A second aspect of an embodiment of the present invention provides a logistics distribution demand forecasting and analysis system based on big data, comprising: The first unit is used to divide the logistics distribution area into multiple regular hexagonal grid cells, calculate the order density feature matrix of the grid cells, dynamically reconstruct the grid cells, perform feature extraction, construct a multi-dimensional feature vector of the grid cells, and calculate the order flow correlation strength matrix between the grid cells based on the multi-dimensional feature vector; The second unit is used to perform wavelet decomposition and soft threshold denoising on the order flow correlation strength matrix to obtain a denoised correlation strength matrix; based on the denoised correlation strength matrix, the entropy of the order flow within the reconstructed grid cells is calculated and an information flow uncertainty index is constructed; The third unit is used to extract time-series order data to construct the order evolution pattern feature vector, concatenate the order evolution pattern feature vector with the multi-dimensional feature vector to obtain the enhanced feature vector, and input the denoised association strength matrix, information flow uncertainty index, and enhanced feature vector into the graph attention network to obtain the order propagation characteristics; The fourth unit is used to obtain external information to calculate the influencing factor vector, and perform nonlinear combination with the order propagation characteristics to obtain the order demand forecast value. The kernel density estimation method is used for correction to obtain the demand forecast distribution, and the demand forecast distribution gradient between grid units is calculated and converted into a motion guidance vector. The optimal delivery path is generated in combination with the real-time traffic status. When the backlog of orders exceeds the threshold, cross-grid scheduling and intelligent allocation are performed.
[0014] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0015] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0016] In this embodiment, the logistics distribution area is divided into regular hexagonal grid cells and a regional order density assessment model is established based on historical order data, achieving dynamic reconstruction of the grid cells. By constructing a multidimensional feature vector and an order flow correlation strength matrix, the order distribution characteristics and delivery demand patterns within the region can be more accurately captured, improving the accuracy of logistics distribution demand forecasts. Wavelet decomposition and adaptive threshold denoising are used to process the order flow correlation strength matrix, and the time series order data is decomposed using the variational mode decomposition method. The order propagation characteristics are extracted through a graph attention network. The influence of external factors such as weather, traffic, and events is also considered, and probability correction is performed using the kernel density estimation method, significantly improving the accuracy and reliability of order demand forecasts. By calculating the gradient of the order demand forecast distribution between grid cells and converting it into a rider movement guidance vector, the optimal delivery path is generated in combination with a time-varying road condition prediction model. In the event of an order backlog, cross-grid scheduling can be performed based on the order propagation characteristics, and intelligent order allocation can be performed based on the rider's historical performance and current load, effectively improving logistics distribution efficiency and reducing distribution costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flow chart of a method for predicting and analyzing logistics distribution demand based on big data according to an embodiment of the present invention; Figure 2 This is a simulation effect diagram of the order flow correlation strength matrix analysis according to an embodiment of the present invention; Figure 3 A histogram showing the collaborative optimization of delivery routes and driving force field partitioning effects according to an embodiment of the present invention; Figure 4 This is a structural diagram of a logistics distribution demand forecasting and analysis system based on big data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0019] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0020] Figure 1 FIG. 1 is a flow chart of a method for predicting and analyzing logistics distribution demand based on big data according to an embodiment of the present invention. Figure 1 As shown, the method includes: The logistics distribution area is divided into multiple regular hexagonal grid cells, the order density feature matrix of the grid cells is calculated, the grid cells are dynamically reconstructed and feature extraction is performed to construct a multi-dimensional feature vector of the grid cells, and the order flow correlation strength matrix between the grid cells is calculated based on the multi-dimensional feature vector. The order flow correlation strength matrix is subjected to wavelet decomposition and soft threshold denoising to obtain a denoised correlation strength matrix. The entropy of the order flow within the reconstructed grid cells is calculated based on the denoised correlation strength matrix, and an information flow uncertainty index is constructed. Extract time-series order data to construct the order evolution pattern feature vector. Concatenate the order evolution pattern feature vector with the multidimensional feature vector to obtain an enhanced feature vector. Input the denoised association strength matrix, information flow uncertainty index, and enhanced feature vector into the graph attention network to obtain order propagation characteristics. External information is obtained to calculate the influencing factor vector, and a nonlinear combination is performed with the order propagation characteristics to obtain the order demand forecast value. The kernel density estimation method is used for correction to obtain the demand forecast distribution. The demand forecast distribution gradient between grid units is calculated and converted into a motion guidance vector. The optimal delivery path is generated in combination with the real-time traffic status. When the backlog of orders exceeds the threshold, cross-grid scheduling and intelligent allocation are performed.
[0021] In an optional embodiment, The logistics distribution area is divided into multiple regular hexagonal grid cells. The order density feature matrix of the grid cells is calculated as follows: Establish a coordinate system based on the geometric center of the logistics distribution area, calculate the length of the regular hexagon based on the area and the number of target grids, generate a grid center coordinate mapping table, use the ray method to determine the boundary point set, identify the adjacent grid identifier set, extract the path to construct a road network connection diagram, establish the mapping relationship between the grid and the road network, and generate a grid structure containing the grid center coordinates, boundary point set, adjacent grid identifiers and the road network connection diagram; Set a time window to collect the starting point coordinate sequence and timestamp sequence of orders within the grid unit. Calculate the order density per unit area based on the ratio of the number of orders to the grid coverage area. Obtain the GPS positioning trajectory of the rider and the delivery range polygon to calculate the rider density per unit area. Calculate the delivery capacity coefficient based on the load limit and historical order completion rate. Collect the real-time traffic volume and average speed of the road section to calculate the road congestion index. The order density per unit area, rider density per unit area, delivery capacity coefficient and road congestion index collected from each grid unit are normalized to their maximum and minimum values to obtain the corresponding standardized eigenvalue sequence. The information entropy is calculated based on each standardized eigenvalue sequence to determine the weight coefficient of each feature. The order density eigenvalue is obtained by weighted summation of the standardized eigenvalue and the corresponding weight coefficient, and organized into a grid unit order density feature matrix according to the grid unit identifier.
[0022] For example, first, in terms of the division of the logistics distribution area, a plane rectangular coordinate system is established with the geometric center of the distribution area as the origin. Assuming that the total area of the distribution area is 100 square kilometers, the target is divided into 400 grid cells. By calculation, it can be obtained that the side length of the regular hexagon is 300 meters. For each grid cell, its center coordinates are determined by row and column indexes. For example, the center coordinates of the grid cell numbered (3, 4) are (900, 1200) meters. The ray method is used to determine whether the grid boundary point is located in the distribution area. Each regular hexagon has 6 vertex coordinates. Adjacent grids are identified by boundary overlap detection. For example, the grid cell numbered (3, 4) is adjacent to grid cells (3, 3), (3, 5), (2, 4), (4, 4), etc.
[0023] To establish road network topology, the shortest paths and travel times between road nodes are extracted. For example, the shortest path from node A to node B is 500 meters long, with a typical travel time of 3 minutes. This information is integrated into a road network connectivity diagram, and a mapping relationship between grid cells and the road network is established. The road node and road section information within each grid cell is recorded for subsequent order density assessment.
[0024] For order data collection, a 30-minute time window is set to record the starting coordinates and time information of orders within the grid cell. For example, if a grid cell has 20 orders within the time window and the grid area is 0.09 square kilometers, the order density per unit area is 222 per square kilometer. At the same time, the GPS tracks of the delivery drivers are obtained, and the polygon of each driver's delivery range is recorded. Assuming there are 5 delivery drivers within a grid cell, the delivery density per unit area is 56 per square kilometer.
[0025] For delivery capacity assessment, the delivery capacity coefficient is calculated based on a maximum rider load of 10 kg and a historical order completion rate of 95%. Real-time traffic flow, such as 300 vehicles per hour and an average speed of 25 kilometers per hour, is collected to calculate the road congestion index.
[0026] During feature standardization, the four eigenvalues of each grid cell are normalized. For example, the normalized order density of a grid cell is 0.85, the rider density is 0.6, the delivery capacity coefficient is 0.75, and the road congestion index is 0.4. The feature weights are determined using the information entropy method, such as a weight of 0.4 for order density, 0.3 for rider density, 0.2 for delivery capacity coefficient, and 0.1 for road congestion index. The final calculated order density eigenvalue for this grid cell is 0.69. The eigenvalues of all grid cells are organized into a matrix for subsequent distribution area optimization analysis.
[0027] In this embodiment, by adopting the regular hexagonal grid division method, uniform coverage of the distribution area is achieved, the spatial discontinuity of the traditional rectangular grid when connected diagonally is avoided, and the accuracy and practicality of the regional division are improved. The comprehensive evaluation model based on multi-dimensional data features fully considers key factors such as order density, rider resources, delivery capacity and traffic conditions, making the regional order density evaluation results more comprehensive and objective, and able to truly reflect the business characteristics of the distribution area. The information entropy weight method is used to determine the importance of features, avoiding the deviation that may be caused by subjective human empowerment. At the same time, the comparability of indicators of different dimensions is guaranteed through feature standardization processing, which improves the reliability and applicability of the evaluation results.
[0028] In an optional embodiment, after dynamically reconstructing the grid cells, feature extraction is performed to construct a multi-dimensional feature vector of the grid cells, and the order flow correlation strength matrix between the grid cells is calculated based on the multi-dimensional feature vector, including: Obtain historical order data and the initial grid unit order density feature matrix, construct a grid unit adjacency relationship graph, use a graph neural network to extract structural features, calculate the dynamic weight coefficients of adjacent grid units, and weightedly fuse them with the order density feature value to obtain grid similarity; Calculating a grid reconstruction effect evaluation value based on historical order distribution, inputting the grid similarity and reconstruction effect evaluation value into a deep reinforcement learning model to obtain a grid reconstruction strategy; performing a grid reconstruction operation according to the grid reconstruction strategy, and constructing the order density distribution of the reconstructed grid unit into a spatiotemporal feature sequence; Based on the spatiotemporal feature sequence, the order change trend characteristics, grid spatial dependency characteristics, and distribution resource characteristics of each grid unit are extracted respectively, and the deep feature vector of the grid unit is obtained by adaptive weight fusion. Based on the distribution patterns of order starting points and destinations in historical order flow data, a knowledge base containing the spatiotemporal patterns of order flows is constructed; the comprehensive correlation weights of the time dimension, space dimension, and resource dimension are calculated based on the deep feature vectors of the grid units; based on the order flow patterns in the knowledge base and the comprehensive correlation weights between the grid units, the order flow correlation strength between the grid units is calculated, and an order flow correlation strength matrix is constructed.
[0029] For example, we first obtain historical order data for the logistics distribution area, including information such as the order's origin and destination, order placement time, and delivery time. We then divide the distribution area into initial grid cells, calculate the order density within each grid cell, and construct an initial order density feature matrix.
[0030] To determine the adjacency of grid cells, an adjacency graph is constructed using the principle of geographic proximity. Edges are established between adjacent grid cells, with the initial edge weight set to 1. A graph neural network is used to extract features from the grid cells. The network consists of multiple graph convolutional layers, each of which aggregates features from adjacent nodes. This feature extraction yields the structural feature vectors of the grid cells.
[0031] When calculating the dynamic weights between adjacent grid cells, we first perform a similarity calculation on the structural feature vectors. We use the cosine similarity method to calculate the degree of similarity between feature vectors, resulting in a similarity value between 0 and 1. We then perform a weighted fusion of this similarity value and the Euclidean distance of the order density feature, using weight coefficients obtained through training with historical data.
[0032] In a deep reinforcement learning model, the state vector contains information such as grid similarity, historical reconstruction records, and coverage constraints. A multi-agent collaborative training approach is employed, with each agent responsible for making grid reconstruction decisions for a subregion. Agents exchange information through a message passing mechanism to jointly optimize the reconstruction strategy. The reward function is designed based on the balance of order distribution and coverage after reconstruction.
[0033] During mesh reconstruction, for mesh merging operations, the order density distribution characteristics of the cells to be merged are extracted, including order quantity, temporal distribution, and spatial distribution. The order distribution characteristics of the merged mesh are calculated using a density-weighted average. For mesh splitting operations, cluster analysis is performed based on the peak values of the order density distribution to determine the optimal splitting solution.
[0034] In the multi-scale feature extraction network, the long short-term memory module uses a bidirectional LSTM structure to extract trend features of order quantity changes over time. The graph convolution module uses a multi-layer graph convolutional network to extract spatial dependencies between grid cells. The recurrent neural network module uses a GRU structure to extract distribution resource utilization features. The three features are adaptively fused using an attention mechanism.
[0035] When building the order flow knowledge base, we analyze the distribution patterns of origins and destinations in historical order data, including periodic patterns in the time dimension and clustering patterns in the spatial dimension. In the spatiotemporal attention network, attention scores for the time, space, and resource dimensions are calculated and weighted together to produce a comprehensive association weight.
[0036] In this embodiment, dynamic grid reconstruction improves the balance of order distribution, reduces the problem of uneven distribution of delivery resources, and improves delivery efficiency. The grid reconstruction strategy can adapt to changes in order distribution and ensure the rationality of grid division. The multi-scale feature extraction network enables in-depth mining of order spatiotemporal characteristics, accurately capturing order change trends and spatial dependencies, and providing reliable feature support for order flow prediction. The order flow correlation strength calculation method based on the knowledge base and spatiotemporal attention mechanism fully considers the influencing factors of the three dimensions of time, space, and resources, improves the accuracy of order flow prediction, and provides an important basis for delivery route planning.
[0037] In an optional embodiment, wavelet decomposition and soft threshold denoising are performed on the order flow correlation strength matrix to obtain a denoised correlation strength matrix; and the entropy of the order flow in the reconstructed grid unit is calculated based on the denoised correlation strength matrix to construct an information flow uncertainty index, including: The order flow correlation intensity matrix is segmented into multiple sliding windows. The number of wavelet decomposition layers is calculated based on the order peak distribution characteristics within each window after segmentation. A two-dimensional discrete wavelet transform is performed to obtain a multi-scale frequency coefficient matrix, and abnormal fluctuation components are identified through energy cluster analysis. Based on the abnormal fluctuation component, a bidirectional shrinkage function having a positive fluctuation shrinkage parameter and a negative fluctuation shrinkage parameter is constructed, a dynamic noise estimate is generated by combining the historical noise variance and the current observation noise variance, and the dynamic noise estimate is input into an adaptive threshold function to generate a threshold parameter; The bidirectional shrinkage function and the threshold parameter are combined to construct a hybrid threshold processing operator, and soft threshold denoising is performed on the multi-scale frequency coefficient matrix. The multi-scale order flow features are extracted from the multi-scale frequency coefficient matrix, and the features are cascaded with the denoised frequency coefficient matrix. The denoised association strength matrix is reconstructed using the attention weight mechanism; A multi-layer probability transfer network is constructed based on the denoised correlation intensity matrix. The order flow entropy of each grid unit is calculated separately. A hierarchical entropy index is formed through inter-layer information transmission and integration. A recursive neural network is used to extract time scale features and perform nonlinear combination to construct an information flow uncertainty index.
[0038] For example, the order flow correlation strength matrix is first segmented using multiple sliding windows in the time and space dimensions. Specifically, time windows and spatial windows of varying sizes can be set. For example, time windows can be set to 1 hour, 3 hours, or 6 hours, and spatial windows can be set to 1 km × 1 km, 2 km × 2 km, or 5 km × 5 km. Using these sliding windows of varying sizes, the order flow correlation strength matrix is segmented into multiple sub-matrices.
[0039] For each segmented submatrix, calculate the order peak distribution characteristics. Statistical indicators such as kurtosis and skewness can be used to describe the order peak distribution. For example, for a 1-hour, 1 km × 1 km window, the calculated kurtosis is 3.5 and the skewness is 0.8, indicating that the order distribution within this window is relatively concentrated and slightly right-skewed.
[0040] The number of wavelet decomposition layers is determined based on the order peak distribution characteristics. A threshold rule can be set. For example, when the kurtosis is greater than 3 and the absolute value of the skewness is greater than 0.5, a 3-layer wavelet decomposition is selected; otherwise, a 2-layer wavelet decomposition is selected. This allows the appropriate number of decomposition layers to be adaptively selected based on the complexity of the order distribution. A two-dimensional discrete wavelet transform is performed on each submatrix to obtain a multi-scale frequency coefficient matrix. Common wavelet basis functions such as Daubechies wavelet or Haar wavelet can be selected here. The frequency coefficient matrix obtained after the transformation contains low-frequency approximation coefficients and high-frequency detail coefficients. The sub-band energy distribution of the multi-scale frequency coefficient matrix is calculated. Specifically, the sum of the squares of each sub-band coefficient can be calculated as the energy value. For example, for the result of a 3-layer wavelet decomposition, the energy values of 10 sub-bands can be obtained.
[0041] Identify abnormal fluctuation components through energy cluster analysis. K-means clustering can be used to cluster subband energy values into "normal" and "abnormal" categories. For example, if a subband's energy value is significantly higher than that of other subbands, it may be identified as an abnormal fluctuation component.
[0042] Based on the identified abnormal fluctuation components, a bidirectional shrinkage function is constructed. This function contains positive and negative shrinkage parameters, which are used to shrink abnormal fluctuations in different directions to varying degrees. For example, the positive shrinkage parameter can be set to 0.8 and the negative shrinkage parameter to 0.6, indicating that the degree of shrinkage for positive abnormal fluctuations is slightly lower than that for negative abnormal fluctuations.
[0043] Calculate the historical noise variance using the order time series data for each grid cell. You can select order data from the past period (e.g., 30 days) and calculate the variance of its fluctuations as the historical noise variance. For example, if the standard deviation of the order volume for a grid cell over the past 30 days is 10, then the historical noise variance is 100.
[0044] The historical noise variance is combined with the current observed noise variance through exponential weighting to generate a dynamic noise estimate. A weighting factor α (0 < α < 1) can be set. For example, if α = 0.7, the dynamic noise estimate is: 0.7 × historical noise variance + 0.3 × current observed noise variance.
[0045] The dynamic noise estimate is fed into an adaptive threshold function to generate a threshold parameter that takes into account the spatial relationships between grid cells and the primary and secondary directions of order flow. A function based on distance decay and directional weighting can be designed to ensure that the threshold parameter reflects both spatial correlation and flow characteristics.
[0046] A hybrid thresholding operator is constructed by combining a bidirectional shrinkage function with a threshold parameter. This operator not only considers the shrinkage of abnormal fluctuations but also incorporates adaptive threshold control, enabling more precise denoising. The hybrid thresholding operator is then used to perform soft thresholding denoising on the multi-scale frequency coefficient matrix, yielding the denoised frequency coefficient matrix. Soft thresholding denoising can smoothly suppress noise while preserving the signal's key features.
[0047] Extract multi-scale order flow features from the multi-scale frequency coefficient matrix. Consider extracting directional and intensity features at each scale. For example, calculate metrics such as flow angle distribution and flow size distribution at different scales. Concatenate the multi-scale order flow features with the denoised frequency coefficient matrix. This step combines the original frequency information with the extracted high-level features to form a richer feature representation.
[0048] The feature concatenation results are reconstructed using an attention weighting mechanism to obtain a denoised correlation strength matrix. The attention mechanism automatically learns the importance of different features, highlighting key information during the reconstruction process. For example, a self-attention mechanism can be used to adaptively assign weights based on the correlation between features. Based on the denoised correlation strength matrix, a multi-layer probabilistic transfer network is constructed. This network comprises local, regional, and global transfer layers, each representing the order flow distribution at different spatial scales.
[0049] At the local transfer layer, the probability of order flow between adjacent grid cells can be considered. For example, for a 5×5 local area, the probability distribution of order flow from the central grid to the surrounding 24 grid cells is calculated. At the regional transfer layer, grid cells can be aggregated into larger regional blocks to calculate the probability of order flow between these blocks. For example, 25 grid cells can be aggregated into a single region, and the probability of order flow between these different regions can be calculated. At the global transfer layer, the probability of order flow between all grid cells within the entire study area is considered. This layer can capture long-range order flow patterns.
[0050] At each transfer layer, the entropy of order flows within a grid cell at the corresponding spatial scale is calculated. Entropy calculation can be based on the definition of Shannon entropy, reflecting the degree of uncertainty in the distribution of order flows. For example, if orders in a grid cell flow evenly in all directions, its entropy will be high; if orders are concentrated in a single direction, the entropy will be low. Through inter-layer information transfer, a hierarchical entropy metric is integrated. A bottom-up information transfer mechanism can be designed to gradually aggregate entropy information from lower layers to higher layers, forming a multi-scale entropy representation. Finally, a recurrent neural network unit is used to extract features and perform nonlinear combinations on the hierarchical entropy metric at different time scales. Recurrent neural networks can capture long-term dependencies in time series data and are suitable for processing time-varying data such as order flows. By modeling entropy values at different time scales (such as hours, days, and weeks), a comprehensive metric reflecting the uncertainty of information flows is constructed.
[0051] Figure 2 This is a simulation effect diagram of the order flow correlation strength matrix analysis according to an embodiment of the present invention. Figure 2 As shown, the performance of three different techniques is demonstrated by multiple sliding window segmentation and order peak distribution feature analysis. This solution demonstrates significant advantages when processing complex order flow data, particularly when the window size is between 15 and 25, achieving a peak identification accuracy of 93.7%. Traditional wavelet decomposition and filtering algorithms achieve only 78.2% and 65.4% accuracy in this range, respectively. At a window size of 35, the performance of this solution decreases slightly, but remains at a high level of 87.5%, while other methods exhibit more significant performance degradation. The figure also demonstrates that this solution significantly outperforms existing techniques across various window sizes, with a standard deviation of only 3.2%. This is attributed to its adaptive multi-scale analysis capabilities and improved dynamic calculation mechanism for the number of wavelet transform layers. Numerically, at a window size of 30, the feature extraction integrity of this solution reaches 95.8%, significantly exceeding the 82.3% of the wavelet decomposition method and the 71.9% of the filtering algorithm, demonstrating its superior performance in capturing order flow correlation features.
[0052] In this embodiment, a multi-scale analysis of the correlation strength of order flows is achieved through multiple sliding windows and wavelet decomposition technology, which can effectively capture the order flow patterns at different time and space scales, and improve the accuracy and comprehensiveness of the analysis. The hybrid threshold processing operator constructed by the adaptive threshold function and the bidirectional shrinkage function can automatically adjust the denoising parameters according to the dynamic characteristics of the order data, effectively remove noise interference, while retaining the key information of the order flow, improving the noise reduction effect and data quality. By constructing a multi-layer probability transfer network and a hierarchical entropy value indicator, combined with recursive neural networks for time series modeling, a multi-scale quantification and dynamic characterization of order flow uncertainty is achieved, providing a reliable decision-making basis for subsequent order forecasting and resource scheduling, and helping to improve operational efficiency and service quality.
[0053] In an optional embodiment, time-series order data is extracted to construct an order evolution pattern feature vector, which is then concatenated with a multidimensional feature vector to obtain an enhanced feature vector. The denoised association strength matrix, information flow uncertainty index, and enhanced feature vector are input into a graph attention network to obtain order propagation features including: Extracting the time series order data of the reconstructed grid cells, calculating an adaptive penalty factor based on the local fluctuation characteristics of the time series order data, constructing an optimization objective function with a penalty term using the adaptive penalty factor, and iteratively solving the optimization objective function to obtain a trend term, a period term, and a random term; The slope feature and acceleration feature of the trend item are calculated and weightedly combined to obtain the trend fluctuation intensity feature; the period item is subjected to a multi-scale Fourier transform to obtain a frequency domain feature sequence, the correlation coefficients between different frequency components are calculated to obtain a frequency correlation matrix, and the main period feature, periodic intensity feature, and periodic stability feature are extracted; and the feature combination forms the order evolution pattern feature vector; Calculating the importance score of each feature component in the order evolution pattern feature vector and the multidimensional feature vector, and selecting feature components that are higher than a preset feature selection threshold for splicing to obtain an enhanced feature vector; The grid units are used as graph network nodes. The spatial adjacency relationship and semantic association relationship are constructed based on the denoised association strength matrix. The spatial distance attention coefficient and feature similarity attention coefficient are calculated. The enhanced feature vector is combined with the information flow uncertainty index to form the initial feature of the node. The order propagation feature is obtained through multi-level information transmission and residual connection fusion.
[0054] For example, reconstructed time series order data is first extracted from the grid cells. This data can be order values that have been aggregated and cleaned based on geographic location, time period, or specific business dimensions. Based on the local fluctuation characteristics of the time series data, an adaptive penalty factor is calculated to penalize abnormal fluctuations. This factor balances the accuracy and smoothness of the data fit. It is introduced into the variational constrained optimization objective function to enhance the ability to capture the characteristics of the time series data.
[0055] During the optimization process, an objective function with a penalty term is employed, and an improved iterative solution method is used to decompose the data. Specifically, a nonlinear shrinkage operator is combined with the alternating direction multiplication method to make the solution more efficient and stable. Ultimately, the time series data is decomposed into three components: a trend term, a cyclic term, and a random term.
[0056] The trend term, representing the overall trend of time series data, is extracted through forward differencing to determine its rate of change, known as the trend slope feature. This feature is then subjected to reverse differencing to further extract the acceleration of the trend change, known as the trend acceleration feature. The trend slope and trend acceleration features are combined and their combined volatility is weighted to form the trend volatility intensity feature. This trend volatility intensity feature reflects the overall magnitude and frequency of changes in order data over time.
[0057] For periodic terms, their frequency domain characteristics are analyzed through multi-scale Fourier transforms to extract information about different frequency components. Based on the correlation between the frequency components in the frequency domain feature sequence, a frequency correlation matrix is calculated. The frequency domain feature sequence is weighted and fused using the frequency correlation matrix to extract the primary periodicity feature, the periodicity intensity feature, and the periodicity stability feature. The primary periodicity feature characterizes the primary periodicity of the data, the periodicity intensity feature reflects the significance of the periodicity, and the periodicity stability feature is used to assess the persistence and consistency of the periodicity.
[0058] The trend and cycle features are combined to form a feature vector for the order evolution pattern. This feature vector includes trend slope, trend acceleration, trend volatility intensity, primary cycle, cyclical intensity, and cyclical stability. Based on this feature combination, a feature selection threshold is set based on the importance scores of different feature components to select important feature components above the threshold. These selected feature components are then combined with other multidimensional features to form an enhanced feature vector.
[0059] Based on the denoised association strength matrix, the spatial adjacency and semantic association relationships of grid cells are constructed. The spatial distance attention coefficient and feature similarity attention coefficient between grid cells are calculated to enhance the modeling of grid cell association information. The enhanced feature vector is combined with the information flow uncertainty index to form the initial features of the grid cells.
[0060] A graph attention network is used to perform multi-level information propagation on the initial features of the grid cells. At each level of propagation, node features are weighted based on the spatial distance attention coefficient and the feature similarity attention coefficient. To avoid information loss during feature representation, residual connections are used to fuse the features after multi-level propagation. Finally, the fused features are weighted and aggregated to generate the order propagation features for the grid cells.
[0061] In this embodiment, by introducing an adaptive penalty factor and an improved solution operator, the accuracy and efficiency of variational mode decomposition are improved, and the trend and periodic characteristics of time-series order data can be captured more accurately. Through multi-scale Fourier transform and frequency correlation matrix analysis, this method can comprehensively extract the characteristics of periodic terms, including the main period, periodic intensity and periodic stability, thereby better characterizing the periodic variation pattern of order data. Using a graph attention network and a multi-level information transmission mechanism, this method can fully utilize the spatial relationship and semantic association between grid cells, effectively integrate multi-source heterogeneous features, and extract more representative and discriminative order propagation features, providing a reliable basis for subsequent order prediction and resource scheduling.
[0062] In an optional embodiment, external information is obtained to calculate the impact factor vector, and a nonlinear combination is performed with the order propagation characteristics to obtain the order demand forecast value, and the kernel density estimation method is used for correction. The demand forecast distribution obtained includes: The influence factor vector is concatenated with the order propagation feature to obtain an initial feature vector, the neighbor clustering of the local features in the initial feature vector is calculated to obtain the self-attention weight, the feature importance is calculated to obtain the cross-attention weight, and the initial feature vector is weightedly combined according to the self-attention weight and the cross-attention weight to obtain the weighted feature; Calculating the mutual information and conditional entropy of the weighted features, determining the branch merging weights and performing nonlinear combination to obtain the initial order demand forecast value of the grid unit; Performing time-frequency decomposition on the historical forecast deviation sequence to obtain multiple frequency components, calculating the fluctuation patterns of the multiple frequency components, calculating information entropy based on the fluctuation patterns, determining the optimal time window size, and performing stratified sampling on the historical forecast deviation sequence to obtain a forecast deviation sample set; Perform kernel density estimation based on the temporal correlation constraint and information divergence of the prediction deviation sample set to obtain an error probability density function; Calculate the statistical characteristics of the error probability density function, extract the local fluctuation characteristics of the initial order demand forecast value, calculate the corrected forecast value based on the statistical characteristics and the local fluctuation characteristics, repeatedly sample the error probability density function at the same time to obtain a multi-component site, evaluate the stability of the multi-component site to obtain a confidence interval, and combine the corrected forecast value with the confidence interval to obtain an order demand forecast distribution with a confidence interval.
[0063] For example, the impact factor vector and order propagation features are first obtained and concatenated to form an initial feature vector. The gradient cosine similarity matrix between features is calculated. When the similarity exceeds a preset threshold (e.g., 0.8), a branch is added to determine the number of branches. With a baseline depth of three layers, each additional branch increases the depth by one layer. For example, with three branches, the network depth is five layers. Dense skip connections are used within each branch, meaning that the output of each layer is directly connected to subsequent layers.
[0064] Perform nearest neighbor clustering analysis on the local features in the initial feature vector. Select a sliding window size of 5 and calculate the Euclidean distance between features. Features with a distance less than a threshold of 0.3 are grouped together. The inverse ratio between the cluster center and the distance between each feature and the cluster center is used as the self-attention weight. Simultaneously, calculate the mutual information between each feature and the target variable as the feature importance score, and normalize this to form the cross-attention weight. Multiply these two weights to obtain a combined weight, which is then weighted on the initial features to obtain a weighted feature vector.
[0065] Calculate the mutual information and conditional entropy for each pair of features in the weighted feature vector. Mutual information indicates the degree of information overlap between features, while conditional entropy indicates the uncertainty of one feature given the knowledge of another. Based on these two metrics, branch merging weights are designed. Features with greater mutual information and smaller conditional entropy are assigned greater branch merging weights. The outputs of different branches are combined using a weighted summation method to obtain the initial order demand forecast.
[0066] Wavelet decomposition is applied to the historical forecast deviation series to obtain different frequency components. Fluctuation characteristics, including peaks, troughs, and periodicity, are extracted for each frequency component. Sample entropy is calculated for each frequency component; higher entropy values indicate more complex fluctuations. Based on the sample entropy, the optimal time window size is determined, with the window size proportional to the maximum entropy value. Using the optimal window size, stratified sampling is performed on the historical deviation series, with 80% of the samples randomly selected from each stratum to form the forecast deviation sample set.
[0067] The Gaussian kernel function is selected as the base kernel function, and the kernel function's bandwidth matrix is initialized as a diagonal matrix. Based on the temporal autocorrelation of the prediction error samples and the KL divergence between samples, the bandwidth matrix is optimized via gradient descent. The kernel density is estimated using the optimized kernel function for the prediction error sample set to obtain the error probability density function.
[0068] Calculate the statistical characteristics of the error probability density function, including mean, variance, skewness, and kurtosis. Extract local fluctuation characteristics of the initial prediction value, such as the rate of change and acceleration between adjacent time points. Input these characteristics into a nonlinear correction model composed of a multi-layer perceptron to obtain the corrected prediction value. Simultaneously, resample the error probability density function 1000 times, extracting quantiles from 0.025 to 0.975 each time. Calculate the standard deviation of these quantiles. The final confidence interval is the interval of quantiles with a standard deviation less than a threshold of 0.1.
[0069] In this embodiment, the combination of a multi-branch feature network and an attention mechanism improves the efficiency of feature extraction and combination, better captures the nonlinear relationship between features, and improves prediction accuracy. The kernel density estimation method is used to model the prediction error, avoiding the limitations of the parameter distribution assumption, and can more accurately characterize the distribution characteristics of the prediction error, providing a reliable uncertainty estimate. The sampling strategy based on time-frequency analysis and information entropy ensures the representativeness and diversity of the samples, improves the stability and reliability of the prediction results, and intuitively displays the uncertainty range of the prediction results in the form of confidence intervals.
[0070] In an optional embodiment, the demand forecast distribution gradient between grid cells is calculated, converted into a motion guidance vector, and combined with the real-time traffic status to generate an optimal delivery path. When the backlog of orders exceeds a threshold, cross-grid scheduling and intelligent allocation are performed, including: Extract the multi-order statistical moments of the order demand forecast distribution of the grid cells, calculate the information entropy difference and kernel density distance between adjacent grid cells, construct an asymmetric migration probability matrix, and obtain the benchmark flow direction characteristics through tensor decomposition; Obtaining the geographic connectivity between grids and historical order flow sequences, extracting periodic flow patterns, and combining them with spatial autocorrelation coefficients to construct an adaptive weight network. The baseline flow features are input into the adaptive weight network to obtain the rider movement guidance vector. The historical traffic time series is obtained and subjected to wavelet packet decomposition to obtain multiple frequency band signals. The fluctuation characteristics and phase information of the multiple frequency band signals are extracted. A recursive state equation is constructed based on the fluctuation characteristics. The phase information is used to design an adaptive observation matrix. The recursive state equation and adaptive observation matrix are used to predict the real-time traffic status. The time-varying road network cost matrix is constructed in combination with the road network topology characteristics. The rider motion guidance vector is mapped into the driving force field to divide the sub-blocks. The overall optimal delivery path is obtained through collaborative optimization. Based on the long-range correlation and scale-dependence characteristics of the historical data of grid backlog orders, a memory decay function and an adaptive bandwidth kernel function are constructed to generate a dynamic threshold surface. When it is detected that the number of grid backlog orders exceeds the dynamic threshold surface, a cross-grid collaborative objective function is constructed based on the rider's historical performance and current load, and an optimization solution is performed based on the propagation entropy matrix and the grid reachability tensor to obtain a cross-grid scheduling strategy for intelligent order allocation.
[0071] For example, we first analyze the predicted distribution of order demand across grid cells. We extract the spatiotemporal distribution characteristics of orders for each grid cell, including order density, temporal distribution patterns, and spatial aggregation. Specifically, we select high-frequency sample points from historical order data and use kernel density estimation to calculate the order distribution of grid cells. For example, we divide a certain area in a city into 500-meter-by-500-meter grid cells. We select historical order data from the past three months and extract the order density distribution of each grid cell at different time periods.
[0072] To characterize information flow between adjacent grid cells, a sliding time window method is used to calculate the probability of order flow between grid cells. By analyzing the distribution patterns of order origins and destinations, an asymmetric migration probability matrix is constructed. For example, for any two adjacent grid cells A and B, the ratio of the number of orders from A to B to the total number of orders within a 15-minute time window is calculated to obtain the probability of order flow between grid cells.
[0073] When obtaining rider movement guidance vectors, we first analyze the geographic connectivity between grids. Based on the road network topology, we calculate inter-grid accessibility metrics. We also perform time series decomposition on historical order flow sequences to extract hourly, daily, and weekly cyclical patterns. For example, a certain grid cell may exhibit stable order flow patterns during weekday morning and evening peak hours.
[0074] For real-time traffic condition prediction, a multi-level time series analysis approach is employed. First, historical travel time data is preprocessed to remove outliers and noise. Then, the time series is decomposed to extract trend terms, cyclical terms, and random terms. Based on these extracted features, a prediction model is constructed to predict traffic conditions for future time periods. For example, historical travel times for a particular road section show that travel times increase by approximately 30% during the morning rush hour on rainy days.
[0075] For route planning, a time-varying weighted graph is first constructed based on the road network topology. Nodes represent intersections, and edges represent road segments. Edge weights are dynamically adjusted based on real-time traffic conditions. A hierarchical routing strategy is employed, first planning the general route at a macro level, then optimizing local segments at a micro level. For example, if a delivery route for an order spans multiple grid cells, the optimal connection between grid cells is first determined, followed by optimizing the specific delivery sequence within the grid.
[0076] To handle order backlogs, we built a dynamic threshold judgment mechanism. We analyzed the spatiotemporal distribution of order backlogs based on historical data, extracting long-term trends and short-term fluctuations. For example, if a certain business district experiences a 50% increase in orders during holidays compared to normal times, we could adjust the backlog judgment threshold accordingly.
[0077] Finally, during order allocation, the rider's historical service quality and current workload are comprehensively considered. Metrics such as rider efficiency and completion rate are extracted from historical delivery data to build a service quality evaluation system. When order backlogs occur, intelligent order allocation is performed based on rider evaluation results and workload status.
[0078] In this embodiment, by effectively predicting the order demand distribution and its changing trend between grid units, the rider movement guidance vector is accurately generated, thereby improving delivery efficiency. By combining the time-varying road condition prediction model with the real-time traffic status, a dynamic optimal delivery path is generated, significantly reducing delivery time and logistics costs. When orders are piled up, the scheduling needs between grids are flexibly judged through the dynamic threshold surface, and cross-grid collaborative scheduling is performed in combination with the order propagation characteristics to achieve efficient allocation and balance of resources. Combining the rider's historical performance and current load, the order allocation strategy is intelligently optimized to improve the rider's service quality and user experience. Adaptive multi-scale analysis and distributed optimization algorithms are used to ensure the robustness and real-time performance of the model in complex scenarios. The overall solution can coordinate grid resources and rider capabilities, improve the overall operating efficiency and responsiveness of the logistics system, and at the same time improve the system's dynamic scheduling flexibility and path planning accuracy.
[0079] Figure 3 This is a histogram of the collaborative optimization of the delivery path and the driving force field partition effect according to the embodiment of the present invention, such as Figure 3As shown in the figure, this compares the performance of different path planning methods in terms of delivery efficiency, path optimality, and computational complexity. This technical solution maps the rider's motion guidance vector into a driving force field and divides it into sub-blocks. This solution, combined with a time-varying road network cost matrix, performs collaborative optimization to obtain the overall optimal delivery path. Data shows that the average delivery time for this technical solution is 18.6 minutes, compared to 24.3 minutes for the traditional Dijkstra algorithm and 21.8 minutes for the A* algorithm. In terms of path length, the average path length for this technical solution is 4.2 kilometers, 5.9 kilometers for the Dijkstra algorithm, and 5.1 kilometers for the A* algorithm. Especially during peak hours, the congestion avoidance success rate for this technical solution reaches 93.7%, compared to only 58.2% for the Dijkstra algorithm and 67.5% for the A* algorithm. In terms of memory consumption, this technical solution requires only 248MB of memory to process 1000 grid nodes, while the Dijkstra algorithm requires 645MB and the A* algorithm requires 512MB. Through collaborative optimization among the 32 sub-blocks, this technical solution achieved a global task allocation balance of 0.96, significantly ahead of the Dijkstra algorithm's 0.71 and the A* algorithm's 0.78. During the actual delivery process, this technical solution considered the real-time traffic status of 86 key intersections and combined the spatiotemporal distance calculation between the rider's current location and the target location, resulting in an actual execution rate of 89.4% for the planned path, far exceeding the 65.7% and 72.9% of other algorithms, effectively resolving the problem of deviation from the theoretical optimal path in actual execution.
[0080] Figure 4 FIG. 1 is a structural diagram of a logistics distribution demand forecasting and analysis system based on big data according to an embodiment of the present invention. Figure 4 As shown, the system includes: The first unit is used to divide the logistics distribution area into multiple regular hexagonal grid cells, calculate the order density feature matrix of the grid cells, dynamically reconstruct the grid cells, perform feature extraction, construct a multi-dimensional feature vector of the grid cells, and calculate the order flow correlation strength matrix between the grid cells based on the multi-dimensional feature vector; The second unit is used to perform wavelet decomposition and soft threshold denoising on the order flow correlation strength matrix to obtain a denoised correlation strength matrix; based on the denoised correlation strength matrix, the entropy of the order flow within the reconstructed grid cells is calculated and an information flow uncertainty index is constructed; The third unit is used to extract time-series order data to construct the order evolution pattern feature vector, concatenate the order evolution pattern feature vector with the multi-dimensional feature vector to obtain the enhanced feature vector, and input the denoised association strength matrix, information flow uncertainty index, and enhanced feature vector into the graph attention network to obtain the order propagation characteristics; The fourth unit is used to obtain external information to calculate the influencing factor vector, and perform nonlinear combination with the order propagation characteristics to obtain the order demand forecast value. The kernel density estimation method is used for correction to obtain the demand forecast distribution, and the demand forecast distribution gradient between grid units is calculated and converted into a motion guidance vector. The optimal delivery path is generated in combination with the real-time traffic status. When the backlog of orders exceeds the threshold, cross-grid scheduling and intelligent allocation are performed.
[0081] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0082] According to a fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.
[0083] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A logistics distribution demand forecasting and analysis method based on big data, characterized in that: include: The logistics distribution area is divided into multiple regular hexagonal grid cells, the order density feature matrix of the grid cells is calculated, the grid cells are dynamically reconstructed and feature extraction is performed to construct a multi-dimensional feature vector of the grid cells, and the order flow correlation strength matrix between the grid cells is calculated based on the multi-dimensional feature vector. The order flow correlation strength matrix is subjected to wavelet decomposition and soft threshold denoising to obtain a denoised correlation strength matrix. The entropy of the order flow within the reconstructed grid cells is calculated based on the denoised correlation strength matrix, and an information flow uncertainty index is constructed. Extract time-series order data to construct the order evolution pattern feature vector. Concatenate the order evolution pattern feature vector with the multidimensional feature vector to obtain an enhanced feature vector. Input the denoised association strength matrix, information flow uncertainty index, and enhanced feature vector into the graph attention network to obtain order propagation characteristics. External information is obtained to calculate the influencing factor vector, and a nonlinear combination is performed with the order propagation characteristics to obtain the order demand forecast value. The kernel density estimation method is used for correction to obtain the demand forecast distribution. The demand forecast distribution gradient between grid units is calculated and converted into a motion guidance vector. The optimal delivery path is generated in combination with the real-time traffic status. When the backlog of orders exceeds the threshold, cross-grid scheduling and intelligent allocation are performed.
2. The method according to claim 1, characterized in that The logistics distribution area is divided into multiple regular hexagonal grid cells. The order density feature matrix of the grid cells is calculated as follows: Establish a coordinate system based on the geometric center of the logistics distribution area, calculate the length of the regular hexagon based on the area and the number of target grids, generate a grid center coordinate mapping table, use the ray method to determine the boundary point set, identify the adjacent grid identifier set, extract the path to construct a road network connection diagram, establish the mapping relationship between the grid and the road network, and generate a grid structure containing the grid center coordinates, boundary point set, adjacent grid identifiers and the road network connection diagram; Set a time window to collect the starting point coordinate sequence and timestamp sequence of orders within the grid unit. Calculate the order density per unit area based on the ratio of the number of orders to the grid coverage area. Obtain the GPS positioning trajectory of the rider and the delivery range polygon to calculate the rider density per unit area. Calculate the delivery capacity coefficient based on the load limit and historical order completion rate. Collect the real-time traffic volume and average speed of the road section to calculate the road congestion index. The order density per unit area, rider density per unit area, delivery capacity coefficient and road congestion index collected from each grid unit are normalized to their maximum and minimum values to obtain the corresponding standardized eigenvalue sequence. The information entropy is calculated based on each standardized eigenvalue sequence to determine the weight coefficient of each feature. The order density eigenvalue is obtained by weighted summation of the standardized eigenvalue and the corresponding weight coefficient, and organized into a grid unit order density feature matrix according to the grid unit identifier.
3. The method according to claim 1, characterized in that After dynamic reconstruction of the grid cells, feature extraction is performed to construct a multi-dimensional feature vector of the grid cells. Based on the multi-dimensional feature vector, the order flow correlation strength matrix between the grid cells is calculated, including: Obtain historical order data and the initial grid unit order density feature matrix, construct a grid unit adjacency relationship graph, use a graph neural network to extract structural features, calculate the dynamic weight coefficients of adjacent grid units, and weightedly fuse them with the order density feature value to obtain grid similarity; Calculating a grid reconstruction effect evaluation value based on historical order distribution, inputting the grid similarity and reconstruction effect evaluation value into a deep reinforcement learning model to obtain a grid reconstruction strategy; performing a grid reconstruction operation according to the grid reconstruction strategy, and constructing the order density distribution of the reconstructed grid unit into a spatiotemporal feature sequence; Based on the spatiotemporal feature sequence, the order change trend characteristics, grid spatial dependency characteristics, and distribution resource characteristics of each grid unit are extracted respectively, and the deep feature vector of the grid unit is obtained by adaptive weight fusion. Based on the distribution patterns of order starting points and destinations in historical order flow data, a knowledge base containing the spatiotemporal patterns of order flows is constructed; the comprehensive correlation weights of the time dimension, space dimension, and resource dimension are calculated based on the deep feature vectors of the grid units; based on the order flow patterns in the knowledge base and the comprehensive correlation weights between the grid units, the order flow correlation strength between the grid units is calculated, and an order flow correlation strength matrix is constructed.
4. The method according to claim 1, wherein Perform wavelet decomposition and soft threshold denoising on the order flow correlation strength matrix to obtain the denoised correlation strength matrix; The entropy of the order flow in the reconstructed grid unit is calculated based on the noise-reduced correlation intensity matrix and the information flow uncertainty index is constructed, including: The order flow correlation intensity matrix is segmented into multiple sliding windows. The number of wavelet decomposition layers is calculated based on the order peak distribution characteristics within each window after segmentation. A two-dimensional discrete wavelet transform is performed to obtain a multi-scale frequency coefficient matrix, and abnormal fluctuation components are identified through energy cluster analysis. Based on the abnormal fluctuation component, a bidirectional shrinkage function having a positive fluctuation shrinkage parameter and a negative fluctuation shrinkage parameter is constructed, a dynamic noise estimate is generated by combining the historical noise variance and the current observation noise variance, and the dynamic noise estimate is input into an adaptive threshold function to generate a threshold parameter; The bidirectional shrinkage function and the threshold parameter are combined to construct a hybrid threshold processing operator, and soft threshold denoising is performed on the multi-scale frequency coefficient matrix. The multi-scale order flow features are extracted from the multi-scale frequency coefficient matrix, and the features are cascaded with the denoised frequency coefficient matrix. The denoised association strength matrix is reconstructed using the attention weight mechanism; A multi-layer probability transfer network is constructed based on the denoised correlation intensity matrix. The order flow entropy of each grid unit is calculated separately. A hierarchical entropy index is formed through inter-layer information transmission and integration. A recursive neural network is used to extract time scale features and perform nonlinear combination to construct an information flow uncertainty index.
5. The method according to claim 1, characterized in that Extract time-series order data to construct the order evolution pattern feature vector. Concatenate the order evolution pattern feature vector with the multi-dimensional feature vector to obtain an enhanced feature vector. Input the denoised association strength matrix, information flow uncertainty index, and enhanced feature vector into the graph attention network to obtain the order propagation features including: Extracting the time series order data of the reconstructed grid cells, calculating an adaptive penalty factor based on the local fluctuation characteristics of the time series order data, constructing an optimization objective function with a penalty term using the adaptive penalty factor, and iteratively solving the optimization objective function to obtain a trend term, a period term, and a random term; The slope feature and acceleration feature of the trend item are calculated and weightedly combined to obtain the trend fluctuation intensity feature; the period item is subjected to a multi-scale Fourier transform to obtain a frequency domain feature sequence, the correlation coefficients between different frequency components are calculated to obtain a frequency correlation matrix, and the main period feature, periodic intensity feature, and periodic stability feature are extracted; and the feature combination forms the order evolution pattern feature vector; Calculating the importance score of each feature component in the order evolution pattern feature vector and the multidimensional feature vector, and selecting feature components that are higher than a preset feature selection threshold for splicing to obtain an enhanced feature vector; The grid units are used as graph network nodes. The spatial adjacency relationship and semantic association relationship are constructed based on the denoised association strength matrix. The spatial distance attention coefficient and feature similarity attention coefficient are calculated. The enhanced feature vector is combined with the information flow uncertainty index to form the initial feature of the node. The order propagation feature is obtained through multi-level information transmission and residual connection fusion.
6. The method according to claim 1, characterized in that Obtain external information to calculate the impact factor vector, and perform nonlinear combination with the order propagation characteristics to obtain the order demand forecast value. Then use the kernel density estimation method to perform correction, and obtain the demand forecast distribution including: The influence factor vector is concatenated with the order propagation feature to obtain an initial feature vector, the neighbor clustering of the local features in the initial feature vector is calculated to obtain the self-attention weight, the feature importance is calculated to obtain the cross-attention weight, and the initial feature vector is weightedly combined according to the self-attention weight and the cross-attention weight to obtain the weighted feature; Calculating the mutual information and conditional entropy of the weighted features, determining the branch merging weights and performing nonlinear combination to obtain the initial order demand forecast value of the grid unit; Performing time-frequency decomposition on the historical forecast deviation sequence to obtain multiple frequency components, calculating the fluctuation patterns of the multiple frequency components, calculating information entropy based on the fluctuation patterns, determining the optimal time window size, and performing stratified sampling on the historical forecast deviation sequence to obtain a forecast deviation sample set; Perform kernel density estimation based on the temporal correlation constraint and information divergence of the prediction deviation sample set to obtain an error probability density function; Calculate the statistical characteristics of the error probability density function, extract the local fluctuation characteristics of the initial order demand forecast value, calculate the corrected forecast value based on the statistical characteristics and the local fluctuation characteristics, repeatedly sample the error probability density function at the same time to obtain a multi-component site, evaluate the stability of the multi-component site to obtain a confidence interval, and combine the corrected forecast value with the confidence interval to obtain an order demand forecast distribution with a confidence interval.
7. The method according to claim 1, characterized in that Calculate the demand forecast distribution gradient between grid cells, convert it into a motion guidance vector, and generate the optimal delivery path based on real-time traffic status. When the backlog exceeds the threshold, cross-grid scheduling and intelligent allocation are performed, including: Extract the multi-order statistical moments of the order demand forecast distribution of the grid cells, calculate the information entropy difference and kernel density distance between adjacent grid cells, construct an asymmetric migration probability matrix, and obtain the benchmark flow direction characteristics through tensor decomposition; Obtaining the geographic connectivity between grids and historical order flow sequences, extracting periodic flow patterns, and combining them with spatial autocorrelation coefficients to construct an adaptive weight network. The baseline flow features are input into the adaptive weight network to obtain the rider movement guidance vector. The historical traffic time series is obtained and subjected to wavelet packet decomposition to obtain multiple frequency band signals. The fluctuation characteristics and phase information of the multiple frequency band signals are extracted. A recursive state equation is constructed based on the fluctuation characteristics. The phase information is used to design an adaptive observation matrix. The recursive state equation and adaptive observation matrix are used to predict the real-time traffic status. The time-varying road network cost matrix is constructed in combination with the road network topology characteristics. The rider motion guidance vector is mapped into the driving force field to divide the sub-blocks. The overall optimal delivery path is obtained through collaborative optimization. Based on the long-range correlation and scale-dependence characteristics of the historical data of grid backlog orders, a memory decay function and an adaptive bandwidth kernel function are constructed to generate a dynamic threshold surface. When it is detected that the number of grid backlog orders exceeds the dynamic threshold surface, a cross-grid collaborative objective function is constructed based on the rider's historical performance and current load, and an optimization solution is performed based on the propagation entropy matrix and the grid reachability tensor to obtain a cross-grid scheduling strategy for intelligent order allocation.
8. A logistics distribution demand forecasting and analysis system based on big data, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to divide the logistics distribution area into multiple regular hexagonal grid cells, calculate the order density feature matrix of the grid cells, dynamically reconstruct the grid cells, perform feature extraction, construct a multi-dimensional feature vector of the grid cells, and calculate the order flow correlation strength matrix between the grid cells based on the multi-dimensional feature vector; The second unit is used to perform wavelet decomposition and soft threshold denoising on the order flow correlation strength matrix to obtain the denoised correlation strength matrix; The entropy of the order flow in the reconstructed grid cells is calculated based on the correlation strength matrix after noise reduction and the information flow uncertainty index is constructed; The third unit is used to extract time-series order data to construct the order evolution pattern feature vector, concatenate the order evolution pattern feature vector with the multi-dimensional feature vector to obtain the enhanced feature vector, and input the denoised association strength matrix, information flow uncertainty index, and enhanced feature vector into the graph attention network to obtain the order propagation characteristics; The fourth unit is used to obtain external information to calculate the influencing factor vector, and perform nonlinear combination with the order propagation characteristics to obtain the order demand forecast value. The kernel density estimation method is used for correction to obtain the demand forecast distribution, and the demand forecast distribution gradient between grid units is calculated and converted into a motion guidance vector. The optimal delivery path is generated in combination with the real-time traffic status. When the backlog of orders exceeds the threshold, cross-grid scheduling and intelligent allocation are performed.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Delivery time determination method and electronic equipment
CN121052730A
Ecological environment condition index evaluation method based on ecological footprint model
CN121352488A
Logistics demand prediction method and system based on artificial intelligence
CN121961135A
Artificial intelligence-based logistics demand prediction method and system
CN121961135B
Regional integrated energy system material allocation prediction method and system based on regression algorithm
CN122047614A