Logistics management method and system for bank money box
By using an adaptive center clustering algorithm and a deep Q-network model, network points are dynamically grouped to generate the optimal escort plan, which solves the problem of low route planning efficiency in the transportation of bank cash boxes and improves transportation costs and security.
Patent Information
- Application Number
- CN202511488265.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-17
AI Technical Summary
In the current technology for managing bank cash box transportation, static route planning leads to vehicles frequently getting stuck in congested areas, resulting in low resource allocation efficiency, high empty running rates, fuel waste, and rising labor costs, making it difficult to adapt to complex and ever-changing real-world needs.
By using an adaptive center clustering algorithm and a deep Q-network model, network points are dynamically grouped to generate the optimal cash box escort plan. Combining historical transportation demand and passenger flow, resource allocation is optimized, and real-time responses to changes in transportation demand are made to reduce cross-group transportation costs and avoid route congestion.
It improves the efficiency and security of bank cash box transportation, dynamically adjusts route planning, reduces transportation costs, reduces risks caused by concentrated demand bursts within the cluster, and automates and improves the efficiency of solution generation.
Smart Images

Figure CN120952653A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of logistics management, and in particular to a logistics management method and system for bank cash boxes. Background Technology
[0002] Bank cash boxes are specialized, sealed containers used by banks for the safe storage and transport of high-value items such as cash, negotiable instruments, precious metals, and securities. They are a crucial tool in the bank's cash management chain, directly impacting fund security and operational efficiency.
[0003] The transportation management of bank cash boxes is a core element in ensuring fund security and business continuity. However, the static route planning method currently widely used in the industry is no longer suitable for the complex and ever-changing real-world needs, and its technical shortcomings are becoming a key bottleneck restricting efficiency and security. In the traditional model, transportation routes are usually pre-set based on historical experience or fixed administrative divisions, such as a cyclical pattern of "Branch A → Vault → Branch B," with strictly fixed departure schedules. This design leads to vehicles frequently getting stuck in congested areas, with single-trip travel time fluctuating by ±50%; at the same time, static planning results in low resource allocation efficiency, with empty-running rates exceeding 30%, and significant problems of fuel waste and rising labor costs.
[0004] Therefore, there is a need to provide logistics management methods and systems for bank cash boxes to improve the efficiency of their transportation. Summary of the Invention
[0005] This invention provides a logistics management method for bank cash boxes, comprising: acquiring historical cash box transportation needs and historical passenger flow of multiple branches; dividing the multiple branches into multiple branch groups based on the historical cash box transportation needs of the multiple branches and the shortest path between any two branches; determining the weights of multiple handover time periods corresponding to various key date types for each branch based on the historical passenger flow of the multiple branches; acquiring the current cash box transportation needs of the multiple branches; and generating an optimal cash box escort plan based on the current cash box transportation needs of the multiple branches, the multiple branch groups, and the weights of multiple handover time periods corresponding to various key date types for each branch, wherein the optimal cash box escort plan includes multiple cash box escort routes, and each logistics transportation route corresponds to cash box transportation of at least one branch.
[0006] Furthermore, based on the historical cash box transportation needs of multiple network points and the shortest path between any two network points, the multiple network points are divided into multiple network point groups, including: calculating the demand similarity between any two network points based on the historical cash box transportation needs of multiple network points; determining similar and dissimilar network points for each network point based on the demand similarity between any two network points; for each network point, calculating the demand correlation coefficient between the network point and each dissimilar network point based on the network point's historical cash box transportation needs and the historical cash box transportation needs of each dissimilar network point; and using an adaptive center clustering algorithm, dividing the multiple network points into multiple network point groups based on the similar and dissimilar network points for each network point, the demand correlation coefficient between the network point and each dissimilar network point, and the shortest path between any two network points.
[0007] Furthermore, using an adaptive center clustering algorithm, based on the similarity and dissimilarity of each node, the demand correlation coefficient between a node and each dissimilar node, and the shortest path between any two nodes, multiple nodes are divided into multiple node groups, including: S11, calculating the demand similarity fluctuation value of each node based on the demand similarity between any two nodes, calculating the demand correlation fluctuation value of each node based on the demand correlation coefficient between a node and each dissimilar node, and calculating the adaptive center value of a node based on the number of dissimilar nodes, the demand similarity fluctuation value, and the demand correlation fluctuation value; S12, determining multiple nodes as cluster centers based on the adaptive center value of each node; S13, for each For nodes that are not cluster centers, if the cluster center is a dissimilar node, then the cluster center is used as a candidate cluster center. Based on the correlation coefficient and shortest path between the node and the candidate cluster center, the cluster distance between the node and the candidate cluster center is calculated, and the node is assigned to the candidate cluster center with the shortest cluster distance. S14: Determine if the inner loop clustering condition is met. If yes, proceed to S16; otherwise, proceed to S15. S15: Update the cluster centers and proceed to S13. S16: Determine if the outer loop clustering condition is met. If yes, complete the clustering and divide multiple nodes into multiple node groups; otherwise, proceed to S17. S17: Determine the multi-objective clustering function, determine the update direction, update the number of cluster centers based on the update direction, and proceed to S12.
[0008] Furthermore, based on the historical passenger flow of multiple outlets, the weights of multiple transition time periods corresponding to various key date types for each outlet are determined, including: determining various date types; calculating the passenger flow similarity between any two date types based on the historical passenger flow of multiple outlets, merging the various date types to determine various key date types; for each key date type, calculating the initial weights of multiple transition time periods corresponding to each key date type for each outlet based on the historical passenger flow of multiple outlets; iteratively correcting the initial weights of multiple transition time periods corresponding to each key date type for each outlet based on the demand correlation coefficient between each outlet and each dissimilar outlet, generating the weights of multiple transition time periods corresponding to various key date types for each outlet.
[0009] Furthermore, based on the historical passenger flow of multiple outlets, the passenger flow similarity between any two date types is calculated, and multiple date types are merged to determine multiple key date types, including: for each outlet, based on the outlet's historical passenger flow, the passenger flow similarity between any two date types corresponding to the outlet is calculated; based on the passenger flow similarity between any two date types corresponding to each outlet, multiple date types are merged to determine multiple key date types.
[0010] Furthermore, based on the demand correlation coefficients between each network point and its similar and dissimilar network points, the initial weights of multiple handover time periods corresponding to key date types for each network point are iteratively corrected to generate weights for multiple handover time periods corresponding to multiple key date types for each network point. This includes: S21, for each network point, based on the demand correlation coefficients between the network point and each dissimilar network point, determining the positively and negatively correlated network points for the network point, and correcting the initial weights of the handover time periods corresponding to the network point based on the initial weights of the similar, positively correlated, and negatively correlated network points for the network point, generating corrected weights for the handover time periods corresponding to the network point; S22, calculating the global corrected distance based on the initial weights and corrected weights of the handover time periods corresponding to each network point; S23, determining whether the iteration termination condition is met based on the global corrected distance. If yes, generating weights for multiple handover time periods corresponding to multiple key date types for each network point; otherwise, proceeding to S24; S24, using the corrected weights of the handover time periods corresponding to the network point as the initial weights of the handover time periods corresponding to the network point, and proceeding to S21.
[0011] Furthermore, based on the current cash and container transportation needs of multiple network points, multiple network point groups, and the weights of multiple handover time periods corresponding to various key date types for each network point, an optimal logistics transportation plan is generated. This includes: modeling the problem as a reinforcement learning task, where the state includes the current cash and container transportation needs of the remaining network points, multiple network point groups, the current location of the vehicle, the current key date type, the weights of multiple handover time periods corresponding to the current key date type for each network point, and the shortest travel time of the road connecting any two network points in the multiple handover time periods; the action is the next network point to visit; establishing and training a deep Q-network model; generating multiple initial network points based on the current cash and container transportation needs of multiple network points and multiple network point groups; and generating an optimal logistics transportation plan based on the multiple initial network points using the deep Q-network model.
[0012] Furthermore, based on the current cash and box transportation needs of multiple network points and multiple network point groups, multiple initial network points are generated. Using a deep Q-network model, an optimal logistics transportation plan is generated based on these initial network points, including: S31, determining the remaining total demand for each network point group based on the current cash and box transportation needs of multiple remaining network points; determining the current network point group based on the remaining total demand of each network point group; and determining the current initial network point based on the current cash and box transportation needs of each remaining network point included in the current network point group; S32, the deep Q-network model outputs the Q-value of each action based on the current state, and selects the action with the highest Q-value; S33, updating the state and determining whether the single-path termination condition is met. If yes, proceed to S34; otherwise, proceed to S32; S34, determining whether the plan generation termination condition is met. If yes, generate the optimal logistics transportation plan; otherwise, proceed to S31.
[0013] Furthermore, the reward function of the deep Q-network model is at least related to whether the next node and the current node belong to the same node group, the shortest travel time from the current node to the next node, and the handover time period of the next node.
[0014] This invention provides a logistics management system for bank cash boxes, used to execute the aforementioned logistics management method for bank cash boxes, comprising: a data acquisition module for acquiring historical cash box transportation needs and historical passenger flow of multiple branches; a branch grouping module for dividing multiple branches into multiple branch groups based on the historical cash box transportation needs of multiple branches and the shortest path between any two branches; a weight determination module for determining the weight of multiple handover time periods corresponding to multiple key date types for each branch based on the historical passenger flow of multiple branches; a demand acquisition module for acquiring the current cash box transportation needs of multiple branches; and a transportation management module for generating an optimal cash box escort plan based on the current cash box transportation needs of multiple branches, the multiple branch groups, and the weight of multiple handover time periods corresponding to multiple key date types for each branch, wherein the optimal cash box escort plan includes multiple cash box escort routes, and each logistics transportation route corresponds to cash box transportation at least one branch.
[0015] Compared with existing technologies, the logistics management method and system for bank cash boxes provided by this invention have at least the following beneficial effects:
[0016] 1. Based on historical cash box transportation needs and the shortest paths between branches, branches are grouped to make them more interconnected in terms of transportation needs and geographical location, reducing cross-group transportation costs and improving overall logistics efficiency. The weight of handover time periods for each branch under different key date types is determined based on historical passenger flow, accurately reflecting the business busyness of each time period. This makes the generated escort plans more aligned with actual business and security needs, optimizing resource allocation. The optimal escort plan is generated by acquiring current cash box transportation needs and combining grouping and weighting information. This allows for real-time response to changes in branch transportation needs, dynamically adjusting route planning to ensure the effectiveness and timeliness of the plan, and improving the bank's cash box logistics management level.
[0017] 2. Multi-objective clustering functions achieve decentralized grouping of demand patterns through dual constraints of demand similarity / relevance and cluster number control, thereby reducing the risk of concentrated demand surges within clusters during transportation scheduling. By minimizing the average demand similarity from points within a cluster to the cluster center, the demand patterns of points within a cluster are forced to differ significantly. Within clusters with dispersed demand patterns, peak demand times for points are staggered, allowing for dynamic adjustment of access order during scheduling to avoid path congestion caused by concentrated access. Traditional clustering may lead to insufficient vehicles due to simultaneous demand surges within clusters, requiring frequent use of backup vehicles; with decentralized demand, such situations are significantly reduced.
[0018] 3. The cash box escort problem is modeled as a reinforcement learning task. The state design comprehensively covers key information such as remaining demand, network group, vehicle location, date type, handover time period weight, and road travel time. This allows for dynamic adaptation to changes in transportation demand under different scenarios, improving the flexibility and accuracy of solution generation. By learning the optimal action (next network point visit) selection strategy through a deep Q-network model and combining it with road travel time predictions based on historical data, the escort route can be effectively optimized, reducing transportation time and costs and improving logistics efficiency. Initial network points are determined based on the total remaining demand of the network group, and then the deep Q-network model generates paths layer by layer. This satisfies both overall transportation demand and the service requirements of individual network points, ensuring comprehensive coverage and reasonable feasibility of the solution. The deep Q-network model automatically outputs the Q-value and selects the optimal action, iteratively updating the state until the termination condition is met, achieving automation and efficiency in solution generation, reducing manual intervention, and minimizing decision-making errors. Attached Figure Description
[0019] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0020] Figure 1 This is a flowchart illustrating a logistics management method for bank cash boxes according to some embodiments of this specification;
[0021] Figure 2 This is a flowchart illustrating the process of dividing multiple outlets into multiple outlet groups according to some embodiments of this specification;
[0022] Figure 3 This is a schematic diagram of a logistics management system for bank cash boxes, shown according to some embodiments of this specification. Detailed Implementation
[0023] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0024] Figure 1 This is a flowchart illustrating a logistics management method for bank cash boxes according to some embodiments of this specification, such as... Figure 1 As shown, the logistics management method for bank cash boxes may include the following steps.
[0025] Step 1: Obtain historical cash box transportation needs and historical passenger flow from multiple locations.
[0026] Specifically, the historical cash box transportation demand of the outlets can include the cash box transportation demand volume of multiple past periods (e.g., daily, weekly, etc.), as shown in Table 1.
[0027] Branch ID date Demand Date type handover period Branch 1 2023-10-01 8 weekdays 06:00-07:00 Branch 2 2023-10-01 3 weekdays 12:00-13:00
[0028] The historical customer traffic of a branch can include the number of customers visiting over multiple historical time periods, as shown in Table 2.
[0029] Table 2
[0030] Step 2: Based on the historical cash box transportation needs of multiple outlets and the shortest path between any two outlets, divide the multiple outlets into multiple outlet groups.
[0031] Specifically, it includes: Based on the historical cash box transportation needs of multiple outlets, calculate the demand similarity between any two outlets; Based on the demand similarity between any two network points, determine the similar and dissimilar network points for each network point; For each branch, based on the branch's historical cash box transportation demand and the historical cash box transportation demand of each dissimilar branch, calculate the demand correlation coefficient between the branch and each dissimilar branch. Using an adaptive center clustering algorithm, multiple network points are divided into multiple network point groups based on the similar and dissimilar network points of each network point, the demand correlation coefficient between each network point and each dissimilar network point, and the shortest path between any two network points.
[0032] The demand similarity between two outlets measures the degree of similarity in their historical cash box transportation demand. The cosine similarity of the historical cash box transportation demand between the two outlets can be calculated as their demand similarity.
[0033] For each site, if there is a site whose demand similarity is greater than the demand similarity threshold (e.g., 0.6), the site is considered a similar site. The demand similarity threshold can be determined through experimental data.
[0034] For each branch, the historical cash box transportation demand of the branch and the historical cash box transportation demand of dissimilar branches can be substituted into the correlation coefficient calculation formula (e.g., Pearson correlation coefficient, Kendall rank correlation coefficient, etc.) to calculate the demand correlation coefficient between the branch and dissimilar branches.
[0035] Figure 2 This is a flowchart illustrating the process of dividing multiple outlets into multiple outlet groups according to some embodiments of this specification, such as... Figure 2 As shown, preferably, an adaptive center clustering algorithm is used to divide multiple network points into multiple network point groups based on the similar and dissimilar network points of each network point, the demand correlation coefficient between each network point and each dissimilar network point, and the shortest path between any two network points. These groups include: S11. Based on the demand similarity between any two network points, calculate the demand similarity fluctuation value of each network point. Based on the demand correlation coefficient between the network point and each dissimilar network point, calculate the demand correlation fluctuation value of each network point. Based on the number of dissimilar network points, the demand similarity fluctuation value, and the demand correlation fluctuation value, calculate the adaptive center value of the network point. S12. Based on the adaptive center value of each point, determine multiple points as cluster centers; S13. For each non-cluster center point, if the cluster center is a dissimilar point to the point, then the cluster center is used as a candidate cluster center. Based on the demand correlation coefficient and shortest path between the point and the candidate cluster center, the cluster distance between the point and the candidate cluster center is calculated, and the point is assigned to the candidate cluster center with the shortest cluster distance. S14. Determine whether the inner loop clustering condition is met. If yes, execute S16; otherwise, execute S15. S15. Update the cluster center, then execute S13; S16. Determine whether the outer loop clustering conditions are met. If yes, complete the clustering and divide the multiple nodes into multiple node groups. If no, execute S17. S17. Determine the multi-objective clustering function, determine the update direction, update the number of cluster centers based on the update direction, and execute S12.
[0036] Specifically, for each branch, the standard deviation of the demand similarity between that branch and any other branch is calculated, and this standard deviation is used as the demand similarity fluctuation value for that branch. The standard deviation of the demand correlation coefficient between that branch and each dissimilar branch is calculated, and this standard deviation is used as the demand correlation fluctuation value for that branch.
[0037] The adaptive center value of the network points can be calculated based on the following formula:
[0038] in, Let i be the adaptive center value of the i-th point. Let i be the similarity fluctuation value of demand at the i-th branch. Let i be the demand-related fluctuation value of the i-th branch. Let i be the number of dissimilar nodes of the i-th node. This represents the total number of outlets.
[0039] Based on the adaptive center value of each point, multiple points are sorted, and the top K points are selected as cluster centers. The value of K can be determined by the elbow rule, contour coefficient, etc.
[0040] The smaller the correlation coefficient between the network points and the candidate cluster centers, the shorter the shortest path, and the shorter the clustering distance between the network points and the candidate cluster centers.
[0041] The inner loop clustering condition can be that the number of inner loops reaches the maximum number of loops or the number of updated cluster centers in multiple loops is less than a number threshold (e.g., 3, etc.).
[0042] For each node within each cluster, calculate the standard deviation of the demand similarity between that node and any other node in the cluster, and use this as the intra-cluster demand similarity fluctuation value for that node. Calculate the standard deviation of the demand correlation coefficient between that node and each dissimilar node in the cluster, and use this as the intra-cluster demand correlation fluctuation value for that node. Based on the intra-cluster demand similarity fluctuation value, the intra-cluster demand correlation fluctuation value, and the number of dissimilar nodes within the cluster, calculate the intra-cluster adaptive centrality value for that node. The calculation methods for the intra-cluster adaptive centrality value and the adaptive centrality value of a node are similar and will not be repeated here. Update the node with the largest intra-cluster adaptive centrality value as the new cluster center.
[0043] The outer loop clustering conditions can be: the number of outer loop iterations reaches the maximum number of iterations, the value of the target clustering function is greater than a preset threshold, or the value of the target clustering function reaches its maximum value.
[0044] Multi-objective clustering functions are used to measure the quality of clustering results. By quantifying multiple conflicting clustering objectives (e.g., intra-cluster compactness, inter-cluster separation, path cost, etc.), they provide optimization directions and termination conditions for the clustering process. As an example only, a multi-objective clustering function may include the multiple parts shown in Table 3.
[0045] Target explain Optimization direction significance Intra-cluster demand similarity target Average demand similarity from intra-cluster nodes to cluster center minimize Ensure that the demand patterns of nodes within the cluster are relatively similar. Inter-cluster similarity target Average demand similarity between different cluster centers maximize Ensure that outlets with similar demand patterns are distributed across different clusters. Cluster-related requirements Average demand correlation coefficient from intra-cluster points to cluster center minimize Ensure that the demand patterns of nodes within the cluster have a low degree of correlation. Inter-cluster requirements related objectives Average demand correlation coefficient between different cluster centers maximize Ensure that outlets with similar demand patterns are distributed across different clusters. Cluster number target Nonlinear penalty for cluster number minimize Prevent over-grouping
[0046] Multi-objective clustering functions achieve decentralized grouping of demand patterns through dual constraints of demand similarity / correlation and cluster number control, thereby reducing the risk of concentrated demand surges within clusters during transportation scheduling. By minimizing the average demand similarity from points within a cluster to the cluster center, it forces significant differences in demand patterns among points within a cluster (e.g., one point has high demand during the day, another at night), avoiding the clustering of "homogeneous" points. By maximizing the average demand similarity between different cluster centers, it ensures that points with similar demand patterns are distributed across different clusters (e.g., all points with "high morning peak" demand are evenly distributed across multiple clusters). Correlation measures the synergy of demand changes (e.g., whether point B's demand increases synchronously when point A's demand rises). By constraining correlation, it further disperses points with synchronous demand fluctuations. Cluster number targets limit the number of clusters, avoiding over-grouping. If the demand patterns of points within a cluster differ significantly (e.g., a mixture of morning peak, evening peak, and stable demand), a single transport operation does not need to simultaneously address the peak demand of all points, resulting in a more balanced vehicle load. Within clusters with dispersed demand patterns, peak demand times for service points are staggered, allowing for dynamic adjustment of access order during scheduling to avoid path congestion caused by concentrated access. Traditional clustering may lead to insufficient vehicles due to simultaneous surges in demand within the cluster, necessitating frequent deployment of backup vehicles; with decentralized demand, such situations are significantly reduced.
[0047] Step 3: Based on the historical passenger flow of multiple outlets, determine the weight of multiple handover time periods corresponding to various key date types for each outlet.
[0048] Preferably, step 3 specifically includes: Define various date types, such as weekdays (Monday to Friday), weekends (Saturday and Sunday), statutory holidays (such as Spring Festival and National Day), and special dates (such as Singles' Day, school winter and summer vacations). Based on the historical passenger flow of multiple outlets, calculate the passenger flow similarity between any two date types, merge multiple date types, and determine multiple key date types; For each key date type, based on the historical customer traffic of multiple outlets, the initial weights of multiple handover time periods corresponding to each key date type are calculated for each outlet. The handover time period can be the time period during which the outlet can carry out cash box handover. Based on the demand correlation coefficients between similar sites and dissimilar sites for each site, the initial weights of multiple handover time periods corresponding to key date types for each site are iteratively corrected to generate weights for multiple handover time periods corresponding to multiple key date types for each site.
[0049] As a preferred approach, based on historical passenger flow data from multiple locations, the passenger flow similarity between any two date types is calculated. Multiple date types are then merged to determine several key date types, including: For each branch, based on the branch's historical passenger flow, calculate the passenger flow similarity for any two date types corresponding to the branch;
[0050] Based on the similarity of passenger flow for any two date types at each branch, multiple date types are merged to determine several key date types.
[0051] For each location, the cosine similarity of the mean passenger flow sequence for each time period of the day under both date types is calculated, and this cosine similarity is used as the passenger flow similarity for the location under the two date types.
[0052] For each branch, the average passenger flow similarity for the two date types is calculated and used as the global passenger flow similarity for the two date types. Date types with a global passenger flow similarity greater than the global passenger flow similarity threshold (e.g., 0.6) are merged into one, and the remaining date types after merging are designated as the key date types.
[0053] For each key date type, the ratio of the difference between the historical average total passenger flow for a day and the historical average passenger flow for a certain handover period to the historical average total passenger flow for a day is calculated and used as the initial weight of the handover period. The smaller the initial weight of a certain handover period, the higher the passenger flow during that period, which may be accompanied by problems such as mixed crowds and noisy environment, increasing the risk of theft, robbery, or surveillance lapses during the handover of cash boxes.
[0054] Preferably, based on the demand correlation coefficients between similar and dissimilar locations for each location, the initial weights of multiple handover time periods corresponding to key date types for each location are iteratively corrected to generate weights for multiple handover time periods corresponding to various key date types for each location, including: S21. For each network point, based on the demand correlation coefficient between the network point and each dissimilar network point, determine the positively correlated network points and negatively correlated network points of the network point. Based on the initial weights of the handover time periods corresponding to the network point's similar network points, positively correlated network points and negatively correlated network points, correct the initial weights of the handover time periods corresponding to the network point, and generate the corrected weights of the network point's handover time periods. S22. Calculate the global corrected distance based on the initial weight and the corrected weight of the handover time period corresponding to each network point; S23. Based on the global correction distance, determine whether the iteration termination condition is met. If yes, generate the weights of multiple handover time periods corresponding to various key date types for each network point. If no, execute S24. S24. Use the corrected weight of the handover time period corresponding to the branch as the initial weight of the handover time period corresponding to the branch, and execute S21.
[0055] Specifically, outlets with a demand correlation coefficient greater than the positive demand correlation coefficient threshold (e.g., 0.5) can be considered as positively correlated outlets, while outlets with a demand correlation coefficient less than the negative demand correlation coefficient threshold (e.g., -0.5) can be considered as negatively correlated outlets.
[0056] If a similar branch has a high initial weight for a certain handover time period, it indicates that this handover time period may better align with regional business patterns; therefore, the weight of the target branch corresponding to this handover time period should be increased. Similarly, if a positively correlated branch has a high initial weight for a certain handover time period, it suggests that this handover time period may better align with regional business patterns; therefore, the weight of the target branch corresponding to this handover time period should be increased. Conversely, if a negatively correlated branch has a high initial weight for a certain handover time period, it may reflect lower demand at the target branch during that handover time period; therefore, the weight of the target branch corresponding to this handover time period should be decreased.
[0057] The initial weights of the corresponding handover time periods of a network point can be corrected in any way based on the initial weights of similar, positively correlated, and negatively correlated network points. For example, a first correction coefficient can be calculated based on the initial weights of similar network points, a second correction coefficient based on the initial weights of positively correlated network points, and a third correction coefficient based on the initial weights of negatively correlated network points. These three correction coefficients are then used to correct the initial weights of the corresponding handover time periods, generating the corrected weights. As an example, the first correction coefficient can be determined based on the initial weights of similar network points according to preset rules. These preset rules could be:
[0058] in, This is the first correction factor. This is the difference between the mean of the initial weights of similar network points corresponding to the handover time period and the initial weights of the network points corresponding to the handover time period. The determination methods for the second and third correction coefficients are similar to those for the first correction coefficient, and will not be repeated here. The first, second, and third correction coefficients can be summed to obtain the overall correction coefficient. The product of the overall correction coefficient and the initial weights of the network points corresponding to the handover time period is calculated as the correction amount. The correction amount is then summed with the initial weights of the network points corresponding to the handover time period to obtain the corrected weights of the network points corresponding to the handover time period.
[0059] For each node, calculate the absolute value of the difference between the initial weight and the corrected weight for the corresponding handover time period, and use this as the corrected distance for the node. Sum the corrected distances of all nodes to obtain the global corrected distance.
[0060] The iteration termination condition can be that the global correction distance is less than the iteration termination condition threshold (e.g., 0.3), or the number of iterations reaches the maximum number of iterations.
[0061] Step 3 calculates and merges the passenger flow similarity between any two date types at each branch, accurately identifying date types with similar passenger flow patterns. This avoids redundant management caused by overly detailed date type classifications and ensures that key date types accurately reflect passenger flow characteristics under different business scenarios, improving the targeting of subsequent weight calculations and escort plan generation. The merged key date types better align with actual business cycle changes, enabling the generated escort plans to better adapt to passenger flow fluctuations on different dates. For example, merging similar weekday types allows for unified planning of weekday security escort strategies, improving the applicability and stability of the plans across different dates. Iteratively correcting the initial weights of target branches based on the initial weights of similar, positively correlated, and negatively correlated branches fully considers the business relevance and demand differences between branches, eliminating the impact of single-branch data deviations on weights and ensuring that the generated weights more accurately reflect the actual security risks at each handover period.
[0062] Step 4: Obtain the current cash box transportation needs of multiple outlets.
[0063] Step 5: Based on the current cash box transportation needs of multiple outlets, the weights of multiple outlet groups and multiple handover time periods corresponding to various key date types for each outlet, generate the optimal cash box escort plan.
[0064] The optimal cash box escort plan includes multiple cash box escort routes, with each logistics transportation route corresponding to cash box transportation at at least one network point.
[0065] Preferably, step 5 specifically includes: The problem is modeled as a reinforcement learning task, where the state includes the current cash box transportation demand of the remaining network points, multiple network point groups, the current location of the vehicle, the current key date type, the weight of multiple handover time periods for each network point corresponding to the current key date type, and the shortest travel time of the road connecting any two network points in multiple handover time periods. The action is the next network point to be visited, the remaining network points are unserved network points, and the shortest travel time of the road connecting any two network points in multiple handover time periods can be predicted based on historical data. Build and train a deep Q-network model;
[0066] Based on the current cash and container transportation needs of multiple network points and multiple network point groups, multiple initial network points are generated. Based on the multiple initial network points, the optimal logistics transportation plan is generated using a deep Q-network model.
[0067] As a preferred approach, based on the current container transportation needs of multiple network points and multiple network point groups, multiple initial network points are generated. Then, using a deep Q-network model, an optimal logistics transportation plan is generated based on these initial network points, including: S31. Based on the current cash box transportation needs of multiple remaining outlets, determine the total remaining needs of each outlet group. Based on the total remaining needs of each outlet group, determine the current outlet group. Based on the current cash box transportation needs of each remaining outlet included in the current outlet group, determine the current initial outlet. S32. The deep Q-network model outputs the Q-value of each action based on the current state and selects the action with the highest Q-value. S33. Update the status and determine whether the single-path termination condition is met. If yes, execute S34; otherwise, execute S32. S34. Determine whether the termination condition for generating the solution is met. If yes, generate the optimal logistics transportation solution. If not, execute S31.
[0068] The remaining total demand of a network group is the sum of the current cash box transportation demands of the remaining network points included in the network group. The network group with the largest remaining total demand can be designated as the current network group, and the remaining network points within the current network group with the largest current cash box transportation demands can be designated as the current initial network points.
[0069] A valid action selected from the remaining locations must satisfy the following conditions: 1. Vehicle capacity constraint, that is, the number of cash boxes loaded in the vehicle shall not exceed the upper limit; 2. Time window constraint, that is, the arrival time at the branch must be within the time period when it can be handed over.
[0070] As a preferred embodiment, the reward function of the deep Q-network model is at least related to whether the next node and the current node belong to the same node group, the shortest travel time from the current node to the next node, and the handover time period of the next node.
[0071] For example, the reward function of a deep Q-network model can be:
[0072] in, As a reward, This is the current state. For action, The state after the action is performed. , , and As weight, , , and Greater than 0, ,For example, , , and The values are 0.2, 0.2, 0.3, and 0.3 respectively. As a reward for belonging to the branch group, The reward is for the shortest travel time; the shorter the shortest travel time from the current branch to the next branch, the larger the reward. The weight of the handover time period to which the arrival time at the next branch belongs. Basic rewards, such as fixed cost penalties or fixed rewards for completing deliveries, This represents the similarity of needs between the current branch and the next branch.
[0073] Status updates can include adding the selected action (next stop) to the current path, and updating status variables such as the remaining set of stops, remaining vehicle capacity, and cumulative cost.
[0074] Single-path termination conditions may include: 1. Vehicle capacity is full; 2. No more sites are available for access; 3. The maximum path length limit is reached (such as the maximum number of network points or the maximum distance).
[0075] One of the conditions must be met, which means the single-path termination condition must be satisfied.
[0076] The termination conditions for plan generation may include: all points of interest are assigned to a certain route, or the total demand for remaining points of interest is less than the total amount that can be transported in a single trip.
[0077] If the termination condition for solution generation is met, output the set of all current paths as candidate solutions.
[0078] Step 5 models the cash box escort problem as a reinforcement learning task. The state design comprehensively covers key information such as remaining demand, network group, vehicle location, date type, handover time period weight, and road travel time. This allows for dynamic adaptation to changes in transportation demand under different scenarios, improving the flexibility and accuracy of solution generation. By learning the optimal action (next network point visit) selection strategy through a deep Q-network model and combining it with road travel time predictions from historical data, the escort route can be effectively optimized, reducing transportation time and costs and improving logistics efficiency. Initial network points are determined based on the total remaining demand of the network group, and then the deep Q-network model generates paths layer by layer. This satisfies both overall transportation demand and the service requirements of individual network points, ensuring comprehensive coverage and feasibility of the solution. The deep Q-network model automatically outputs the Q-value and selects the optimal action, iteratively updating the state until the termination condition is met, achieving automation and efficiency in solution generation, reducing manual intervention, and minimizing decision-making errors.
[0079] Figure 3 This is a schematic diagram of a logistics management system for bank cash boxes, shown according to some embodiments of this specification, such as... Figure 3 As shown, the logistics management system for bank cash boxes may include a data acquisition module, a branch grouping module, a weight determination module, a demand acquisition module, and a transportation management module.
[0080] The data acquisition module is used to acquire historical cash box transportation needs and historical passenger flow from multiple outlets; The branch grouping module is used to divide multiple branches into multiple branch groups based on the historical cash box transportation needs of multiple branches and the shortest path between any two branches; The weight determination module is used to determine the weight of each outlet for multiple handover time periods corresponding to various key date types based on the historical passenger flow of multiple outlets. The demand acquisition module is used to acquire the current cash box transportation demand from multiple outlets; The transportation management module is used to generate the optimal cash box escort plan based on the current cash box transportation needs of multiple outlets, multiple outlet groups, and the weights of multiple handover time periods corresponding to various key date types for each outlet. The optimal cash box escort plan includes multiple cash box escort routes, and each logistics transportation route corresponds to cash box transportation at least one outlet.
[0081] The logistics management system for bank cash boxes can be used to implement logistics management methods for bank cash boxes, which will not be elaborated here.
[0082] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. A logistics management method for bank cash boxes, characterized in that, include: Obtain historical cash box transportation needs and historical passenger flow data from multiple locations; Based on the historical cash box transportation needs of multiple outlets and the shortest path between any two outlets, the multiple outlets are divided into multiple outlet groups; Based on the historical passenger flow of multiple outlets, the weights of multiple handover time periods corresponding to various key date types are determined for each outlet; Obtain the current cash box transportation needs of multiple outlets; Based on the current cash box transportation needs of multiple outlets, the weights of multiple outlet groups and multiple handover time periods corresponding to various key date types for each outlet, an optimal cash box escort plan is generated. The optimal cash box escort plan includes multiple cash box escort routes, and each logistics transportation route corresponds to cash box transportation at least one outlet.
2. The logistics management method for bank cash boxes according to claim 1, characterized in that, Based on the historical cash box transportation needs of multiple network points and the shortest path between any two network points, the multiple network points are divided into multiple network point groups, including: Based on the historical cash box transportation needs of multiple outlets, calculate the demand similarity between any two outlets; Based on the demand similarity between any two network points, determine the similar and dissimilar network points for each network point; For each branch, based on the branch's historical cash box transportation demand and the historical cash box transportation demand of each dissimilar branch, calculate the demand correlation coefficient between the branch and each dissimilar branch. Using an adaptive center clustering algorithm, multiple network points are divided into multiple network point groups based on the similar and dissimilar network points of each network point, the demand correlation coefficient between each network point and each dissimilar network point, and the shortest path between any two network points.
3. The logistics management method for bank cash boxes according to claim 2, characterized in that, Using an adaptive center clustering algorithm, based on similar and dissimilar nodes for each node, the demand correlation coefficient between a node and each dissimilar node, and the shortest path between any two nodes, multiple nodes are divided into multiple node groups, including: S11. Based on the demand similarity between any two network points, calculate the demand similarity fluctuation value of each network point. Based on the demand correlation coefficient between the network point and each dissimilar network point, calculate the demand correlation fluctuation value of each network point. Based on the number of dissimilar network points, the demand similarity fluctuation value, and the demand correlation fluctuation value, calculate the adaptive center value of the network point. S12. Based on the adaptive center value of each point, determine multiple points as cluster centers; S13. For each non-cluster center point, if the cluster center is a dissimilar point to the point, then the cluster center is used as a candidate cluster center. Based on the demand correlation coefficient and shortest path between the point and the candidate cluster center, the cluster distance between the point and the candidate cluster center is calculated, and the point is assigned to the candidate cluster center with the shortest cluster distance. S14. Determine whether the inner loop clustering condition is met. If yes, execute S16; otherwise, execute S15. S15. Update the cluster center, then execute S13; S16. Determine whether the outer loop clustering conditions are met. If yes, complete the clustering and divide the multiple nodes into multiple node groups. If no, execute S17. S17. Determine the multi-objective clustering function, determine the update direction, update the number of cluster centers based on the update direction, and execute S12.
4. The logistics management method for bank cash boxes according to claim 2, characterized in that, Based on historical passenger traffic at multiple locations, the weights of various handover time periods corresponding to different key date types are determined for each location, including: Determine multiple date types; Based on the historical passenger flow of multiple outlets, calculate the passenger flow similarity between any two date types, merge multiple date types, and determine multiple key date types; For each key date type, based on the historical passenger flow of multiple outlets, calculate the initial weights of multiple handover time periods corresponding to the key date type for each outlet; Based on the demand correlation coefficients between similar sites and dissimilar sites for each site, the initial weights of multiple handover time periods corresponding to key date types for each site are iteratively corrected to generate weights for multiple handover time periods corresponding to multiple key date types for each site.
5. The logistics management method for bank cash boxes according to claim 4, characterized in that, Based on historical passenger flow data from multiple locations, the similarity of passenger flow between any two date types is calculated. Multiple date types are then merged to identify several key date types, including: For each branch, based on the branch's historical passenger flow, calculate the passenger flow similarity for any two date types corresponding to the branch; Based on the similarity of passenger flow for any two date types at each branch, multiple date types are merged to determine several key date types.
6. The logistics management method for bank cash boxes according to claim 5, characterized in that, Based on the demand correlation coefficients between similar and dissimilar locations for each location, the initial weights of multiple handover time periods corresponding to key date types for each location are iteratively corrected to generate weights for multiple handover time periods corresponding to various key date types for each location, including: S21. For each network point, based on the demand correlation coefficient between the network point and each dissimilar network point, determine the positively correlated network points and negatively correlated network points of the network point. Based on the initial weights of the handover time periods corresponding to the network point's similar network points, positively correlated network points and negatively correlated network points, correct the initial weights of the handover time periods corresponding to the network point, and generate the corrected weights of the network point's handover time periods. S22. Calculate the global corrected distance based on the initial weight and the corrected weight of the handover time period corresponding to each network point; S23. Based on the global correction distance, determine whether the iteration termination condition is met. If yes, generate the weights of multiple handover time periods corresponding to various key date types for each network point. If no, execute S24. S24. Use the corrected weight of the handover time period corresponding to the branch as the initial weight of the handover time period corresponding to the branch, and execute S21.
7. The logistics management method for bank cash boxes according to any one of claims 1-6, characterized in that, Based on the current cash container transportation needs of multiple network points, the weights of multiple network point groups and multiple handover time periods corresponding to various key date types for each network point, an optimal logistics transportation plan is generated, including: The problem is modeled as a reinforcement learning task, where the state includes the current cash box transportation demand of the remaining network points, multiple network point groups, the current location of the vehicle, the current key date type, the weight of multiple handover time periods for each network point corresponding to the current key date type, and the shortest travel time of the road connecting any two network points in multiple handover time periods, and the action is the next network point to be visited. Build and train a deep Q-network model; Based on the current cash and container transportation needs of multiple network points and multiple network point groups, multiple initial network points are generated. Based on the multiple initial network points, the optimal logistics transportation plan is generated using a deep Q-network model.
8. The logistics management method for bank cash boxes according to claim 7, characterized in that, Based on the current cash container transportation needs of multiple network points and multiple network point groups, multiple initial network points are generated. Then, using a deep Q-network model, an optimal logistics transportation plan is generated based on these initial network points, including: S31. Based on the current cash box transportation needs of multiple remaining outlets, determine the total remaining needs of each outlet group. Based on the total remaining needs of each outlet group, determine the current outlet group. Based on the current cash box transportation needs of each remaining outlet included in the current outlet group, determine the current initial outlet. S32. The deep Q-network model outputs the Q-value of each action based on the current state and selects the action with the highest Q-value. S33. Update the status and determine whether the single-path termination condition is met. If yes, execute S34; otherwise, execute S32. S34. Determine whether the termination condition for generating the solution is met. If yes, generate the optimal logistics transportation solution. If not, execute S31.
9. The logistics management method for bank cash boxes according to claim 7, characterized in that, The reward function of the deep Q-network model is related to at least whether the next node and the current node belong to the same node group, the shortest travel time from the current node to the next node, and the handover time period of the next node.
10. A logistics management system for bank cash boxes, characterized in that, The logistics management method for bank cash boxes according to any one of claims 1-9 includes: The data acquisition module is used to acquire historical cash box transportation needs and historical passenger flow from multiple outlets; The branch grouping module is used to divide multiple branches into multiple branch groups based on the historical cash box transportation needs of multiple branches and the shortest path between any two branches; The weight determination module is used to determine the weight of each outlet for multiple handover time periods corresponding to various key date types based on the historical passenger flow of multiple outlets. The demand acquisition module is used to acquire the current cash box transportation demand from multiple outlets; The transportation management module is used to generate an optimal cash box escort plan based on the current cash box transportation needs of multiple outlets, multiple outlet groups, and the weights of multiple handover time periods corresponding to various key date types for each outlet. The optimal cash box escort plan includes multiple cash box escort routes, and each logistics transportation route corresponds to cash box transportation at least one outlet.
Citation Information
Patent Citations
Distribution path optimization method and device
CN112785085A
Smart bank outlet scheduling management method and system and storage medium
CN116151950A
Ecort route planning method and device, electronic equipment and storage medium
CN117010786A
Armor cash carrier path planning method, and training method and device of path prediction model
CN117035207A