Logistics management method and system for banknote cases

By using an adaptive center clustering algorithm and a deep Q-network model, network points are dynamically grouped to generate the optimal escort plan, which solves the problem of low efficiency in the transportation management of bank cash boxes and achieves efficient and flexible logistics management.

CN120952653BActive Publication Date: 2025-12-23SICHUAN JINTOU FINANCIAL ECONOMIC SERVICE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511488265.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-12-23
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing technologies for managing bank cash box transportation suffer from inefficiencies due to static route planning, unreasonable resource allocation, high empty-running rates, fuel waste, and safety risks, making them difficult to adapt to complex and ever-changing transportation needs.

Method used

By using an adaptive center clustering algorithm and a deep Q-network model, network points are dynamically grouped to generate the optimal cash box escort plan. Combined with historical transportation demand and passenger flow, resource allocation and route planning are optimized.

Benefits of technology

It has improved transportation efficiency, reduced transportation costs, mitigated the risk of route congestion, enabled greater flexibility and accuracy in logistics management, and optimized resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952653B_ABST
    Figure CN120952653B_ABST
Patent Text Reader

Abstract

The application provides a logistics management method and system for bank cash boxes, and relates to the field of logistics management.The method comprises the following steps: obtaining historical cash box transportation demands and historical passenger flows of multiple sites; dividing the multiple sites into multiple site groups based on the historical cash box transportation demands of the multiple sites and the shortest paths between any two sites; determining the weights of multiple handover time periods corresponding to multiple key date types of each site based on the historical passenger flows of the multiple sites; obtaining current cash box transportation demands of the multiple sites; and generating an optimal cash box escort scheme based on the current cash box transportation demands of the multiple sites, the multiple site groups and the weights of the multiple handover time periods corresponding to multiple key date types of each site, wherein the optimal cash box escort scheme comprises multiple cash box escort paths, each logistics transportation path corresponds to cash box transportation of at least one site, and has the advantage of improving the transportation efficiency of bank cash boxes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of logistics management, in particular to a logistics management method and system for bank money boxes. BACKGROUND

[0002] A bank money box is a special sealed container used by banks in the process of cash and important goods circulation, mainly used for safe storage and transportation of cash, bills, precious metals, valuable securities and other high-value goods. It is a key tool in the bank's fund management chain, directly related to the safety of funds and operational efficiency.

[0003] The transportation management of bank money boxes is a key link to ensure the safety of funds and business continuity. However, the static path planning method currently widely used in the industry has been difficult to adapt to the complex and changing real needs, and its technical defects have become a key bottleneck restricting efficiency and safety. In the traditional mode, the transportation route is usually pre-set based on historical experience or fixed administrative divisions, such as the circular mode of "branch A → vault → branch B", and the departure schedule is strictly fixed. This design leads to vehicles frequently getting stuck in congested sections, with a single trip in-transit time fluctuating in the range of ±50%; at the same time, the static planning of resource allocation is inefficient, with an empty running rate of more than 30%, highlighting the problems of fuel waste and rising labor costs.

[0004] Therefore, it is necessary to provide a logistics management method and system for bank money boxes to improve the efficiency of bank money box transportation. SUMMARY

[0005] The present application provides a logistics management method for bank money boxes, comprising: obtaining historical money box transportation demand and historical passenger flow of a plurality of branches; dividing the plurality of branches into a plurality of branch groups based on the historical money box transportation demand of the plurality of branches and the shortest path between any two branches; determining the weight of each branch corresponding to a plurality of transfer time periods of a plurality of key date types based on the historical passenger flow of the plurality of branches; obtaining the current money box transportation demand of the plurality of branches; generating an optimal money box escort scheme based on the current money box transportation demand of the plurality of branches, the plurality of branch groups and the weight of each branch corresponding to a plurality of transfer time periods of a plurality of key date types, wherein the optimal money box escort scheme comprises a plurality of money box escort paths, and each logistics transportation path corresponds to the money box transportation of at least one branch.

[0006] Further, the plurality of network points are divided into a plurality of network point groups based on historical case box transportation demands of the plurality of network points and shortest paths between any two network points, including: calculating demand similarity between any two network points based on historical case box transportation demands of the plurality of network points; determining similar network points and dissimilar network points of each network point based on the demand similarity between any two network points; for each network point, calculating demand correlation coefficients between the network point and each dissimilar network point based on historical case box transportation demands of the network point and historical case box transportation demands of each dissimilar network point; and dividing the plurality of network points into the plurality of network point groups based on the similar network points and the dissimilar network points of each network point, the demand correlation coefficients between the network point and each dissimilar network point, and the shortest paths between any two network points through a self-adaptive center clustering algorithm.

[0007] Further, the plurality of network points are divided into a plurality of network point groups based on the similar network points and the dissimilar network points of each network point, the demand correlation coefficients between the network point and each dissimilar network point, and the shortest paths between any two network points through a self-adaptive center clustering algorithm, including: S11, calculating a demand similarity fluctuation value of each network point based on the demand similarity between any two network points, calculating a demand correlation fluctuation value of each network point based on the demand correlation coefficients between the network point and each dissimilar network point, and calculating a self-adaptive center value of the network point based on a number of dissimilar network points of the network point, the demand similarity fluctuation value, and the demand correlation fluctuation value; S12, determining the plurality of network points as cluster centers based on the self-adaptive center values of each network point; S13, for each network point that is not a cluster center, if a cluster center is a dissimilar network point of the network point, taking the cluster center as a candidate cluster center, calculating a clustering distance between the network point and the candidate cluster center based on the demand correlation coefficients between the network point and the candidate cluster center and the shortest paths, and assigning the network point to the candidate cluster center with the shortest clustering distance; S14, judging whether an inner loop clustering condition is met, if yes, performing S16, if not, performing S15; S15, updating the cluster centers and performing S13; S16, judging whether an outer loop clustering condition is met, if yes, completing clustering and dividing the plurality of network points into the plurality of network point groups, if not, performing S17; and S17, determining a multi-objective clustering function, determining an update direction, updating a number of cluster centers based on the update direction, and performing S12.

[0008] Further, based on the historical passenger flow of the plurality of network points, weights of a plurality of handover time periods corresponding to a plurality of key date types of each network point are determined, including: determining a plurality of date types; based on the historical passenger flow of the plurality of network points, calculating the passenger flow similarity of any two date types, merging the plurality of date types, and determining a plurality of key date types; for each key date type, based on the historical passenger flow of the plurality of network points, calculating the initial weight of the plurality of handover time periods corresponding to the key date type of each network point; based on the similar network points of each network point and the demand correlation coefficient of the network point and each dissimilar network point, the initial weight of the plurality of handover time periods corresponding to the key date type of each network point is iteratively corrected to generate the weight of the plurality of handover time periods corresponding to the plurality of key date types of each network point.

[0009] Further, based on the historical passenger flow of the plurality of network points, the passenger flow similarity of any two date types is calculated, the plurality of date types is merged, and a plurality of key date types is determined, including: for each network point, based on the historical passenger flow of the network point, the passenger flow similarity of the network point corresponding to any two date types is calculated; based on the passenger flow similarity of any two date types corresponding to each network point, the plurality of date types is merged to determine a plurality of key date types.

[0010] Further, based on the similar network points of each network point and the demand correlation coefficient of the network point and each dissimilar network point, the initial weight of the plurality of handover time periods corresponding to the key date type of each network point is iteratively corrected to generate the weight of the plurality of handover time periods corresponding to the plurality of key date types of each network point, including: S21, for each network point, based on the demand correlation coefficient of the network point and each dissimilar network point, the positively correlated network point and the negatively correlated network point of the network point are determined, and based on the initial weight of the handover time period corresponding to the similar network point, the positively correlated network point and the negatively correlated network point of the network point, the initial weight of the handover time period corresponding to the network point is corrected to generate the corrected weight of the handover time period corresponding to the network point; S22, according to the initial weight and the corrected weight of the handover time period corresponding to each network point, the global correction distance is calculated; S23, according to the global correction distance, it is judged whether the iteration end condition is met, if yes, the weight of the plurality of handover time periods corresponding to the plurality of key date types of each network point is generated, if not, S24 is executed; S24, the corrected weight of the handover time period corresponding to the network point is taken as the initial weight of the handover time period corresponding to the network point, and S21 is executed.

[0011] Further, based on the current cash box transportation demand of the plurality of network points, the plurality of network point groups, and the weight of the plurality of transfer time periods corresponding to each network point of the plurality of key date types, an optimal logistics transportation scheme is generated, including: modeling the problem as a reinforcement learning task, wherein the state includes the current cash box transportation demand of the remaining network points, the plurality of network point groups, the current position of the vehicle, the current key date type, the weight of the plurality of transfer time periods corresponding to each network point of the current key date type, and the shortest travel time of the road connecting any two network points in the plurality of transfer time periods, and the action is the next visited network point; a deep Q network model is established and trained; based on the current cash box transportation demand of the plurality of network points and the plurality of network point groups, a plurality of initial network points are generated, and based on the plurality of initial network points, the deep Q network model is used to generate the optimal logistics transportation scheme.

[0012] Further, based on the current cash box transportation demand of the plurality of network points and the plurality of network point groups, a plurality of initial network points are generated, and based on the plurality of initial network points, the deep Q network model is used to generate the optimal logistics transportation scheme, including: S31, based on the current cash box transportation demand of the plurality of remaining network points, determining the remaining total demand of each network point group, determining the current network point group according to the remaining total demand of each network point group, and determining the current initial network point according to the current cash box transportation demand of each remaining network point included in the current network point group; S32, the deep Q network model outputs the Q value of each action according to the current state, and selects the action with the highest Q value; S33, update the state, and judge whether the single path termination condition is met, if yes, execute S34, if not, execute S32; S34, judge whether the scheme generation termination condition is met, if yes, generate the optimal logistics transportation scheme, if not, execute S31.

[0013] Further, the reward function of the deep Q network model is related to at least whether the next network point and the current network point belong to the same network point group, the shortest travel time from the current network point to the next network point, and the transfer time period of the next network point.

[0014] The application provides a logistics management system for bank cash boxes for implementing the logistics management method for bank cash boxes, comprising: a data acquisition module for acquiring historical cash box transportation demands and historical customer flows of multiple sites; a site grouping module for grouping the multiple sites into multiple site groups based on the historical cash box transportation demands of the multiple sites and the shortest paths between any two sites; a weight determination module for determining the weights of multiple handover time periods corresponding to multiple key date types of each site based on the historical customer flows of the multiple sites; a demand acquisition module for acquiring current cash box transportation demands of the multiple sites; and a transportation management module for generating an optimal cash box transportation plan based on the current cash box transportation demands of the multiple sites, the multiple site groups and the weights of the multiple handover time periods corresponding to multiple key date types of each site, wherein the optimal cash box transportation plan comprises multiple cash box transportation paths, and each logistics transportation path corresponds to the cash box transportation of at least one site.

[0015] Compared with the prior art, the logistics management method and system for bank cash boxes provided by the application have at least the following beneficial effects:

[0016] 1. The site grouping is performed based on the historical cash box transportation demands and the shortest paths between sites, so that the sites in the group are more relevant in terms of transportation demand and geographical location, the cross-group transportation cost is reduced, and the overall logistics efficiency is improved. The weights of the handover time periods of each site under different key date types are determined according to the historical customer flow, which can accurately reflect the business busy degree of each time period, so that the generated transportation plan is more suitable for actual business demand and safety demand, and the resource allocation is optimized. The current cash box transportation demand is acquired, and the optimal transportation plan is generated in combination with the grouping and weight information, so that the transportation demand change of the site can be responded in real time, the path planning can be dynamically adjusted, the effectiveness and timeliness of the plan can be ensured, and the bank cash box logistics management level is improved.

[0017] 2. The multi-objective clustering function realizes the dispersed grouping of demand patterns through the dual constraints of demand similarity / correlation and cluster number control, so as to reduce the risk caused by the concentrated outbreak of demand in the cluster. By minimizing the average demand similarity of the sites in the cluster to the cluster center, the demand patterns of the sites in the cluster are forced to be significantly different. In the cluster with dispersed demand patterns, the peak time of the demand of the sites is staggered, so that the access order can be dynamically adjusted during scheduling to avoid path congestion caused by concentrated access. In the traditional clustering, vehicles may be insufficient due to the simultaneous outbreak of demand in the cluster, and backup vehicles need to be frequently called; after the demand is dispersed, such situations are significantly reduced.

[0018] 3. The cash box escort problem is modeled as a reinforcement learning task. The state design comprehensively covers key information such as remaining demand, network group, vehicle location, date type, handover time period weight, and road travel time. This allows for dynamic adaptation to changes in transportation demand under different scenarios, improving the flexibility and accuracy of solution generation. By learning the optimal action (next network point visit) selection strategy through a deep Q-network model and combining it with road travel time predictions based on historical data, the escort route can be effectively optimized, reducing transportation time and costs and improving logistics efficiency. Initial network points are determined based on the total remaining demand of the network group, and then the deep Q-network model generates paths layer by layer. This satisfies both overall transportation demand and the service requirements of individual network points, ensuring comprehensive coverage and reasonable feasibility of the solution. The deep Q-network model automatically outputs the Q-value and selects the optimal action, iteratively updating the state until the termination condition is met, achieving automation and efficiency in solution generation, reducing manual intervention, and minimizing decision-making errors. Attached Figure Description

[0019] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0020] Figure 1 This is a flowchart illustrating a logistics management method for bank cash boxes according to some embodiments of this specification;

[0021] Figure 2 This is a flowchart illustrating the process of dividing multiple outlets into multiple outlet groups according to some embodiments of this specification;

[0022] Figure 3 This is a schematic diagram of a logistics management system for bank cash boxes, shown according to some embodiments of this specification. Detailed Implementation

[0023] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0024] Figure 1 This is a flowchart illustrating a logistics management method for bank cash boxes according to some embodiments of this specification, such as... Figure 1 As shown, the logistics management method for bank cash boxes may include the following steps.

[0025] Step 1, obtaining historical case transportation demands and historical customer flows of multiple sites.

[0026] Specifically, the historical case transportation demands of a site can include the case transportation demand amounts of multiple periods (e.g., every day, every week, etc.) in the past, as shown in Table 1.

[0027] Netpoint ID Date Demand Date Type Handover Time Period Netpoint 1 2023-10-01 8 Weekday 06:00-07:00 Netpoint 2 2023-10-01 3 Weekday 12:00-13:00

[0028] The historical customer flows of a site can include the numbers of visiting customers in multiple historical time periods, as shown in Table 2.

[0029] Table 2

[0030] Step 2, dividing the multiple sites into multiple site groups based on the historical case transportation demands of the multiple sites and the shortest paths between any two sites.

[0031] Specifically, it includes:

[0032] calculating the demand similarity between any two sites based on the historical case transportation demands of the multiple sites;

[0033] determining the similar sites and the dissimilar sites of each site based on the demand similarity between any two sites;

[0034] for each site, calculating the demand correlation coefficient between the site and each dissimilar site based on the historical case transportation demand of the site and the historical case transportation demand of each dissimilar site;

[0035] dividing the multiple sites into multiple site groups based on the similar sites and the dissimilar sites of each site, the demand correlation coefficient between the site and each dissimilar site, and the shortest paths between any two sites through a self-adaptive clustering algorithm.

[0036] The demand similarity between two sites is used to measure the closeness of the demand patterns of the two sites in the historical case transportation demands. The cosine similarity of the historical case transportation demands of the two sites can be calculated as the demand similarity between the two sites.

[0037] For each site, when there is a site with a demand similarity greater than a demand similarity threshold (e.g., 0.6, etc.) with the site, the site is regarded as a similar site. The demand similarity threshold can be determined through experimental data.

[0038] For each site, the historical case transportation demand of the site and the historical case transportation demand of the dissimilar site can be substituted into a correlation coefficient calculation formula (e.g., Pearson correlation coefficient, Kendall rank correlation coefficient, etc.) to calculate the demand correlation coefficient between the site and the dissimilar site.

[0039] Figure 2 is a flowchart of dividing a plurality of nodes into a plurality of node groups according to some embodiments of the present specification, as shown in Figure 2 As shown, as preferred, a plurality of nodes are divided into a plurality of node groups by a clustering algorithm of an adaptive center based on similar nodes and dissimilar nodes of each node, a demand correlation coefficient of the node and each dissimilar node, and a shortest path between any two nodes, including:

[0040] S11, based on the demand similarity of any two nodes, calculating a demand similarity fluctuation value of each node, based on the demand correlation coefficient of the node and each dissimilar node, calculating a demand correlation fluctuation value of each node, based on the number of dissimilar nodes of the node, the demand similarity fluctuation value and the demand correlation fluctuation value, calculating an adaptive center value of the node;

[0041] S12, based on the adaptive center value of each node, determining a plurality of nodes as cluster centers;

[0042] S13, for each non-cluster center node, if the cluster center is a dissimilar node of the node, the cluster center is taken as a candidate cluster center, the clustering distance of the node and the candidate cluster center is calculated according to the demand correlation coefficient and the shortest path of the node and the candidate cluster center, and the node is assigned to the candidate cluster center with the shortest clustering distance;

[0043] S14, judging whether the inner loop clustering condition is met, if yes, executing S16, if no, executing S15;

[0044] S15, updating the cluster center and executing S13;

[0045] S16, judging whether the outer loop clustering condition is met, if yes, completing clustering and dividing a plurality of nodes into a plurality of node groups, if no, executing S17;

[0046] S17, determining a multi-objective clustering function, determining an update direction, updating the number of cluster centers based on the update direction, and executing S12.

[0047] Wherein, for each node, the standard deviation of the demand similarity of the node and any other node is calculated as the demand similarity fluctuation value of the node. The standard deviation of the demand correlation coefficient of the node and each dissimilar node is calculated as the demand correlation fluctuation value of the node.

[0048] The adaptive center value of the node can be calculated based on the following formula:

[0049] Wherein, is the adaptive center value of the i-th node, is the demand similarity fluctuation value of the i-th node, a demand correlation fluctuation value of the i th node, a number of dissimilar nodes of the i th node, a total number of nodes.

[0050] According to the adaptive center value of each node, the plurality of nodes are sorted, and the first K nodes are selected as cluster centers, wherein the value of K can be determined by the elbow rule, the silhouette coefficient, etc.

[0051] The smaller the demand correlation coefficient of the node and the candidate cluster center, the shorter the shortest path, and the shorter the clustering distance of the node and the candidate cluster center.

[0052] The inner loop clustering condition can be that the number of inner loop times reaches the maximum number of loop times or the number of updated cluster centers in multiple loops is less than a number threshold (for example, 3, etc.).

[0053] For each node included in each cluster, the standard deviation of the demand similarity between the node and any other node included in the cluster is calculated as the intra-cluster demand similarity fluctuation value of the node. The standard deviation of the demand correlation coefficient between the node and each dissimilar node included in the cluster is calculated as the intra-cluster demand correlation fluctuation value of the node. Based on the intra-cluster demand similarity fluctuation value of the node, the intra-cluster demand correlation fluctuation value of the node, and the number of intra-cluster dissimilar nodes of the node, the intra-cluster adaptive center value of the node is calculated, wherein the calculation method of the intra-cluster adaptive center value of the node and the adaptive center value of the node is similar, which will not be repeated here. The node with the largest intra-cluster adaptive center value is updated as the new cluster center.

[0054] The outer loop clustering condition can be that the number of outer loop times reaches the maximum number of loop times, the value of the target clustering function is greater than a preset threshold, or the value of the target clustering function reaches a maximum value.

[0055] The multi-objective clustering function is used to measure the quality of the clustering result, by quantifying multiple conflicting clustering objectives (for example, intra-cluster compactness, inter-cluster separation, path cost, etc.), to provide an optimization direction and termination condition for the clustering process. For example only, the multi-objective clustering function can include multiple parts shown in Table 3.

[0056] Objective Explanation Optimization Direction Significance Intra-cluster demand similarity objective Average demand similarity between intra-cluster netpoints and cluster center Minimize Ensure low degree of similarity in demand patterns among intra-cluster netpoints Inter-cluster demand similarity objective Average demand similarity between different cluster centers Maximize Ensure distribution of netpoints with similar demand patterns in different clusters Intra-cluster demand correlation objective Average demand correlation coefficient between intra-cluster netpoints and cluster center Minimize Ensure low degree of correlation in demand patterns among intra-cluster netpoints Inter-cluster demand correlation objective Average demand correlation coefficient between different cluster centers Maximize Ensure distribution of netpoints with similar demand patterns in different clusters Cluster number objective Non-linear penalty on cluster number Minimize Prevent over-clustering

[0057] The multi-objective clustering function realizes the dispersion grouping of demand patterns through the dual constraints of demand similarity / correlation and cluster number control, thereby reducing the risk caused by the concentrated outbreak of demand in the cluster in the transportation scheduling. By minimizing the average demand similarity of the nodes in the cluster to the cluster center, the demand patterns of the nodes in the cluster are forced to be significantly different (such as one node has high daytime demand and another node has high nighttime demand), avoiding the aggregation of “homogeneous” nodes. By maximizing the average demand similarity between different cluster centers, it is ensured that the nodes with similar demand patterns are dispersed to different clusters (such as all “high-peak-morning” nodes are evenly distributed to multiple clusters). Correlation measures the synchronicity of demand changes (such as whether node B demand rises when node A demand rises). By constraining correlation, further dispersion of nodes with synchronized demand fluctuations is achieved. The cluster number is limited by the cluster number target to avoid excessive grouping. If the demand patterns of the nodes in the cluster are significantly different (such as early peak, late peak, and stable mixed), there is no need to simultaneously cope with the peak demand of all nodes in a single transportation, and the vehicle load is more balanced. In the cluster with dispersed demand patterns, the peak demand time of the nodes is staggered, and the access order can be dynamically adjusted during scheduling to avoid path congestion caused by concentrated access. Traditional clustering may cause vehicle shortage due to simultaneous demand outbreak in the cluster, and may need to frequently call backup vehicles; after demand dispersion, such situations are significantly reduced.

[0058] Step 3, based on the historical passenger flow of multiple nodes, determine the weight of each node corresponding to multiple transfer time periods of multiple key date types.

[0059] As preferred, step 3 specifically includes:

[0060] Determine multiple date types, for example, weekdays (Monday to Friday), weekends (Saturday and Sunday), statutory holidays (such as Spring Festival and National Day), special dates (such as Double Eleven and school winter and summer vacation), etc.;

[0061] Based on the historical passenger flow of multiple nodes, calculate the passenger flow similarity of any two date types, merge multiple date types, and determine multiple key date types;

[0062] For each key date type, based on the historical passenger flow of multiple nodes, calculate the initial weight of each node corresponding to multiple transfer time periods of the key date type, wherein the transfer time period can be the time period during which the node can perform case box transfer;

[0063] Based on the similar nodes of each node and the demand correlation coefficient between the node and each non-similar node, the initial weight of each node corresponding to multiple transfer time periods of the key date type is iteratively corrected to generate the weight of each node corresponding to multiple transfer time periods of multiple key date types.

[0064] As preferred, based on the historical passenger flow of multiple network points, the passenger flow similarity of any two date types is calculated, the multiple date types are merged, and the multiple key date types are determined, including:

[0065] For each network point, based on the historical passenger flow of the network point, the passenger flow similarity of the network point corresponding to any two date types is calculated;

[0066] Based on the passenger flow similarity of each network point corresponding to any two date types, the multiple date types are merged, and the multiple key date types are determined.

[0067] Among them, for each network point, the cosine similarity of the average sequence of passenger flow of the network point in each time period of the whole day of the two date types is calculated as the passenger flow similarity of the network point corresponding to the two date types.

[0068] The average of the passenger flow similarity of each network point corresponding to two date types is taken as the global passenger flow similarity of the two date types. The two date types with a global passenger flow similarity greater than a global passenger flow similarity threshold (for example, 0.6) are merged into one, and after the merging is completed, the remaining date types are taken as key date types.

[0069] For each key date type, the difference between the average of the total historical passenger flow of a day and the average of the historical passenger flow of a certain handover time period and the ratio of the average of the total historical passenger flow of a day are taken as the initial weight of the handover time period. If the initial weight of a certain handover time period is smaller, it means that the passenger flow of the handover time period is larger, which may be accompanied by problems such as personnel mixing, noisy environment, etc., increasing the risk of being stolen, robbed or missed by monitoring during the transfer of the case.

[0070] As preferred, based on the similar network points of each network point and the demand correlation coefficient of the network point and each dissimilar network point, the initial weight of each network point corresponding to multiple handover time periods of the key date type is iteratively corrected to generate the weight of each network point corresponding to multiple handover time periods of multiple key date types, including:

[0071] S21, for each network point, based on the demand correlation coefficient of the network point and each dissimilar network point, the positively correlated network point and the negatively correlated network point of the network point are determined, and the initial weight of the network point corresponding to the handover time period is corrected based on the initial weight of the handover time period corresponding to the similar network point, the positively correlated network point and the negatively correlated network point of the network point to generate the corrected weight of the network point corresponding to the handover time period;

[0072] S22, according to the initial weight and the corrected weight of each network point corresponding to the handover time period, the global corrected distance is calculated;

[0073] S23, judging whether the iteration end condition is met according to the global correction distance, if yes, generating the weight of each network point corresponding to multiple handover time periods of multiple key date types, if no, performing S24;

[0074] S24, taking the corrected weight of the network point corresponding to the handover time period as the initial weight of the network point corresponding to the handover time period, and performing S21.

[0075] The network point with a demand correlation coefficient greater than a positive demand correlation coefficient threshold (for example, 0.5) can be regarded as a positive correlation network point, and the network point with a demand correlation coefficient less than a negative demand correlation coefficient threshold (for example, -0.5) can be regarded as a negative correlation network point.

[0076] If the initial weight of a certain handover time period of the similar network point is high, it indicates that the handover time period is more likely to conform to the regional business rules, and therefore the weight of the corresponding handover time period of the target network point is increased. If the weight of a certain handover time period of the positive correlation network point is high, it indicates that the handover time period is more likely to conform to the regional business rules, and therefore the weight of the corresponding handover time period of the target network point is increased. If the initial weight of a certain handover time period of the negative correlation network point is high, it may reflect that the demand of the target network point in the handover time period is low, and therefore the weight of the corresponding handover time period of the target network point is reduced.

[0077] The initial weight of the network point corresponding to the handover time period can be corrected in any manner based on the initial weight of the similar network point, the positive correlation network point and the negative correlation network point of the network point corresponding to the handover time period, for example, a first correction coefficient is calculated based on the initial weight of the similar network point corresponding to the handover time period, a second correction coefficient is calculated based on the initial weight of the positive correlation network point corresponding to the handover time period, a third correction coefficient is calculated based on the initial weight of the negative correlation network point corresponding to the handover time period, and the initial weight of the network point corresponding to the handover time period is corrected by the first correction coefficient, the second correction coefficient and the third correction coefficient to generate the corrected weight of the network point corresponding to the handover time period. For example only, the first correction coefficient can be determined based on the initial weight of the similar network point corresponding to the handover time period according to a preset rule, wherein the preset rule can be:

[0078] wherein, is the first correction coefficient, The difference between the average value of the initial weight of the corresponding handover time period of the similar node of the node and the initial weight of the corresponding handover time period of the node. The determination manner of the second correction coefficient and the third correction coefficient is similar to that of the first correction coefficient, which will not be described here. The first correction coefficient, the second correction coefficient and the third correction coefficient can be summed as a global correction coefficient, the product of the global correction coefficient and the initial weight of the corresponding handover time period of the node is calculated as a correction amount, and the sum of the correction amount and the initial weight of the corresponding handover time period of the node is obtained as the corrected weight of the corresponding handover time period of the node.

[0079] For each node, the absolute value of the difference between the initial weight and the corrected weight of the corresponding handover time period of the node is calculated as the correction distance of the node, and the sum of the correction distances of all nodes is obtained as the global correction distance.

[0080] The iteration end condition can be that the global correction distance is less than an iteration end condition threshold (for example, 0.3), or the number of iterations reaches a maximum iteration number.

[0081] Step 3 can accurately identify date types with similar passenger flow patterns by calculating the passenger flow similarity of any two date types of each node and merging them, avoid redundant management caused by too fine date type division, and ensure that key date types can truly reflect passenger flow characteristics in different business scenarios, improving the pertinence of subsequent weight calculation and escort scheme generation. The merged key date types are more in line with actual business cycle changes, so that the generated escort scheme can better adapt to passenger flow fluctuations on different dates, for example, after merging similar weekday types, the safety escort strategy of weekdays can be uniformly planned, improving the applicability and stability of the scheme on different dates. The initial weight of the target node is iteratively corrected based on the initial weights of the similar nodes, the positively correlated nodes and the negatively correlated nodes, fully considering the business correlation and demand difference between nodes, which can eliminate the influence of single node data deviation on the weight, and make the generated weight more accurately reflect the actual safety risk of each handover time period.

[0082] Step 4, obtaining the current cash box transportation demand of the plurality of nodes.

[0083] Step 5, based on the current cash box transportation demand of the plurality of nodes, the plurality of node groups and the weight of the plurality of handover time periods corresponding to each node for a plurality of key date types, generating an optimal cash box escort scheme.

[0084] Among them, the optimal cash box escort scheme includes a plurality of cash box escort paths, and each logistics transportation path corresponds to at least one cash box transportation of a node.

[0085] As preferred, step 5 specifically includes:

[0086] Model the problem as a reinforcement learning task, where the state includes the current demand for cash box transportation of the remaining nodes, multiple node groups, the current location of the vehicle, the current key date type, the weight of each node corresponding to the multiple delivery time periods of the current key date type, and the shortest travel time of the road connecting any two nodes in multiple delivery time periods, the action is the next visited node, the remaining nodes are the nodes that have not been served, and the shortest travel time of the road connecting any two nodes in multiple delivery time periods can be predicted based on historical data.

[0087] Establish and train a deep Q network model.

[0088] Based on the current demand for cash box transportation of multiple nodes and multiple node groups, generate multiple initial nodes, and based on the deep Q network model, generate an optimal logistics transportation scheme based on the multiple initial nodes.

[0089] As a preferred embodiment, based on the current demand for cash box transportation of multiple nodes and multiple node groups, generate multiple initial nodes, and based on the deep Q network model, generate an optimal logistics transportation scheme based on the multiple initial nodes, comprising:

[0090] S31, based on the current demand for cash box transportation of the remaining nodes, determine the remaining total demand of each node group, and according to the remaining total demand of each node group, determine the current node group, and according to the current demand for cash box transportation of each remaining node included in the current node group, determine the current initial node;

[0091] S32, the deep Q network model outputs the Q value of each action according to the current state, and selects the action with the highest Q value;

[0092] S33, update the state, and judge whether the single path termination condition is met, if yes, execute S34, if not, execute S32;

[0093] S34, judge whether the scheme generation termination condition is met, if yes, generate the optimal logistics transportation scheme, if not, execute S31.

[0094] The remaining total demand of the node group is the sum of the current demand for cash box transportation of the remaining nodes included in the node group. The node group with the largest remaining total demand can be selected as the current node group, and the remaining node with the largest current demand for cash box transportation included in the current node group can be selected as the current initial node.

[0095] The legal action selected from the remaining nodes must satisfy:

[0096] 1. Vehicle capacity constraint, i.e. the number of cash boxes loaded by the vehicle does not exceed the upper limit;

[0097] 2. Time window constraint, i.e. the arrival node needs to be within its deliverable time period.

[0098] As preferred, the reward function of the deep Q network model is at least related to whether the next node and the current node belong to the same node group, the shortest travel time from the current node to the next node, and the handover time period of the next node.

[0099] For example, the reward function of the deep Q network model can be:

[0100] wherein, is the reward, is the current state, is the action, is the state after the action is performed, , , and is the weight, , , and is greater than 0, for example, , , and are respectively 0.2, 0.2, 0.3, 0.3, is the node group attribution reward, is the shortest travel time reward, the shorter the shortest travel time from the current node to the next node, the greater the shortest travel time reward, is the weight of the handover time period to which the time of reaching the next node belongs, is the basic reward, such as a fixed cost penalty or a fixed reward for completing the delivery, is the demand similarity of the current node and the next node.

[0101] The state update can include: adding the selected action (next node) to the current path, updating the remaining node set, vehicle remaining capacity, cumulative cost, and other state variables.

[0102] The single-path termination condition can include:

[0103] 1. The vehicle capacity is full;

[0104] 2. There are no more accessible remaining nodes;

[0105] 3. The maximum path length limit (such as the upper limit of the number of nodes or distance) is reached.

[0106] One of them is met, i.e. the single-path termination condition is met.

[0107] The solution generation termination condition can include: all nodes are assigned to a path, or the total demand of the remaining nodes is less than the total amount that can be transported at a time.

[0108] If the scheme generation termination condition is met, output the current set of all paths as a candidate scheme.

[0109] Step 5 models the cash box transportation problem as a reinforcement learning task. The state design comprehensively covers key information such as remaining demand, group of sites, vehicle location, date type, weight of transfer time period, and road travel time, which can dynamically adapt to changes in transportation demand in different scenarios, improving the flexibility and accuracy of scheme generation. Through the deep Q network model, the optimal action (next access site) selection strategy is learned, combined with the predicted road travel time based on historical data, which can effectively optimize the transportation path, reduce transportation time and cost, and improve logistics efficiency. Based on the total remaining demand of the group of sites, the initial site is determined, and then the path is generated layer by layer by the deep Q network model, which can meet the overall transportation demand and also consider the service requirements of individual sites, ensuring that the scheme is comprehensive, reasonable and feasible. Through the deep Q network model, the Q value is automatically output and the optimal action is selected, and the state is iteratively updated until the termination condition is met, realizing the automation and efficiency of scheme generation, reducing manual intervention and decision-making errors.

[0110] Figure 3 is a schematic diagram of a module of a logistics management system for bank cash boxes according to some embodiments of the present specification, as shown in Figure 3 The logistics management system for bank cash boxes can include a data acquisition module, a site grouping module, a weight determination module, a demand acquisition module, and a transportation management module.

[0111] The data acquisition module is configured to acquire historical cash box transportation demand and historical passenger flow of a plurality of sites.

[0112] The site grouping module is configured to group the plurality of sites into a plurality of site groups based on the historical cash box transportation demand of the plurality of sites and the shortest path between any two sites.

[0113] The weight determination module is configured to determine the weight of each site corresponding to a plurality of transfer time periods of a plurality of key date types based on the historical passenger flow of the plurality of sites.

[0114] The demand acquisition module is configured to acquire current cash box transportation demand of the plurality of sites.

[0115] The transportation management module is configured to generate an optimal cash box transportation scheme based on the current cash box transportation demand of the plurality of sites, the plurality of site groups, and the weight of each site corresponding to a plurality of transfer time periods of a plurality of key date types, wherein the optimal cash box transportation scheme includes a plurality of cash box transportation paths, and each logistics transportation path corresponds to cash box transportation of at least one site.

[0116] The logistics management system for bank cash boxes can be used to perform the logistics management method for bank cash boxes, which will not be described here.

[0117] Finally, it should be understood that the embodiments described herein are only given by way of example. Other variations might fall within the scope of the present description. Accordingly, the present description is not limited to that precisely as shown and described.

Claims

1. A logistics management method for a banknote cassette, characterized by, The method comprises the following steps: obtaining historical cash box transportation demand and historical passenger flow of multiple sites; dividing the multiple sites into multiple site groups based on the historical cash box transportation demand of the multiple sites and the shortest path between any two sites; determining the weight of each site corresponding to multiple transfer time periods of multiple key date types based on the historical passenger flow of the multiple sites; obtaining current cash box transportation demand of the multiple sites; generating an optimal cash box transport scheme based on the current cash box transportation demand of the multiple sites, the multiple site groups and the weight of each site corresponding to multiple transfer time periods of multiple key date types, wherein the optimal cash box transport scheme comprises multiple cash box transport paths, each of which corresponds to cash box transportation of at least one site, and specifically, generating multiple initial sites based on the current cash box transportation demand of the multiple sites and the multiple site groups, and generating an optimal logistics transportation scheme based on the multiple initial sites by a deep Q network model, wherein a reward function of the deep Q network model is related to at least whether the next site and the current site belong to the same site group, the shortest travel time from the current site to the next site and the transfer time period of the next site.

2. The logistics management method for a banknote case according to claim 1, characterized by, The method for dividing the multiple sites into multiple site groups based on the historical cash box transportation demand of the multiple sites and the shortest path between any two sites comprises the following steps: calculating the demand similarity between any two sites based on the historical cash box transportation demand of the multiple sites; determining the similar sites and the dissimilar sites of each site based on the demand similarity between any two sites; calculating the demand correlation coefficient between the site and each dissimilar site based on the historical cash box transportation demand of the site and the historical cash box transportation demand of each dissimilar site; dividing the multiple sites into multiple site groups based on the similar sites and the dissimilar sites of each site, the demand correlation coefficient between the site and each dissimilar site and the shortest path between any two sites by an adaptive center clustering algorithm.

3. The logistics management method for a banknote case according to claim 2, characterized by, The method for dividing the multiple sites into multiple site groups by an adaptive center clustering algorithm based on the similar sites and the dissimilar sites of each site, the demand correlation coefficient between the site and each dissimilar site and the shortest path between any two sites comprises the following steps: S11, calculating the demand similarity fluctuation value of each site based on the demand similarity between any two sites, calculating the demand correlation fluctuation value of each site based on the demand correlation coefficient between the site and each dissimilar site, and calculating the adaptive center value of the site based on the number of dissimilar sites of the site, the demand similarity fluctuation value and the demand correlation fluctuation value; S12, determining multiple sites as cluster centers based on the adaptive center value of each site; S13, for each non-cluster center site, if the cluster center is a dissimilar site of the site, the cluster center is taken as a candidate cluster center, the clustering distance between the site and the candidate cluster center is calculated according to the demand correlation coefficient and the shortest path between the site and the candidate cluster center, and the site is assigned to the candidate cluster center with the shortest clustering distance; S14, judging whether the inner loop clustering condition is met, if yes, executing S16, and if no, executing S15; S15, updating the cluster center and executing S13; S16, judging whether the outer loop clustering condition is met, if yes, completing clustering, dividing the plurality of sites into a plurality of site groups, if no, executing S17; S17, determining a multi-objective clustering function, determining an update direction, updating the number of cluster centers based on the update direction, and executing S12.

4. The logistics management method for a banknote case according to claim 2, characterized by, Based on the historical passenger flow of the plurality of sites, the weight of each site corresponding to a plurality of transfer time periods of a plurality of key date types is determined, including: determining a plurality of date types; Based on the historical passenger flow of the plurality of sites, the passenger flow similarity of any two date types is calculated, the plurality of date types is merged, and a plurality of key date types is determined. For each key date type, based on the historical passenger flow of the plurality of sites, the initial weight of each site corresponding to a plurality of transfer time periods of the key date type is calculated. Based on the similar sites of each site and the demand correlation coefficient of the site and each dissimilar site, the initial weight of each site corresponding to a plurality of transfer time periods of the key date type is iteratively corrected to generate the weight of each site corresponding to a plurality of transfer time periods of a plurality of key date types.

5. The logistics management method for a banknote cassette according to claim 4, characterized by, Based on the historical passenger flow of the plurality of sites, the passenger flow similarity of any two date types is calculated, the plurality of date types is merged, and a plurality of key date types is determined, including: For each site, based on the historical passenger flow of the site, the passenger flow similarity of the site corresponding to any two date types is calculated. Based on the passenger flow similarity of each site corresponding to any two date types, the plurality of date types is merged to determine a plurality of key date types.

6. The logistics management method for a banknote case according to claim 5, characterized by, Based on the similar sites of each site and the demand correlation coefficient of the site and each dissimilar site, the initial weight of each site corresponding to a plurality of transfer time periods of the key date type is iteratively corrected to generate the weight of each site corresponding to a plurality of transfer time periods of a plurality of key date types, including: S21, for each site, based on the demand correlation coefficient of the site and each dissimilar site, the positively correlated sites and the negatively correlated sites of the site are determined, and based on the initial weight of the similar sites, the positively correlated sites and the negatively correlated sites of the site corresponding to the transfer time period, the initial weight of the site corresponding to the transfer time period is corrected to generate the corrected weight of the site corresponding to the transfer time period; S22, according to the initial weight and the corrected weight of each site corresponding to the transfer time period, the global correction distance is calculated; S23, according to the global correction distance, whether the iteration end condition is met is judged, if yes, the weight of each site corresponding to a plurality of transfer time periods of a plurality of key date types is generated, if no, S24 is executed; S24, the corrected weight of the site corresponding to the transfer time period is taken as the initial weight of the site corresponding to the transfer time period, and S21 is executed.

7. The logistics management method for a banknote case according to any one of claims 1 to 6, characterized by, Based on the current case box transportation demand of the plurality of sites, the plurality of site groups and the weight of each site corresponding to a plurality of transfer time periods of a plurality of key date types, an optimal logistics transportation scheme is generated, including: Model the problem as a reinforcement learning task, where the state includes the current cash box transportation demand of the remaining network points, multiple network point groups, the current location of the vehicle, the current key date type, the weight of each network point corresponding to multiple delivery time periods of the current key date type, and the shortest travel time of the road connecting any two network points in multiple delivery time periods, and the action is the next visited network point; Establish and train a deep Q network model; Based on the current cash box transportation demand of multiple network points and multiple network point groups, generate multiple initial network points, and based on the deep Q network model, generate an optimal logistics transportation scheme based on the multiple initial network points.

8. The logistics management method for a banknote case according to claim 7, characterized by, Based on the current cash box transportation demand of multiple network points and multiple network point groups, generate multiple initial network points, and based on the deep Q network model, generate an optimal logistics transportation scheme based on the multiple initial network points, including: S31, based on the current cash box transportation demand of the remaining network points, determine the remaining total demand of each network point group, and according to the remaining total demand of each network point group, determine the current network point group, and according to the current cash box transportation demand of each remaining network point included in the current network point group, determine the current initial network point; S32, the deep Q network model outputs the Q value of each action according to the current state, and selects the action with the highest Q value; S33, update the state, and judge whether the single path termination condition is met, if yes, execute S34, if not, execute S32; S34, judge whether the scheme generation termination condition is met, if yes, generate the optimal logistics transportation scheme, if not, execute S31.

9. Logistics management system for banknote cases, characterized in that The logistics management method for bank cash boxes of any one of claims 1-8, comprising: A data acquisition module for acquiring historical cash box transportation demand and historical passenger flow of multiple network points; A network point grouping module for grouping multiple network points into multiple network point groups based on the historical cash box transportation demand of multiple network points and the shortest path between any two network points; A weight determination module for determining the weight of each network point corresponding to multiple delivery time periods of multiple key date types based on the historical passenger flow of multiple network points; A demand acquisition module for acquiring the current cash box transportation demand of multiple network points; A transportation management module for generating an optimal cash box transportation scheme based on the current cash box transportation demand of multiple network points, multiple network point groups, and the weight of each network point corresponding to multiple delivery time periods of multiple key date types, wherein the optimal cash box transportation scheme includes multiple cash box transportation paths, and each logistics transportation path corresponds to the cash box transportation of at least one network point.

Citation Information

Patent Citations

  • Distribution path optimization method and device

    CN112785085A

  • Armor cash carrier path planning method, and training method and device of path prediction model

    CN117035207A