A Joint Optimization Method for Value Mining and Dynamic Scheduling in Logistics Data Space

CN122572976APending Publication Date: 2026-08-14WUXI PROFESSIONAL COLLEGE OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610715732.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

这种串行架构存在两个核心缺陷:其一,挖掘模型仅以预测准确率为目标,未考虑其对后续调度收益的实际贡献,导致挖掘结果可能具备统计学意义但缺乏调度实用价值;其二,调度模型仅被动接收挖掘结果,无法主动引导挖掘方向,也无法对挖掘结果的置信度偏差进行鲁棒性应对,当挖掘结果出现较大偏差时,调度决策质量将严重下降

Benefits of technology

1、实现挖掘与调度深度协同:本发明通过构建调度收益导向的时空联邦挖掘模型,将挖掘结果的调度贡献直接融入损失函数,使挖掘过程不再以预测准确率为唯一目标,而是面向调度收益进行优化,解决了挖掘与调度流程脱节导致挖掘结果缺乏实用价值的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122572976A_ABST
    Figure CN122572976A_ABST
Patent Text Reader

Abstract

This invention discloses a joint optimization method for value mining and dynamic scheduling in the logistics data space, belonging to the field of intelligent logistics scheduling technology, and is used for cross-entity collaborative resource scheduling across the entire logistics chain. First, this invention constructs a scheduling benefit-oriented spatiotemporal federated mining model to achieve compliant cross-entity data use and high-value mining results output. Second, based on historical scheduling contributions, it calculates the confidence level of the mining results and filters invalid results, constructing a multi-agent scheduling model with robust confidence penalties to generate globally optimal scheduling decisions. Finally, through a dual-loop mechanism of minute-level online adjustment and day-level offline iteration, it achieves bidirectional joint optimization of mining and scheduling. This invention effectively solves problems such as the disconnect between mining and scheduling and data silos, and can significantly reduce logistics operating costs, improve delivery timeliness, and ensure data usage compliance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent logistics scheduling technology, specifically involving a joint optimization method for value mining and dynamic scheduling in logistics data space. Background Technology

[0002] With the rapid development of the logistics industry, logistics systems are transforming from single-entity, closed-loop operations to a multi-entity collaborative, data-driven open ecosystem. Against this backdrop, the logistics data space, as an infrastructure that aggregates operational data from different logistics entities, provides new possibilities for improving end-to-end scheduling efficiency by integrating heterogeneous data from multiple sources such as orders, transportation capacity, road conditions, and weather. How to efficiently utilize the data resources in the logistics data space and mine high-value information to guide dynamic scheduling decisions has become a research hotspot in the field of intelligent logistics scheduling.

[0003] However, existing logistics scheduling technologies still have the following shortcomings in terms of collaborative data mining and scheduling, and cross-entity data value utilization: First, the data mining and scheduling processes are disconnected, lacking value-oriented collaborative optimization. Existing technologies typically separate data mining and scheduling decisions into two independent stages: first, the mining model generates prediction results, and then the scheduling model generates decisions based on these static results. This sequential architecture has two core flaws: First, the mining model only aims at prediction accuracy, without considering its actual contribution to subsequent scheduling benefits, resulting in mining results that may have statistical significance but lack practical scheduling value; second, the scheduling model only passively receives mining results, unable to actively guide the mining direction or robustly handle confidence biases in the mining results. When the mining results show significant deviations, the quality of scheduling decisions will severely decline.

[0004] Second, cross-entity data fusion faces data silos and compliance bottlenecks. In the context of end-to-end logistics collaboration, key data such as orders, transportation capacity, and road conditions are often scattered across different entities. Due to concerns about trade secrets and data security, these entities cannot directly share the original data. If a centralized data fusion approach is adopted, it not only faces data compliance risks but also makes it difficult to mobilize the active participation of data resources from various entities.

[0005] Third, there is a lack of quantitative feedback and closed-loop optimization mechanisms for mining value. After the existing system deploys the mining and scheduling models, it lacks a continuous evaluation and feedback mechanism for the actual scheduling contribution of the mining results. On the one hand, it is impossible to quantify the marginal contribution of mining features of different dimensions to scheduling utility, resulting in the inability to accurately optimize the mining model. On the other hand, the scheduling model lacks dynamic adaptability to the confidence level of the mining results, and cannot automatically adjust its dependence on the mining results when the mining quality fluctuates, resulting in insufficient robustness.

[0006] In summary, there is an urgent need for a logistics scheduling method that can connect data mining and scheduling in a closed loop, achieve compliant cross-entity data collaboration, and possess value quantification feedback and robust optimization capabilities. This would address issues such as the disconnect between data mining and scheduling, data silos, and the lack of a value feedback loop in existing technologies, thereby improving the overall effectiveness of logistics scheduling across the entire supply chain. Summary of the Invention

[0007] To address the problems existing in the background technology, this invention proposes a joint optimization method for value mining and dynamic scheduling in the logistics data space. This method constructs a scheduling benefit-oriented spatiotemporal federated mining mechanism, a confidence filtering and robust penalty mechanism based on historical scheduling contributions, and a dual-loop feedback mechanism for bidirectional joint optimization of mining and scheduling. This addresses core technical issues in existing technologies, such as the disconnect between mining and scheduling processes leading to a lack of practical scheduling value in mining results, data silos and compliance bottlenecks in cross-entity data fusion, and a lack of quantitative feedback and closed-loop optimization capabilities for mining value. The method achieves joint optimization of data usage compliance, scheduling contribution of mining results, robustness of scheduling decisions, and global net utility in the entire logistics chain scheduling process.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: A joint optimization method for value mining and dynamic scheduling in logistics data space includes the following steps: S1. Obtain the input logistics scheduling task requirements and the cross-entity scheduling resource set in the access logistics data space, complete the standardized preprocessing of the data, and define the global optimization objective of maximizing the total utility of the entire link; S2. Construct a spatiotemporal federated mining model, receive directed data requests from the scheduling layer, perform feature encoding on multi-source heterogeneous data in the logistics data space, and generate candidate mining results oriented towards scheduling needs. S3. Based on the scheduling contribution of historical mining results, calculate the confidence of the current candidate mining results, filter out invalid results below the confidence threshold, and output the final mining result set. S4. Based on real-time logistics operation data, mining result sets, and various resource call costs, construct the state space and action space of the multi-agent scheduling model; S5. Introduce a robust penalty mechanism based on the confidence level of the mining results, and generate and execute the globally optimal scheduling decision through multi-agent reinforcement learning; S6. Collect the actual execution effectiveness of the current scheduling decision and feed it back to the spatiotemporal federated mining model to calculate the actual value contribution of the mining results; S7. Based on the actual value contribution of the mining results, perform online parameter adjustments at the minute level and offline model iterations at the day level to complete the bidirectional joint optimization of mining and scheduling.

[0009] Specifically, step S1 includes: S11. Obtain the input logistics scheduling task requirements, which are a set of logistics orders to be scheduled. Each order includes six attributes: starting location grid code, destination location grid code, promised delivery time, service fee rate, damage compensation coefficient, and goods type. S12. Obtain the cross-entity scheduling resource set of the access logistics data space. The cross-entity scheduling resource set is specifically a transportation capacity resource set. Each transportation capacity resource includes three attributes: transportation capacity identifier, current location grid code, and unit time operating cost. Transportation capacity resources are divided into two categories: self-owned transportation capacity and cross-entity shared transportation capacity. S13. Perform standardized preprocessing on the multi-source heterogeneous data accessed in the logistics data space. The multi-source heterogeneous data includes three categories: real-time road condition data, real-time weather data, and historical dispatch records. Add a four-tuple tag to all accessed multi-source heterogeneous data, including the entity identifier, permission scope, validity period, and unit call cost. Then perform spatiotemporal alignment on the multi-source heterogeneous data, uniformly converting the time dimension to a Unix timestamp and the spatial dimension to a grid code in the GCJ-02 coordinate system. S14. Define the global optimization objective of maximizing the total utility across the entire link: ; in, The core objective of overall optimization is to maximize the net efficiency of the entire logistics chain. The total revenue for the entire logistics service chain is the sum of the service rates for all orders corresponding to the logistics scheduling task requirements. The total cost of operation across the entire value chain is calculated using the following formula: ; in, This represents the total number of schedulable transport capacity sets. For the first The unit time operating cost of the transport capacity For the first The actual working hours of the transport capacity; The calculation formula for the end-to-end timeout and damage penalty cost is as follows: ; in, This represents the total number of orders corresponding to the logistics scheduling task requirements. For the first Timeout penalty coefficient for each order. For the first The actual delivery time for each order. For the first Promised delivery time for each order For the first Damage compensation coefficient for each order. For the first Damage status indicator for each order. A value of 1 indicates that the order was damaged, and a value of 0 indicates that the order was not damaged. The total cost of using the data is calculated using the following formula: ; in, The total number of data types called. For the first The unit cost of accessing class data. For the first The actual number of times the class data is accessed.

[0010] Specifically, step S2 includes: S21. Receive the directional data request from the scheduling layer. The directional data request includes three attributes: the time range to be covered, the spatial grid range to be covered, and the type of mining results. The mining result type includes two options: grid-level order volume and grid-level average delivery time. S22. Construct the inference architecture of the spatiotemporal federated mining model. Adopt a horizontal federated learning framework to adapt to cross-subject data security requirements. Each data participant in the cross-subject deploys a spatiotemporal long short-term memory network sub-model locally. Each sub-model only calls multi-source heterogeneous data within the scope of the subject's authority. During the inference process, the model gradient is only uploaded to the global aggregation node of the logistics data space. S23. Perform feature encoding on the multi-source heterogeneous data of the input sub-model. The encoding process includes: a. Extract time dimension features. Four types of sub-features are derived from the Unix timestamps in historical scheduling records: hourly time period, weekday attribute, holiday attribute, and monthly attribute. Each type of sub-feature is concatenated after one-hot encoding to obtain the time dimension features, which are used to capture the time cycle pattern of logistics data. b. Extract spatial dimension features. Three types of sub-features are derived from the grid codes in historical scheduling records: the administrative region to which the grid belongs, road network density, and delivery difficulty. The three types of sub-features are spliced ​​together to obtain spatial dimension features, which are used to capture the spatial correlation patterns of logistics data. c. Extract attribute dimension features, including four types of sub-features: order attributes, transportation cost attributes, road condition attributes, and weather attributes. Among them, order attributes are taken from logistics scheduling task requirements, transportation cost attributes are taken from cross-entity scheduling resource sets, road condition attributes are taken from real-time road condition data, and weather attributes are taken from real-time weather data. All sub-features are concatenated after min-max standardization to obtain attribute dimension features, which are used to capture the business attribute patterns of logistics data. d. Concatenate the three types of features into a standardized feature vector: ; in, To standardize the feature vector, Features in the time dimension Spatial dimensional features, For attribute dimension features; S24. The global aggregation node in the logistics data space receives the model inference gradients uploaded by all data participants, and aggregates them by weighted average based on the effective sample size to obtain the global model parameters. The aggregation process satisfies: ; in, This is the global parameter set for the spatiotemporal federated mining model; The total number of data subjects participating in federal computing; For the first The number of valid historical samples for each data subject within the current time and space range of the request; The total number of valid historical samples for all data subjects; For the first A set of local sub-model inference parameters for each data subject; S25. The aggregated global model parameters are distributed to the local sub-models of each data participant. Reasoning is performed based on the attributes of the targeted data request to generate candidate mining results. The candidate mining results include the predicted order volume and the predicted average delivery time for each grid in the next few hours.

[0011] Specifically, the training process of the spatiotemporal federated mining model is as follows: a. Each data participant's local sub-model undergoes iterative training based on local anonymized historical data. The training process employs a loss function that integrates scheduling benefits, specifically: ; in, The total loss of the spatiotemporal federated mining model is defined as the value to be minimized during the training phase. , , Let be the weighting coefficient, satisfying , To predict losses, the baseline accuracy used to constrain the mining results is calculated using the following formula: ; in, This represents the total number of training samples. For the first Mining prediction value for each sample, For the first The true value corresponding to each sample is taken from the actual order volume and actual delivery time data within the same spatial range in the historical scheduling records; The normalized value contributes to the scheduling in the current mining scenario, which is used to characterize the ability of mining results to improve scheduling benefits; The federal consistency loss is represented by the value of the first iteration. Local submodel parameter set for each data subject With the global model parameter set L2 distance: ; b. After each data participant completes a preset number of local iterations of training, it uploads the model gradient to the global aggregation node of the logistics data space once. The global aggregation node updates the global model parameters according to the aggregation rules in step S24 and then distributes them to each local sub-model to complete one round of federated training iteration. c. When the decrease in total loss over a preset number of iterations is less than a predetermined threshold, training is stopped, and the final spatiotemporal federated mining model is obtained. The process by which the spatiotemporal federated mining model generates mining results during the inference phase is as follows: (1) After receiving the targeted data request, extract the multi-source heterogeneous data within the time range and spatial grid range covered by the request, and generate a standardized feature vector through feature encoding in step S23; (2) Input the standardized feature vector into the trained spatiotemporal federated mining model, extract the time series periodicity of the stream data by the time gate of the spatiotemporal long short-term memory network, extract the spatial correlation of adjacent grids by the spatial gate, and obtain the predicted output value after mapping through the fully connected layer. The predicted output value is an unlabeled set of original values, which includes the original predicted values ​​of order volume and average delivery time corresponding to the mining result type of the targeted data request. (3) Format the predicted output value according to the three attributes of the targeted data request to generate candidate mining results.

[0012] Specifically, step S3 includes: S31. Receive candidate mining results from the independent inference output of the local sub-models of each data participant, and form a set of candidate mining results; S32. For each group of candidate mining results, match historical scheduling data of the same scenario; S33. Based on the scheduling contribution of historical mining results, calculate the confidence level of each group of candidate mining results. The calculation formula is as follows: ; in, The confidence level of the current candidate mining results. The weight for historical scheduling is determined by the average net utility improvement rate brought about by similar mining results in the past 30 days. The output value of the current candidate mining result generated in step S2; This represents the average of the historical true values ​​of the corresponding mining results in the historical scheduling data of the same scenario. Minimum protection threshold, used to avoid Division by zero error when the result is 0; S34. Set a confidence threshold and perform two-level filtering on the candidate mining result set: The first level filters candidate results with confidence scores below the threshold and directly removes invalid results; the second level retains the group with the highest confidence score as the valid result for the same spatiotemporal dimension unit and the same mining result type of the remaining candidate results. S35. Integrate the valid results of all spatiotemporal dimension units to generate the final mining result set.

[0013] Specifically, step S4 includes: S41. Collect three types of input data required for constructing the state space and action space. The first type is mining feature data, which is taken from the final mining result set generated in step S3 and the standardized feature vector generated in step S2. The second type is real-time operation data, including the real-time order volume, transportation resource location, real-time road condition data, and real-time weather data of each spatial grid. The third type is cost feature data, which is taken from the unit call cost in the quadruple label of the multi-source heterogeneous data in step S1 and the unit time operation cost of the cross-subject scheduling resource set. S42. Perform min-max standardization on all input features. The standardization process satisfies the following expression: ; in, The original values ​​of the input features. , The minimum and maximum values ​​of this type of feature in the historical scheduling records over the past 30 days. As the minimum protection threshold, These are the standardized feature values; S43. Concatenate the standardized three types of features into a state space vector: ; in, This represents the state space vector of a multi-agent scheduling model. To extract feature subvectors, To run feature subvectors in real time, For cost feature vectors; S44. Construct the action space of the multi-agent scheduling model. The global action space is defined by the following expression: ; in, For the global action space set, Let be the total number of schedulable capacity units, and let each schedulable capacity unit correspond to an independent intelligent agent. Action vectors of each agent For 2D vectors: ; in, This is the task acceptance decision dimension, with a value of 0 or 1. 0 indicates that the capacity rejects the currently assigned order task, and 1 indicates that the capacity accepts the currently assigned order task. A dimension is selected for the driving path, with a normalized value ranging from 0 to 1. The normalization process satisfies the following conditions: ,in, This represents the total number of alternative routes generated based on real-time traffic data. This is an index for alternative paths.

[0014] Specifically, step S5 includes: S51. Build a multi-agent deep deterministic policy gradient scheduling framework, including two stages: offline training and online inference. S52. Define the target reward function for the offline training phase. The target reward function integrates the robustness penalty for mining confidence and the operational constraint penalty, specifically as follows: ; in, for The reward value corresponding to the scheduling decision at each moment is used to iterate the model parameters during the offline training phase with the goal of maximizing the cumulative reward value. for The total net utility across the entire link corresponding to the time-sharing decision. This is the robustness penalty coefficient, used to control the severity of the penalty for deviations in the mining results. This is a predicted value based on the excavation results from the previous moment. This represents the actual value collected after the previous scheduling execution. This represents the confidence level label for the mining result at the previous time step. This is the operational constraint penalty coefficient, used to control the severity of penalties for violating operational rules; The operational constraint penalty value is calculated as follows: ; in, This represents the total number of schedulable transport capacity sets. For the first The cumulative working hours of each transport capacity This is the maximum daily working hours. The total number of orders to be scheduled. For the first The estimated delivery time for each order is calculated based on real-time traffic data. For the first Promised delivery time for each order; S53. Initial scheduling decision is generated during the online inference stage: The state space vector constructed in step S4 is input into the trained multi-agent scheduling model. The model outputs a corresponding 2D action vector for each capacity agent independently. The action vectors of all agents are combined to form the initial scheduling decision. S54. Perform global constraint verification on the initial scheduling decision. The verification rules must satisfy: ; in, For the first The capacity of the first transport to the first Decision-making for receiving individual orders; S55. Execute quantitative guaranteed minimum scheduling trigger judgment and emergency handling, trigger conditions are met: ; in, The number of scheduling cycles for continuous monitoring. For the first The average confidence level of the mining results over a scheduling cycle The confidence threshold; When the triggering condition is met, the backup scheduling mode is triggered, and the following operations are performed: a. Reconstruct the state space vector, setting all mined feature vectors to 0, retaining only the real-time running feature vector and the cost feature vector. The reconstructed state space satisfies: ; b. Route selection directly adopts the alternative route with the shortest travel time under real-time traffic conditions, without considering the prediction of future traffic conditions; c. During the scheduling process, a preset amount of idle transportation capacity is always reserved as an emergency reserve to deal with unexpected sudden increases in order volume or abnormal road conditions; S56. The scheduling decision in the output verification passed or the backup scheduling mode is taken as the global optimal scheduling decision and executed.

[0015] Specifically, step S6 involves: collecting actual operational data after the scheduling decision is executed in step S5, and calculating the actual total net utility of the current scheduling. And based on actual total net utility The overall actual value contribution of the mining results is calculated using the following formula: ; in, This contributes to the overall value of the data collected during this excavation. The average net scheduling utility of the baseline without using mining results in the same scenario is taken from historical scheduling records; Then, the marginal value contribution of a single-class spatiotemporal attribute feature is calculated, satisfying the following expression: ; in, For the first The actual value contribution of spatiotemporal attribute features Corresponding to time dimension features, Corresponding spatial dimensional features, Corresponding attribute dimension features; To make the first After setting all spatiotemporal attribute features to 0, the simulated net utility is obtained using the same scheduling model and real-time data inference.

[0016] Specifically, in step S7, the online parameter adjustment involves adjusting the input weights of the three types of features in the spatiotemporal federated mining model. The adjustment rules are as follows: ; in, For the adjusted number The input weights for the class-mining features are uniformly distributed initially. To adjust the previous number The current weights of the class-mined features. This is the online learning rate coefficient, to avoid excessive weight fluctuations affecting the stability of the mining process; The offline iteration cycle of the model is during the daily off-peak logistics period, updating the core parameters of the spatiotemporal federated mining model and the multi-agent scheduling model respectively, including: (1) Update the weights of the scheduling revenue term in the loss function of the spatiotemporal federated mining model, satisfying the following expression: ; in, The updated scheduling benefit item weights, The weights of the scheduling benefit items before the update. For the corresponding offline update coefficient, The average value contribution to the overall mining value across all scheduling cycles on that day; (2) Update the robustness penalty coefficient of the reward function of the multi-agent scheduling model to satisfy the following expression: ; in, The updated robustness penalty coefficient, The robustness penalty coefficient before the update. For the corresponding offline update coefficient, The average confidence level of all excavation results for that day; After the iteration is completed, the updated parameters are sent to the spatiotemporal federated mining model and the multi-agent scheduling model respectively, for the next round of mining inference and scheduling decision-making, thus completing the bidirectional joint optimization closed loop of mining and scheduling.

[0017] In summary, the beneficial technical effects of the present invention are as follows: 1. Achieving deep collaboration between mining and scheduling: This invention constructs a spatiotemporal federated mining model oriented towards scheduling benefits, directly integrating the scheduling contribution of mining results into the loss function. This makes the mining process no longer solely focused on prediction accuracy, but rather optimized for scheduling benefits, thus solving the problem of mining results lacking practical value due to the disconnect between mining and scheduling processes.

[0018] 2. Achieving compliant cross-entity data collaboration: This invention adopts a horizontal federated learning framework, in which each data participant only uploads model gradients without sharing the original data. Under the premise of meeting data security and compliance requirements, it achieves effective integration of cross-entity data value and solves the data silos and compliance bottlenecks in multi-entity collaboration scenarios.

[0019] 3. Achieving robust scheduling driven by confidence, ensuring decision stability: This invention introduces confidence calculation based on historical scheduling contribution and a two-level filtering mechanism to eliminate invalid mining results, and integrates confidence-based robust penalty into the multi-agent scheduling model. When the mining quality fluctuates, a guaranteed scheduling mode is automatically triggered, effectively improving the robustness of scheduling decisions to mining deviations.

[0020] 4. Achieving a two-way joint optimization closed loop: This invention constructs a dual-loop feedback mechanism of minute-level online adjustment and day-level offline iteration. By quantifying the actual value contribution of the mined features, it dynamically optimizes the feature weights of the mining model and the robustness penalty coefficient of the scheduling model, realizing the two-way collaborative evolution of mining and scheduling, and enabling the system to have the ability to continuously self-optimize. Attached Figure Description

[0021] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0022] To make the technical means, creative features, objectives and effects of this invention clearer and easier to understand, the invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0023] Example like Figure 1As shown, the joint optimization method for value mining and dynamic scheduling in logistics data space provided by this invention specifically includes the following steps: S1. Obtain the input logistics scheduling task requirements and the cross-entity scheduling resource set in the access logistics data space, complete the standardized preprocessing of the data, and define the global optimization objective of maximizing the total utility of the entire link; this step specifically includes: S11. Obtain the input logistics scheduling task requirements. The logistics scheduling task requirements are a set of logistics orders to be scheduled. Each order includes six attributes: starting location grid code, destination location grid code, promised delivery time, service fee rate, damage compensation coefficient, and goods type. S12. Obtain the cross-entity scheduling resource set of the access logistics data space. The cross-entity scheduling resource set is specifically a transportation capacity resource set. Each transportation capacity resource includes three attributes: transportation capacity identifier, current location grid code, and unit time operating cost. Transportation capacity resources are divided into two categories: owned transportation capacity and cross-entity shared transportation capacity. S13. Perform standardized preprocessing on the multi-source heterogeneous data accessed by the logistics data space. The multi-source heterogeneous data includes three categories: real-time road condition data, real-time weather data, and historical dispatch records. Add a four-tuple tag to all accessed multi-source heterogeneous data, including the entity identifier, permission scope, validity period, and unit call cost. All data calls must be verified through the logistics data space authorization interface to ensure compliance with the permission scope. Requests that fail the verification will be blocked. Subsequently, spatiotemporal alignment was performed on the multi-source heterogeneous data. The original times of all data, such as 2024-10-01 18:00 / 10 / 1 / 24 6PM and other different formats, were uniformly converted into 13-bit integer Unix timestamps with a precision of 1 minute. At the same time, the original locations of all data were uniformly mapped to a 100m×100m grid in the GCJ-02 coordinate system. Each grid has a unique numerical code, thus making the time standard and location standard of all data completely unified. Furthermore, the 3σ principle is used to filter out outlier data. For continuous features such as delivery time and order volume in historical scheduling records, the mean and standard deviation are calculated to remove outlier data whose values ​​are outside the range. S14. Define the global optimization objective of maximizing the total utility across the entire link: ; in, The core objective of overall optimization is to maximize the net efficiency of the entire logistics chain. The total revenue for the entire logistics service chain is the sum of the service rates for all orders corresponding to the logistics scheduling task requirements. The total cost of operation across the entire value chain is calculated using the following formula: ; in, This represents the total number of schedulable transport capacity sets. For the first The unit time operating cost of the transport capacity For the first The actual working hours of the transport capacity; The calculation formula for the end-to-end timeout and damage penalty cost is as follows: ; in, This represents the total number of orders corresponding to the logistics scheduling task requirements. For the first Timeout penalty coefficient for each order. For the first The actual delivery time for each order. For the first Promised delivery time for each order For the first Damage compensation coefficient for each order. For the first Damage status indicator for each order. A value of 1 indicates that the order was damaged, and a value of 0 indicates that the order was not damaged. The total cost of using the data is calculated using the following formula: ; in, The total number of data types called. For the first The unit cost of accessing class data. For the first The actual number of times the class data is accessed.

[0024] S2. Construct a spatiotemporal federated mining model, receive targeted data requests from the scheduling layer, perform feature encoding on multi-source heterogeneous data in the logistics data space, and generate candidate mining results oriented towards scheduling needs; this step specifically includes: S21. The scheduling layer generates targeted data requests based on the current logistics scheduling task requirements, including three attributes: the time range covered by the request, the spatial grid range covered by the request, and the type of mining results. The mining result type specifically includes two options: grid-level order volume and grid-level average delivery time. This step is used to limit the spatiotemporal range of mining and the output content. S22. Construct the inference architecture of the spatiotemporal federated mining model. Adopt a horizontal federated learning framework to adapt to cross-subject data security requirements. Each data participant in the cross-subject deploys a spatiotemporal long short-term memory network sub-model locally. Each sub-model only calls multi-source heterogeneous data within the scope of the subject's authority. During the inference process, only the model gradient is uploaded to the global aggregation node of the logistics data space, and no original business data is transmitted. S23. Perform feature encoding on the multi-source heterogeneous data of the input sub-model. The encoding process includes: a. Extract time-dimensional features. Using the Unix timestamp inherent in each historical scheduling sample as the anchor point, four types of sub-features are derived: hourly time period, weekday attribute, holiday attribute, and monthly attribute. The hourly time period is divided into 4 groups according to different times. The weekday attribute is divided into 2 groups according to weekdays and non-weekdays. The holiday attribute is divided into 2 groups according to weekdays and holidays. The monthly attribute is divided into 4 groups according to quarters. After one-hot encoding of each atomic feature, the time-dimensional features are concatenated to capture the time cycle pattern of logistics data. b. Extract spatial dimension features. Using the 100-meter grid code inherent in each historical scheduling sample as the anchor point, three types of sub-features are obtained: the administrative region to which the grid belongs, road network density, and delivery difficulty. The administrative region to which the grid belongs is matched from the publicly available administrative division mapping table in the logistics data space. The road network density is matched from the static road network basic data published by the transportation department, and the value is taken as the proportion of road area within the grid and standardized by min-max. The delivery difficulty is taken from the normalized value of the average delivery time of the corresponding grid in the historical scheduling record. The three types of sub-features are concatenated to obtain spatial dimension features, which are used to capture the spatial correlation patterns of logistics data. c. Extract attribute dimension features, including four types of sub-features: order attributes, transportation cost attributes, road condition attributes, and weather attributes. Among them, order attributes are taken from logistics scheduling task requirements, transportation cost attributes are taken from cross-entity scheduling resource sets, road condition attributes are taken from real-time road condition data, and weather attributes are taken from real-time weather data. All sub-features are concatenated after min-max standardization to obtain attribute dimension features, which are used to capture the business attribute patterns of logistics data. d. Concatenate the three types of features into a standardized feature vector: ; in, To standardize the feature vector, Features in the time dimension Spatial dimensional features, For attribute dimension features; S24. The global aggregation node in the logistics data space receives model inference gradients uploaded by all core data participants, including e-commerce platforms and express delivery companies. It then aggregates these gradients using a weighted average based on the effective sample size to obtain the global model parameters. The aggregation process satisfies the following: ; in, This is the global parameter set for the spatiotemporal federated mining model; The total number of data subjects participating in federal computing; For the first The number of valid historical samples for each data subject within the current time and space range of the request; The total number of valid historical samples for all data subjects; For the first A set of local sub-model inference parameters for each data subject; This step uses a sample size-weighted aggregation rule to ensure that subjects with larger sample sizes and richer data have a greater impact on the global model. S25. Distribute the aggregated global model parameters to the local sub-models of each data participant, perform inference based on the attributes of the targeted data request, and generate candidate mining results. The candidate mining results include the predicted order volume and the predicted average delivery time for each grid in the next few hours.

[0025] The training process of the above spatiotemporal federated mining model is as follows: a. Each data participant's local sub-model undergoes iterative training based on local anonymized historical data. The training process employs a loss function that integrates scheduling benefits, specifically: ; in, The total loss of the spatiotemporal federated mining model is defined as the value to be minimized during the training phase. , , Let be the weighting coefficient, satisfying , To predict losses, the baseline accuracy used to constrain the mining results is calculated using the following formula: ; in, This represents the total number of training samples. For the first Mining prediction value for each sample, For the first The true value corresponding to each sample is taken from the actual order volume and actual delivery time data within the same spatial range in the historical scheduling records; The normalized value of the scheduling contribution in the current mining scenario is used to characterize the ability of the mining results to improve the scheduling benefits; the calculation process is uniformly executed by the global scheduling feedback module of the logistics data space, specifically: (1) Matching historical scheduling records that are completely consistent with the current mining scenario, the matching rules are aligned with the three attributes of the directional data request in step S21, including time range deviation of less than 7 days, spatial grid range overlap of more than 90%, and consistent mining result type; (2) Calculating the average net scheduling utility of using this type of mining result in the historical records obtained by matching. Average net utility of baseline scheduling without using mining results (3) Calculate the normalized value of the scheduling contribution, which satisfies the following expression: ; in, The highest net scheduling utility in the same scenario over the past 30 days is used for normalization to ensure... The value always falls within the range of 0 to 1; if If the value is less than 0, it is 0, which means that this type of mining result cannot bring about an increase in revenue; The federal consistency loss is represented by the value of the first iteration. Local submodel parameter set for each data subject With the global model parameter set L2 distance: ; b. After each data participant completes a preset number of local iterations of training, it uploads the model gradient to the global aggregation node of the logistics data space once. The global aggregation node updates the global model parameters according to the aggregation rules in step S24 and then distributes them to each local sub-model to complete one round of federated training iteration. c. When the decrease in total loss over a preset number of iterations is less than a predetermined threshold, training is stopped, and the final spatiotemporal federated mining model is obtained. The process by which the spatiotemporal federated mining model generates mining results during the inference phase is as follows: (1) After receiving the targeted data request, extract the multi-source heterogeneous data within the time range and spatial grid range covered by the request, and generate a standardized feature vector through feature encoding in step S23; (2) Input the standardized feature vector into the trained spatiotemporal federated mining model. Extract the time series periodicity of the stream data from the time gate of the spatiotemporal long short-term memory network, and extract the spatial correlation of adjacent grids from the spatial gate. After mapping through the fully connected layer, the predicted output value is obtained. The predicted output value is an unlabeled set of original values, which includes the original predicted values ​​of order volume and average delivery time corresponding to the mining result type of the targeted data request. (3) Format the predicted output value according to the three attributes of the targeted data request to generate candidate mining results.

[0026] S3. Based on the scheduling contribution of historical mining results, calculate the confidence level of the current candidate mining results, filter out invalid results below the confidence threshold, and output the final mining result set; the specific steps include: S31. Receive candidate mining results from the independent inference output of the local sub-models of each data participant, and form a set of candidate mining results; where each set of candidate results corresponds to a spatiotemporal dimension unit and a mining result type. The spatiotemporal dimension unit is a combination of a 1-hour time slice and a 100-meter spatial grid. The mining result type is grid-level order volume or grid-level average delivery time. S32. For each group of candidate mining results, match historical scheduling data in the same scenario; the matching rules are aligned with the three attributes of the directional data request in step S2, namely, the time range deviation is less than 7 days, the spatial grid range overlap is greater than 90%, and the mining result type is consistent, and the historical scheduling data in the same scenario is taken from the historical scheduling records. S33. Based on the scheduling contribution of historical mining results, calculate the confidence level of each group of candidate mining results. The calculation formula is as follows: ; in, The confidence level of the current candidate mining results. The weight for contributing to historical scheduling is determined by the average net utility improvement rate brought by the same type of mining results in the past 30 days. The average net utility improvement rate is specifically the difference between the net utility of scheduling using this type of mining results and the baseline net utility of scheduling without using mining results, and the ratio to the baseline net utility of scheduling. The output value of the current candidate mining result generated in step S2; This is the average of the historical true values ​​of the corresponding mining results in the historical scheduling data of the same scenario, taken from the actual order volume and actual delivery time data of the historical scheduling records. Minimum protection threshold, used to avoid Division by zero error when the result is 0; S34. Set a confidence threshold and perform two-level filtering on the candidate mining result set: The first level filters candidate results with confidence scores below the threshold and directly removes invalid results; the second level retains the group with the highest confidence score as the valid result for the same spatiotemporal dimension unit and the same mining result type of the remaining candidate results. S35. Integrate the valid results of all spatiotemporal dimension units to generate the final mining result set; each result in the final mining result set is accompanied by a corresponding confidence label.

[0027] S4. Based on real-time logistics operation data, mining result sets, and various resource call costs, construct the state space and action space of the multi-agent scheduling model; the specific steps include: S41. Collect three types of input data required for constructing the state space and action space. The first type is mining feature data, which is taken from the final mining result set generated in step S3 and the standardized feature vector generated in step S2. The second type is real-time operation data, including the real-time order volume, transportation resource location, real-time road condition data, and real-time weather data of each spatial grid. The third type is cost feature data, which is taken from the unit call cost in the quadruple label of the multi-source heterogeneous data in step S1 and the unit time operation cost of the cross-subject scheduling resource set. S42. Perform min-max standardization on all input features to uniformly map feature values ​​to the interval between 0 and 1, adapting to the input requirements of multi-agent reinforcement learning models. The standardization process satisfies the following expression: ; in, The original values ​​of the input features. , The minimum and maximum values ​​of this type of feature in the historical scheduling records over the past 30 days. As the minimum protection threshold, These are the standardized feature values; S43. Concatenate the standardized three types of features into a state space vector: ; in, This represents the state space vector of a multi-agent scheduling model. To extract feature subvectors, To run feature subvectors in real time, For cost feature vectors; S44. Construct the action space of the multi-agent scheduling model. The global action space is defined by the following expression: ; in, For the global action space set, Let be the total number of schedulable capacity units, and let each schedulable capacity unit correspond to an independent intelligent agent. Action vectors of each agent For 2D vectors: ; in, This is the task acceptance decision dimension, with a value of 0 or 1. 0 indicates that the capacity rejects the currently assigned order task, and 1 indicates that the capacity accepts the currently assigned order task. A dimension is selected for the driving path, with a normalized value ranging from 0 to 1. The normalization process satisfies the following conditions: ,in, The total number of alternative routes generated based on real-time traffic data is fixed at 3. The alternative routes are calculated based on the congestion level of real-time traffic conditions. An index for alternative paths; S45. The spatiotemporal granularity of the unified state space and action space is 1 minute time granularity and 100-meter spatial grid granularity, so as to be completely aligned with the spatiotemporal standard preprocessed in step S1.

[0028] S5. Introduce a robust penalty mechanism based on the confidence level of the mining results, and generate and execute the globally optimal scheduling decision through multi-agent reinforcement learning; the specific steps include: S51. Build a multi-agent deep deterministic policy gradient scheduling framework, including two stages: offline training and online inference. S52. Define the target reward function for the offline training phase. The target reward function integrates the robustness penalty for mining confidence and the operational constraint penalty, specifically as follows: ; in, for The reward value corresponding to the scheduling decision at each moment is used to iterate the model parameters during the offline training phase with the goal of maximizing the cumulative reward value. for The total net utility across the entire link corresponding to the time-sharing decision. This is the robustness penalty coefficient, used to control the severity of the penalty for deviations in the mining results. This is a predicted value based on the excavation results from the previous moment. This represents the actual value collected after the previous scheduling execution. This represents the confidence level label for the mining result at the previous time step. This is the operational constraint penalty coefficient, used to control the severity of penalties for violating operational rules; The operational constraint penalty value is calculated as follows: ; in, This represents the total number of schedulable transport capacity sets. For the first The cumulative working hours of each transport capacity This is the maximum daily working hours. The total number of orders to be scheduled. For the first The estimated delivery time for each order is calculated based on real-time traffic data. For the first Promised delivery time for each order; S53. Initial scheduling decision is generated during the online inference stage: The state space vector constructed in step S4 is input into the trained multi-agent scheduling model. The model outputs a corresponding 2D action vector for each capacity agent independently. The action vectors of all agents are combined to form the initial scheduling decision. S54. Perform global constraint verification on the initial scheduling decision. The verification rules must satisfy: ; in, For the first The capacity of the first transport to the first The task acceptance decision for each order; this formula indicates that each order can be assigned to at most one capacity, avoiding duplicate assignment; S55. Execute quantitative guaranteed minimum scheduling trigger judgment and emergency handling, trigger conditions are met: ; in, The number of scheduling cycles for continuous monitoring. For the first The average confidence level of the mining results over a scheduling cycle The confidence threshold; When the triggering condition is met, the backup scheduling mode is triggered, and the following operations are performed: a. Reconstruct the state space vector, setting all mined feature vectors to 0, retaining only the real-time running feature vector and the cost feature vector. The reconstructed state space satisfies: ; b. Route selection directly adopts the alternative route with the shortest travel time under real-time traffic conditions, without considering the prediction of future traffic conditions; c. During the scheduling process, a preset amount of idle transportation capacity is always reserved as an emergency reserve to deal with unexpected sudden increases in order volume or abnormal road conditions; S56. The scheduling decision that passes the output verification or is in the minimum guarantee scheduling mode is taken as the global optimal scheduling decision and executed. All running data during the scheduling execution process is synchronously uploaded to the subsequent step S6 for utility calculation and model feedback.

[0029] S6. Collect the actual execution utility of the current scheduling decision and feed it back to the spatiotemporal federated mining model to calculate the actual value contribution of the mining results; the specific steps are: collect the actual operation data after the execution of the scheduling decision in step S5, and calculate the actual total net utility of the current scheduling. And based on actual total net utility The overall actual value contribution of the mining results is calculated using the following formula: ; in, This contributes to the overall value of the data collected during this excavation. The average net scheduling utility of the baseline without using mining results in the same scenario is taken from historical scheduling records; Subsequently, the marginal value contribution of single-class spatiotemporal attribute features is calculated. Specifically, the spatiotemporal attribute features are the time dimension features, spatial dimension features, and attribute dimension features in the standardized feature vector of step S23. The marginal value contribution calculation satisfies the following expression: ; in, For the first The actual value contribution of spatiotemporal attribute features Corresponding to time dimension features, Corresponding spatial dimensional features, Corresponding attribute dimension features; To make the first After setting all spatiotemporal attribute features to 0, the simulated net utility obtained by using the same scheduling model and real-time data inference is used. The simulation process only replaces the feature inputs, and the remaining parameters are the same as the real scheduling.

[0030] S7. Based on the actual value contribution of the mining results, perform minute-level online parameter adjustments and day-level offline model iterations to complete the bidirectional joint optimization of mining and scheduling; specifically, the online parameter adjustment involves adjusting the input weights of the three types of features in the spatiotemporal federated mining model, with the following adjustment rules: ; in, For the adjusted number The input weights for the class-mining features are uniformly distributed initially. To adjust the previous number The current weights of the class-mined features. This is the online learning rate coefficient, to avoid excessive weight fluctuations affecting the stability of mining; this formula means that the higher the value contribution of the mined feature, the greater the weight increase, and the higher the feature priority in subsequent mining inference; The offline iteration cycle of the model is during the daily off-peak logistics period, updating the core parameters of the spatiotemporal federated mining model and the multi-agent scheduling model respectively, including: (1) Update the weights of the scheduling revenue term in the loss function of the spatiotemporal federated mining model, satisfying the following expression: ; in, The updated scheduling benefit item weights, The weights of the scheduling benefit items before the update. For the corresponding offline update coefficient, This is the average value contribution of mining across all scheduling cycles on that day. This formula indicates that the higher the value contribution of the mining results on that day, the more likely the subsequent mining model will output high-yield mining results. The upper limit of the weight is set to 0.5 to avoid excessive deviation from the basic prediction accuracy requirements.

[0031] (2) Update the robustness penalty coefficient of the reward function of the multi-agent scheduling model to satisfy the following expression: ; in, The updated robustness penalty coefficient, The robustness penalty coefficient before the update. For the corresponding offline update coefficient, The average confidence level of all mining results for the day; this formula indicates that the higher the confidence level of the mining results for the day, the lower the penalty of the scheduling model on the mining results, and the more trusting the mining output. The lower limit of the coefficient is set to 0.05 to avoid completely losing the robustness constraint. After the iteration is completed, the updated parameters are sent to the spatiotemporal federated mining model and the multi-agent scheduling model respectively, for the next round of mining inference and scheduling decision-making, thus completing the bidirectional joint optimization closed loop of mining and scheduling.

[0032] Therefore, the value mining and dynamic scheduling joint optimization method for logistics data space provided by this invention achieves cross-entity data compliance collaboration and high-value mining result output by constructing a scheduling benefit-oriented spatiotemporal federated mining mechanism. Based on the confidence filtering and robust penalty mechanism of historical scheduling contribution, invalid mining results are eliminated and the stability of scheduling decisions is improved. Finally, through a dual-loop feedback mechanism of minute-level online adjustment and day-level offline iteration, bidirectional joint optimization of mining and scheduling is achieved. This solves the pain points of existing technologies such as the disconnect between mining and scheduling processes, data silos and compliance bottlenecks in cross-entity data fusion, and the lack of quantitative feedback and closed-loop optimization capabilities for mining value. It realizes the joint optimization of data usage compliance, scheduling contribution of mining results, robustness of scheduling decisions, and global net utility in logistics full-link scheduling, providing a high-value, robust, and continuously self-optimizing solution for multi-entity collaborative intelligent logistics scheduling.

[0033] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A joint optimization method for value mining and dynamic scheduling in logistics data space, applied to cross-entity collaborative logistics end-to-end resource scheduling scenarios, characterized by: Includes the following steps: S1. Obtain the input logistics scheduling task requirements and the cross-entity scheduling resource set in the access logistics data space, complete the standardized preprocessing of the data, and define the global optimization objective of maximizing the total utility of the entire link; S2. Construct a spatiotemporal federated mining model, receive directed data requests from the scheduling layer, perform feature encoding on multi-source heterogeneous data in the logistics data space, and generate candidate mining results oriented towards scheduling needs. S3. Based on the scheduling contribution of historical mining results, calculate the confidence of the current candidate mining results, filter out invalid results below the confidence threshold, and output the final mining result set. S4. Based on real-time logistics operation data, mining result sets, and various resource call costs, construct the state space and action space of the multi-agent scheduling model; S5. Introduce a robust penalty mechanism based on the confidence level of the mining results, and generate and execute the globally optimal scheduling decision through multi-agent reinforcement learning; S6. Collect the actual execution effectiveness of the current scheduling decision and feed it back to the spatiotemporal federated mining model to calculate the actual value contribution of the mining results; S7. Based on the actual value contribution of the mining results, perform online parameter adjustments at the minute level and offline model iterations at the day level to complete the bidirectional joint optimization of mining and scheduling.

2. The joint optimization method for value mining and dynamic scheduling in logistics data space according to claim 1, characterized in that, Step S1 specifically includes: S11. Obtain the input logistics scheduling task requirements, which are a set of logistics orders to be scheduled. Each order includes six attributes: starting location grid code, destination location grid code, promised delivery time, service fee rate, damage compensation coefficient, and goods type. S12. Obtain the cross-entity scheduling resource set of the access logistics data space. The cross-entity scheduling resource set is specifically a transportation capacity resource set. Each transportation capacity resource includes three attributes: transportation capacity identifier, current location grid code, and unit time operating cost. Transportation capacity resources are divided into two categories: self-owned transportation capacity and cross-entity shared transportation capacity. S13. Perform standardized preprocessing on the multi-source heterogeneous data accessed in the logistics data space. The multi-source heterogeneous data includes three categories: real-time road condition data, real-time weather data, and historical dispatch records. Add a four-tuple tag to all accessed multi-source heterogeneous data, including the entity identifier, permission scope, validity period, and unit call cost. Then perform spatiotemporal alignment on the multi-source heterogeneous data, uniformly converting the time dimension to a Unix timestamp and the spatial dimension to a grid code in the GCJ-02 coordinate system. S14. Define the global optimization objective of maximizing the total utility across the entire link: ; in, The core objective of overall optimization is to maximize the net efficiency of the entire logistics chain. The total revenue for the entire logistics service chain is the sum of the service rates for all orders corresponding to the logistics scheduling task requirements. The total cost of operation across the entire value chain is calculated using the following formula: ; in, This represents the total number of schedulable transport capacity sets. For the first The unit time operating cost of the transport capacity For the first The actual working hours of the transport capacity; The calculation formula for the end-to-end timeout and damage penalty cost is as follows: ; in, This represents the total number of orders corresponding to the logistics scheduling task requirements. For the first Timeout penalty coefficient for each order. For the first The actual delivery time for each order. For the first Promised delivery time for each order For the first Damage compensation coefficient for each order. For the first Damage status indicator for each order. A value of 1 indicates that the order was damaged, and a value of 0 indicates that the order was not damaged. The total cost of using the data is calculated using the following formula: ; in, The total number of data types called. For the first The unit cost of accessing class data. For the first The actual number of times the class data is accessed.

3. The joint optimization method for value mining and dynamic scheduling in logistics data space according to claim 2, characterized in that, Step S2 specifically includes: S21. Receive the directional data request from the scheduling layer. The directional data request includes three attributes: the time range to be covered, the spatial grid range to be covered, and the type of mining results. The mining result type includes two options: grid-level order volume and grid-level average delivery time. S22. Construct the inference architecture of the spatiotemporal federated mining model. Adopt a horizontal federated learning framework to adapt to cross-subject data security requirements. Each data participant in the cross-subject deploys a spatiotemporal long short-term memory network sub-model locally. Each sub-model only calls multi-source heterogeneous data within the scope of the subject's authority. During the inference process, the model gradient is only uploaded to the global aggregation node of the logistics data space. S23. Perform feature encoding on the multi-source heterogeneous data of the input sub-model. The encoding process includes: a. Extract time dimension features. Four types of sub-features are derived from the Unix timestamps in historical scheduling records: hourly time period, weekday attribute, holiday attribute, and monthly attribute. Each type of sub-feature is concatenated after one-hot encoding to obtain the time dimension features, which are used to capture the time cycle pattern of logistics data. b. Extract spatial dimension features. Three types of sub-features are derived from the grid codes in historical scheduling records: the administrative region to which the grid belongs, road network density, and delivery difficulty. The three types of sub-features are spliced ​​together to obtain spatial dimension features, which are used to capture the spatial correlation patterns of logistics data. c. Extract attribute dimension features, including four types of sub-features: order attributes, transportation cost attributes, road condition attributes, and weather attributes. Among them, order attributes are taken from logistics scheduling task requirements, transportation cost attributes are taken from cross-entity scheduling resource sets, road condition attributes are taken from real-time road condition data, and weather attributes are taken from real-time weather data. All sub-features are concatenated after min-max standardization to obtain attribute dimension features, which are used to capture the business attribute patterns of logistics data. d. Concatenate the three types of features into a standardized feature vector: ; in, To standardize the feature vector, Features in the time dimension Spatial dimensional features, For attribute dimension features; S24. The global aggregation node in the logistics data space receives the model inference gradients uploaded by all data participants, and aggregates them by weighted average based on the effective sample size to obtain the global model parameters. The aggregation process satisfies: ; in, This is the global parameter set for the spatiotemporal federated mining model; The total number of data subjects participating in federal computing; For the first The number of valid historical samples for each data subject within the current time and space range of the request; The total number of valid historical samples for all data subjects; For the first A set of local sub-model inference parameters for each data subject; S25. The aggregated global model parameters are distributed to the local sub-models of each data participant. Reasoning is performed based on the attributes of the targeted data request to generate candidate mining results. The candidate mining results include the predicted order volume and the predicted average delivery time for each grid in the next few hours.

4. The joint optimization method for value mining and dynamic scheduling in logistics data space according to claim 3, characterized in that, The training process of the spatiotemporal federated mining model is as follows: a. Each data participant's local sub-model undergoes iterative training based on local anonymized historical data. The training process employs a loss function that integrates scheduling benefits, specifically: ; in, The total loss of the spatiotemporal federated mining model is defined as the value to be minimized during the training phase. , , Let be the weighting coefficient, satisfying , To predict losses, the baseline accuracy used to constrain the mining results is calculated using the following formula: ; in, This represents the total number of training samples. For the first Mining prediction value for each sample, For the first The true value corresponding to each sample is taken from the actual order volume and actual delivery time data within the same spatial range in the historical scheduling records; The normalized value contributes to the scheduling in the current mining scenario, which is used to characterize the ability of mining results to improve scheduling benefits; The federal consistency loss is represented by the value of the first iteration. Local submodel parameter set for each data subject With the global model parameter set L2 distance: ; b. After each data participant completes a preset number of local iterations of training, it uploads the model gradient to the global aggregation node of the logistics data space once. The global aggregation node updates the global model parameters according to the aggregation rules in step S24 and then distributes them to each local sub-model to complete one round of federated training iteration. c. When the decrease in total loss over a preset number of iterations is less than a predetermined threshold, training is stopped, and the final spatiotemporal federated mining model is obtained. The process by which the spatiotemporal federated mining model generates mining results during the inference phase is as follows: (1) After receiving the targeted data request, extract the multi-source heterogeneous data within the time range and spatial grid range covered by the request, and generate a standardized feature vector through feature encoding in step S23; (2) Input the standardized feature vector into the trained spatiotemporal federated mining model, extract the time series periodicity of the stream data by the time gate of the spatiotemporal long short-term memory network, extract the spatial correlation of adjacent grids by the spatial gate, and obtain the predicted output value after mapping through the fully connected layer. The predicted output value is an unlabeled set of original values, which includes the original predicted values ​​of order volume and average delivery time corresponding to the mining result type of the targeted data request. (3) Format the predicted output value according to the three attributes of the targeted data request to generate candidate mining results.

5. The joint optimization method for value mining and dynamic scheduling in logistics data space according to claim 4, characterized in that, Step S3 specifically includes: S31. Receive candidate mining results from the independent inference output of the local sub-models of each data participant, and form a set of candidate mining results; S32. For each group of candidate mining results, match historical scheduling data of the same scenario; S33. Based on the scheduling contribution of historical mining results, calculate the confidence level of each group of candidate mining results. The calculation formula is as follows: ; in, The confidence level of the current candidate mining results. The weight for historical scheduling is determined by the average net utility improvement rate brought about by similar mining results in the past 30 days. The output value of the current candidate mining result generated in step S2; This represents the average of the historical true values ​​of the corresponding mining results in the historical scheduling data of the same scenario. Minimum protection threshold, used to avoid Division by zero error when the result is 0; S34. Set a confidence threshold and perform two-level filtering on the candidate mining result set: The first level filters candidate results with confidence scores below the threshold and directly removes invalid results; the second level retains the group with the highest confidence score as the valid result for the same spatiotemporal dimension unit and the same mining result type of the remaining candidate results. S35. Integrate the valid results of all spatiotemporal dimension units to generate the final mining result set.

6. The joint optimization method for value mining and dynamic scheduling in logistics data space according to claim 5, characterized in that, Step S4 specifically includes: S41. Collect three types of input data required for constructing the state space and action space. The first type is mining feature data, which is taken from the final mining result set generated in step S3 and the standardized feature vector generated in step S2. The second type is real-time operation data, including the real-time order volume, transportation resource location, real-time road condition data, and real-time weather data of each spatial grid. The third type is cost feature data, which is taken from the unit call cost in the quadruple label of the multi-source heterogeneous data in step S1 and the unit time operation cost of the cross-subject scheduling resource set. S42. Perform min-max standardization on all input features. The standardization process satisfies the following expression: ; in, The original values ​​of the input features. , The minimum and maximum values ​​of this type of feature in the historical scheduling records over the past 30 days. As the minimum protection threshold, These are the standardized feature values; S43. Concatenate the standardized three types of features into a state space vector: ; in, This represents the state space vector of a multi-agent scheduling model. To extract feature subvectors, To run feature subvectors in real time, For cost feature vectors; S44. Construct the action space of the multi-agent scheduling model. The global action space is defined by the following expression: ; in, For the global action space set, Let be the total number of schedulable capacity units, and let each schedulable capacity unit correspond to an independent intelligent agent. Action vectors of each agent For 2D vectors: ; in, This is the task acceptance decision dimension, with a value of 0 or 1. 0 indicates that the capacity rejects the currently assigned order task, and 1 indicates that the capacity accepts the currently assigned order task. A dimension is selected for the driving path, with a normalized value ranging from 0 to 1. The normalization process satisfies the following conditions: ,in, This represents the total number of alternative routes generated based on real-time traffic data. This is an index for alternative paths.

7. The joint optimization method for value mining and dynamic scheduling in logistics data space according to claim 6, characterized in that, Step S5 specifically includes: S51. Build a multi-agent deep deterministic policy gradient scheduling framework, including two stages: offline training and online inference. S52. Define the target reward function for the offline training phase. The target reward function integrates the robustness penalty for mining confidence and the operational constraint penalty, specifically as follows: ; in, for The reward value corresponding to the scheduling decision at each moment is used to iterate the model parameters during the offline training phase with the goal of maximizing the cumulative reward value. for The total net utility across the entire link corresponding to the time-sharing decision. This is the robustness penalty coefficient, used to control the severity of the penalty for deviations in the mining results. This is a predicted value based on the excavation results from the previous moment. This represents the actual value collected after the previous scheduling execution. This represents the confidence level label for the mining result at the previous time step. This is the operational constraint penalty coefficient, used to control the severity of penalties for violating operational rules; The operational constraint penalty value is calculated as follows: ; in, This represents the total number of schedulable transport capacity sets. For the first The cumulative working hours of each transport capacity This is the maximum daily working hours. The total number of orders to be scheduled. For the first The estimated delivery time for each order is calculated based on real-time traffic data. For the first Promised delivery time for each order; S53. Initial scheduling decision is generated during the online inference stage: The state space vector constructed in step S4 is input into the trained multi-agent scheduling model. The model outputs a corresponding 2D action vector for each capacity agent independently. The action vectors of all agents are combined to form the initial scheduling decision. S54. Perform global constraint verification on the initial scheduling decision. The verification rules must satisfy: ; in, For the first The capacity of the first transport to the first Decision-making for receiving individual orders; S55. Execute quantitative guaranteed minimum scheduling trigger judgment and emergency handling, trigger conditions are met: ; in, The number of scheduling cycles for continuous monitoring. For the first The average confidence level of the mining results over a scheduling cycle The confidence threshold; When the triggering condition is met, the backup scheduling mode is triggered, and the following operations are performed: a. Reconstruct the state space vector, setting all mined feature vectors to 0, retaining only the real-time running feature vector and the cost feature vector. The reconstructed state space satisfies: ; b. Route selection directly adopts the alternative route with the shortest travel time under real-time traffic conditions, without considering the prediction of future traffic conditions; c. During the scheduling process, a preset amount of idle transportation capacity is always reserved as an emergency reserve to deal with unexpected sudden increases in order volume or abnormal road conditions; S56. The scheduling decision in the output verification passed or the backup scheduling mode is taken as the global optimal scheduling decision and executed.

8. The joint optimization method for value mining and dynamic scheduling in logistics data space according to claim 7, characterized in that, Step S6 specifically involves: collecting actual operational data after the scheduling decision in step S5 is executed, and calculating the actual total net utility of the current scheduling. And based on actual total net utility The overall actual value contribution of the mining results is calculated using the following formula: ; in, This contributes to the overall value of the data collected during this excavation. The average net scheduling utility of the baseline without using mining results in the same scenario is taken from historical scheduling records; Then, the marginal value contribution of a single-class spatiotemporal attribute feature is calculated, satisfying the following expression: ; in, For the first The actual value contribution of spatiotemporal attribute features Corresponding to time dimension features, Corresponding spatial dimensional features, Corresponding attribute dimension features; To make the first After setting all spatiotemporal attribute features to 0, the simulated net utility is obtained using the same scheduling model and real-time data inference.

9. The joint optimization method for value mining and dynamic scheduling in logistics data space according to claim 8, characterized in that, In step S7, the online parameter adjustment specifically involves adjusting the input weights of the three types of features in the spatiotemporal federated mining model. The adjustment rules are as follows: ; in, For the adjusted number The input weights for the class-mining features are uniformly distributed initially. To adjust the previous number The current weights of the class-mined features. This is the online learning rate coefficient, to avoid excessive weight fluctuations affecting the stability of the mining process; The offline iteration cycle of the model is during the daily off-peak logistics period, updating the core parameters of the spatiotemporal federated mining model and the multi-agent scheduling model respectively, including: (1) Update the weights of the scheduling revenue term in the loss function of the spatiotemporal federated mining model, satisfying the following expression: ; in, The updated scheduling benefit item weights, The weights of the scheduling benefit items before the update. For the corresponding offline update coefficient, The average value contribution to the overall mining value across all scheduling cycles on that day; (2) Update the robustness penalty coefficient of the reward function of the multi-agent scheduling model to satisfy the following expression: ; in, The updated robustness penalty coefficient, The robustness penalty coefficient before the update. For the corresponding offline update coefficient, The average confidence level of all excavation results for that day; After the iteration is completed, the updated parameters are sent to the spatiotemporal federated mining model and the multi-agent scheduling model respectively, for the next round of mining inference and scheduling decision-making, thus completing the bidirectional joint optimization closed loop of mining and scheduling.